Dynamic Programming III: LIS, LCS & Edit Distance

This deck sets up three classic sequence dynamic programs from scratch: Longest Increasing Subsequence, Longest Common Subsequence, and Edit Distance. It emphasizes defining the state in words, choosing the correct base row and column, writing the match-or-mismatch recurrence, and tracing the grid to read back the actual answer. It targets confusing a subsequence with a substring, mis-initializing the Edit Distance base case to zero, tracing the answer back from the wrong cell or in the wrong direction, and thinking that LIS requires contiguous elements.

Subject: CS3000 Algorithms · 131 slides · symbolic lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. What you will be able to do

Objectives

Today you'll build three related dynamic-programming algorithms that all compare or order sequences. By the end you can:

  1. Explain what a subsequence is, and why it is not the same as a substring.
  2. Set up the Longest Increasing Subsequence (LIS) DP: define the state in words, write its recurrence, and trace it on a small array.
  1. Set up the Longest Common Subsequence (LCS) DP on two strings: the state, the base row/column, and the match/mismatch recurrence.
  2. Set up the Edit Distance DP: its base row/column, its insert/delete/replace recurrence, and how it differs from LCS's base case.
  1. Trace back through any of these three DP tables to read out the actual answer - not just its length or cost.

2. What survived from Dynamic Programming II: Knapsack, Coin Change & Rod Cutting?

Warm-up

Discussion prompt

Before we open Dynamic Programming III: LIS, LCS & Edit Distance: without looking back, what was the main idea of Dynamic Programming II: Knapsack, Coin Change & Rod Cutting, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

0/1 knapsack, minimum-coin and combination-counting coin change, and rod cutting, taught with the setup made explicit every time: define the state in words, write the take-or-skip (or best-choice) recurrence, pin the base case, fill the table in dependency order, and reconstruct the actual items, coins, or cuts chosen. Targets four real misconceptions: using a best-ratio greedy strategy on 0/1 knapsack, accidentally reusing a single-copy item, using a greedy coin heuristic on a coin system where it overshoots the true minimum, and an off-by-one in the table's capacity or amount dimension.

3. Toolkit check-in: name them before you look

Concept

Before any new material: cover the screen.

You have named 13 reusable moves so far. Say as many as you can out loud, by number, from memory.

Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.

Here they are. Score yourself.

Today adds no new moves. Every proof in this lesson is built out of the list above. That is the whole point of the list.

The question that starts every proof from here on is not how do I begin. It is which of these applies here?

4. Break it if you can: Toolkit check-in: name them before you look

Counterexample

Discussion prompt

You have named 13 reusable moves so far. Say as many as you can out loud, by number, from memory.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Do not advance until you have actually tried. Getting four of eight is information; skipping the exercise is not.

5. Recap: what makes a problem dynamic programming

Concept

A problem is a good fit for dynamic programming when it has two properties: overlapping subproblems (the same smaller question gets asked over and over) and optimal substructure (the best answer to the big question is built from the best answers to smaller ones).

Instead of recomputing a subproblem every time it comes up, we solve each one once, store it in a table, and reuse it. All three algorithms in this lesson are exactly that: a table of small answers that builds up to one big answer.

6. What a subsequence is

Concept

A subsequence of a list is what's left after you delete zero or more elements, without changing the order of what remains. The elements you keep do not need to be next to each other.

subsequence — A sequence obtained by deleting some (possibly zero) elements from another sequence, without reordering the ones that remain. For example, "ACE" is a subsequence of "ABCDE": keep positions 1, 3, and 5, and drop the rest.

7. By analogy: What a subsequence is

Analogy

Discussion prompt

Explain What a subsequence is by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A subsequence of a list is what's left after you delete zero or more elements, without changing the order of what remains. The elements you keep do not need to be next to each other.

8. Subsequence is not substring

Picture it

Animation

Shows: Subsequence is not substring — a rendered Manim animation.

Rendered with Manim.

Takeaway: LCS allows gaps, which is why a table beats a scan.

9. Something is wrong here: subsequence vs substring

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student is asked whether "ACE" is a subsequence of "ABCDE", and instead checks whether "ACE" appears as a contiguous block of letters inside "ABCDE".

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This tests for a substring - a contiguous chunk - which is a stricter, different requirement than a subsequence.

Check the subsequence condition instead: can you find A, then C, then E, each appearing later in the string than the last one, without requiring them to touch?

Why: This tests for a substring - a contiguous chunk - which is a stricter, different requirement than a subsequence.

10. Trap: subsequence vs substring

Trap

The trap

A student is asked whether "ACE" is a subsequence of "ABCDE", and instead checks whether "ACE" appears as a contiguous block of letters inside "ABCDE".

Search for "ACE" as a run of adjacent letters

Why: This tests for a substring - a contiguous chunk - which is a stricter, different requirement than a subsequence.

Conclude "ACE" is not there

Why: "ABCDE" contains no contiguous run of letters spelling "ACE" (the letters B and D sit between them), so a substring search wrongly reports failure.

The fix

Check the subsequence condition instead: can you find A, then C, then E, each appearing later in the string than the last one, without requiring them to touch?

Scan left to right, matching one target letter at a time

Why: A is at position 1, C is at position 3 (after A), and E is at position 5 (after C). Order is preserved even though B and D are skipped.

Conclude "ACE" is a valid subsequence

Why: Every letter of "ACE" was found in order, just not adjacently. Substrings must be contiguous; subsequences only need to preserve order.

11. Why this distinction matters for today's three DPs

Concept

All three algorithms today - Longest Increasing Subsequence, Longest Common Subsequence, and Edit Distance - reason about subsequences, not substrings. Every recurrence you're about to see is built on the idea of skipping elements freely while keeping order.

If you accidentally think in terms of contiguous runs, the recurrences will look wrong and the trace tables won't match what the DP is actually counting.

12. Teach it back: Why this distinction matters for today's three DPs

Explain it

Discussion prompt

Explain Why this distinction matters for today's three DPs to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

If you accidentally think in terms of contiguous runs, the recurrences will look wrong and the trace tables won't match what the DP is actually counting.

13. Where these three DPs show up

Concept

Longest Increasing Subsequence: finding the longest stretch of improving values buried in noisy data, like the best run of increasing days in a stock's price history.

Longest Common Subsequence: the basis of file-diff tools (what changed between two versions of a file) and comparing DNA sequences for similarity.

Edit Distance: spell-checkers and search-suggestion systems use it to measure how many keystrokes separate what you typed from a real word.

14. The Longest Increasing Subsequence problem

Concept

You're given a list of numbers. You want the longest subsequence whose values are strictly increasing from left to right.

\[ a_1, a_2, \ldots, a_n \]

Symbols: a is the array; a with a subscript is one of its numbers, and n is how many numbers there are, read in the given order.

For example, in the list 5, 2, 8, 6, 3, 6, 9, 7 one increasing subsequence is 2, 3, 6, 9 - each later value is bigger than the one before it, even though they are not next to each other in the original list.

15. Why checking every subsequence is too slow

Intuition

A list of n numbers has an enormous number of subsequences - each element is either kept or dropped, independently of every other. Checking every single one for being increasing, then keeping the longest, blows up impossibly fast as n grows.

\[ 2^{n} \text{ subsequences to check} \]

16. Why picking greedily doesn't work

Intuition

A tempting shortcut: scan left to right and greedily keep a number whenever it's bigger than the last one you kept. This fails, because an early greedy pick can block a much longer chain later.

For example, keeping 5 first (since it's the very first number) can lock you out of the longer chain 2, 3, 6, 9, which only works by skipping 5 entirely. You cannot know which number to commit to without looking ahead.

17. The fix: track the best subsequence ending at each spot

Intuition

Instead of guessing which numbers to commit to, answer a smaller question for every position: if my increasing subsequence has to end exactly at this number, how long can it be?

Once you know that answer for every earlier position, extending to the current position is easy: look back at every smaller, earlier number and build onto the best one.

18. LIS state: dp[i] in words

Concept

Define one small answer per position. Let i be a position in the array, and let dp[i] be the length of the longest increasing subsequence that ends exactly at position i - it must use a with subscript i as its last element.

\[ \text{dp}[i] = \text{length of the longest increasing subsequence ending at position } i \]

This is the single most important setup choice: dp[i] is not 'the best answer using the first i numbers' - it is pinned to end exactly at position i.

19. The DP Setup skeleton

Concept

Every proof of this kind has the same five or six moves in the same order. The order is not something you rediscover each time.

It is on the right. It will stay on the right through the worked examples that follow.

Why this matters: the structure is now handled. You are not spending working memory on what comes next — you are spending all of it on the one hard step.

Step 1 is where every failed DP fails. A subproblem you cannot say in one sentence is a subproblem you cannot write a recurrence for, and steps 2 through 6 will not rescue it.

20. What feels wrong about the other definition?

Intuition

What feels wrong about this?

Two candidate sentences for the LIS table entry:

\[ (a) \;\; dp[i] = \text{longest increasing subsequence ending exactly at } i \]

\[ (b) \;\; dp[i] = \text{longest increasing subsequence among the first } i \text{ numbers} \]

Definition (b) answers the question you actually care about. Definition (a) does not.

_Plain English only. No notation, no algebra. Just say what bothers you._

The feeling: definition (b) tells you a length but not what the sequence ends with, and without that you cannot tell whether the next number is allowed to extend it.

That feeling is the proof. It is not a substitute for the proof — it is the thing the proof writes down.

This is the same trap as Kadane's best-ending-here. The entry that answers the final question directly is usually the one you cannot recurse on — you want the entry that carries the fact the next step needs.

21. Something is wrong here: dp[i] is not "best subsequence among the first i…

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student defines dp[i] as "the length of the longest increasing subsequence using any of the first i numbers," and expects dp[i] to always be at least as large as dp[i-1].

It is wrong. Say what breaks — and say it before you turn the page.

Correct: If dp[i] means 'best among the first i,' then having one more candidate number can never make the best answer worse.

dp[i] is pinned to subsequences that end exactly at position i. It can go up or down compared to dp[i-1]; there's no rule that dp[2] has to be at least dp[1].

Why: If dp[i] means 'best among the first i,' then having one more candidate number can never make the best answer worse.

22. Trap: dp[i] is not "best subsequence among the first i numbers"

Trap

The trap

A student defines dp[i] as "the length of the longest increasing subsequence using any of the first i numbers," and expects dp[i] to always be at least as large as dp[i-1].

Assume dp[i] only ever grows as i increases

Why: If dp[i] means 'best among the first i,' then having one more candidate number can never make the best answer worse.

Try to fill dp[2] this way for the array 5, 2, 8, ...

Why: Under this wrong definition, dp[2] would just copy dp[1] whenever a with subscript 2 doesn't extend anything, blurring together answers that actually end at different positions.

The fix

dp[i] is pinned to subsequences that end exactly at position i. It can go up or down compared to dp[i-1]; there's no rule that dp[2] has to be at least dp[1].

Check the real values for 5, 2, 8, ...

Why: dp[1] = 1 (just the number 5, alone). dp[2] = 1 too (just the number 2, alone) - not because it 'copies' dp[1], but because no earlier number is smaller than 2.

Keep the two questions separate

Why: The answer to the WHOLE problem is the maximum over all dp[i]; each individual dp[i] only answers the narrower 'ending here' question.

23. Decision point: the LIS entry is defined. Get the recurrence.

Intuition

What move should we make next?

The sentence, fixed:

\[ dp[i] = \text{length of the longest increasing subsequence ending exactly at index } i \]

The last element of that subsequence is forced — it is the number at index i.

So what is the last decision? Careful: unlike the stairs, this one does not have two cases.

_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.

24. LIS recurrence: extend from a smaller, earlier best

Concept

To build a subsequence ending at position i, look at every earlier position j whose value is smaller than a with subscript i. Any increasing subsequence ending at j can be extended by tacking a with subscript i onto the end.

\[ \text{dp}[i] = 1 + \max_{\substack{j < i \\ a_j < a_i}} \text{dp}[j] \]

In words: dp[i] is one more than the best dp[j] among all earlier positions j whose value is strictly smaller than a with subscript i.

25. Complete the line: A case split with many branches, not two

Fill the middle

Fill in the blanks

From A case split with many branches, not two — finish the line. Write what belongs on the right of the equals sign before you look.

dp[i] = 1 + \max\{\, dp[j] \;:\; j < i \text{ and } A[j] < A[i] \,\}

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Any earlier index j whose value is smaller is a legal predecessor.

26. A case split with many branches, not two

Concept

The move: #12 (Case-split on the last decision), then #13 (Name the subproblem).

The last decision is which element came immediately before index i

Why: Any earlier index j whose value is smaller is a legal predecessor. That is not two cases — it is up to i cases, and you take the best.

\[ dp[i] = 1 + \max\{\, dp[j] \;:\; j < i \text{ and } A[j] < A[i] \,\} \]

The base case is the empty maximum

Why: If no earlier element is smaller, the maximum is over an empty set and dp[i] is just 1 — the element standing alone.

The answer is not dp of n

Why: Because entries mean ending exactly here, the answer is the largest entry anywhere in the table. Forgetting this is the single most common LIS error.

Case splits are not always binary. Take-or-skip had two branches; this one has as many branches as there are legal predecessors. The move is the same either way: split on the last decision, take the best.

27. Decode the notation: A case split with many branches, not two

Notation

Annotate

From A case split with many branches, not two — read this one piece at a time. What is each part doing?

On: \( dp[i] = 1 + \max\{\, dp[j] \;:\; j < i \text{ and } A[j] < A[i] \,\} \)

  • Any earlier index j whose value is smaller is a legal predecessor. That is not two cases — it is up to i cases, and you take the best.
  • If no earlier element is smaller, the maximum is over an empty set and dp[i] is just 1 — the element standing alone.
  • Because entries mean ending exactly here, the answer is the largest entry anywhere in the table. Forgetting this is the single most common LIS error.

28. LIS base case: every number is a subsequence of length 1 by itself

Concept

If no earlier number is smaller than a with subscript i, position i cannot extend anything - it starts its own one-element chain.

\[ \text{dp}[i] = 1 \quad \text{if no } j < i \text{ has } a_j < a_i \]

This is really the same rule as the general recurrence: when there is no valid j to extend, the 'best dp[j]' contributes nothing, so dp[i] falls back to 1.

29. What has to happen first: Worked example: LIS - start filling the table

Ranking

Put in order

Put the moves of Worked example: LIS - start filling the table into the order they have to happen.

  1. Compute dp[1]
  2. Compute dp[2]
  3. Compute dp[3]
  4. Compute dp[4]
  5. Check these four values by hand against the array

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Position 1 has nothing before it, so it starts its own chain of length 1.

30. Worked example: LIS - start filling the table

Worked example

Find the longest increasing subsequence of this array.

\[ a:\quad 5,\ 2,\ 8,\ 6,\ 3,\ 6,\ 9,\ 7 \qquad (\text{positions } 1 \text{ to } 8) \]

Compute dp[1]

Why: Position 1 has nothing before it, so it starts its own chain of length 1.

\[ \text{dp}[1] = 1 \]

Compute dp[2]

Why: Value at position 2 is 2. The only earlier value is 5, which is not smaller than 2, so position 2 cannot extend anything.

\[ \text{dp}[2] = 1 \]

Compute dp[3]

Why: Value at position 3 is 8. Both earlier values (5 and 2) are smaller than 8, so position 3 can extend either chain. The better of dp[1]=1 and dp[2]=1 is 1, plus one more for position 3 itself.

\[ \text{dp}[3] = 1 + \max(\text{dp}[1], \text{dp}[2]) = 1 + 1 = 2 \]

Compute dp[4]

Why: Value at position 4 is 6. Earlier smaller values are at positions 1 (value 5, dp=1) and 2 (value 2, dp=1); position 3's value 8 is excluded since it is not smaller than 6.

\[ \text{dp}[4] = 1 + \max(\text{dp}[1], \text{dp}[2]) = 1 + 1 = 2 \]

Check these four values by hand against the array

Why: dp[3]=2 should reflect a genuine 2-element increasing run ending at 8, such as 5 then 8, or 2 then 8 - both are real runs, confirming the value.

ivaluedp[i]
151
221
382
462

31. Worked example: LIS - finish the table and find the answer

Worked example

Compute dp[5]

Why: Value at position 5 is 3. The only earlier smaller value is at position 2 (value 2, dp=1); positions with values 5, 8, and 6 are all bigger than 3, so they don't count.

\[ \text{dp}[5] = 1 + \text{dp}[2] = 1 + 1 = 2 \]

Compute dp[6]

Why: Value at position 6 is 6. Earlier smaller values are at positions 1 (dp=1), 2 (dp=1), and 5 (dp=2); positions with values 8 and 6 are excluded (8 is bigger, and 6 is not strictly smaller than 6). The best of these is dp[5]=2.

\[ \text{dp}[6] = 1 + \max(\text{dp}[1], \text{dp}[2], \text{dp}[5]) = 1 + 2 = 3 \]

Compute dp[7]

Why: Value at position 7 is 9. Every earlier value (5, 2, 8, 6, 3, 6) is smaller than 9, so all six dp values are candidates. The largest of them is dp[6]=3.

\[ \text{dp}[7] = 1 + \max(\text{dp}[1..6]) = 1 + 3 = 4 \]

Compute dp[8]

Why: Value at position 8 is 7. Earlier smaller values are at positions 1, 2, 4, 5, and 6 (position 3's value 8 is excluded, and position 7 comes after). The best among those dp values is dp[6]=3.

\[ \text{dp}[8] = 1 + \max(\text{dp}[1], \text{dp}[2], \text{dp}[4], \text{dp}[5], \text{dp}[6]) = 1 + 3 = 4 \]

Verify the answer is the maximum entry in the table

Why: The largest dp value in the completed table is 4, reached at both position 7 and position 8, so the longest increasing subsequence has length 4.

ivaluedp[i]
151
221
382
462
532
663
794
874

32. The table holds the answer; traceback holds the solution

Picture it

Animation

Shows: The table holds the answer; traceback holds the solution — a rendered Manim animation.

Rendered with Manim.

Takeaway: Most exam questions want the second, and it costs one extra pass.

33. Plan first: Worked example: LIS - read back the actual subsequence

Step zero

Discussion prompt

Worked example: LIS - read back the actual subsequence — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Start at position 7 (dp=4, value 9)

Answer:

  1. Start at position 7 (dp=4, value 9)
  2. Find which earlier dp value produced dp[7]
  3. Continue back from position 6
  4. Continue back from position 5
  5. Verify the reconstructed subsequence is strictly increasing

34. Worked example: LIS - read back the actual subsequence

Worked example

The dp array tells us the best length is 4, but not which numbers form it. Trace backward from a position that achieved the max.

Start at position 7 (dp=4, value 9)

Why: This is one of the positions achieving the maximum; any increasing subsequence achieving the max is a valid longest one.

Find which earlier dp value produced dp[7]

Why: dp[7] = 1 + dp[6], and dp[6]=3 is the value that was used, so the chain continues at position 6 (value 6).

Continue back from position 6

Why: dp[6] = 1 + dp[5], and dp[5]=2 is the value that was used, so the chain continues at position 5 (value 3).

Continue back from position 5

Why: dp[5] = 1 + dp[2], and dp[2]=1 is the value that was used, so the chain continues at position 2 (value 2), which is a base case with no predecessor.

Verify the reconstructed subsequence is strictly increasing

Why: Reading the chain forward: position 2 (value 2), position 5 (value 3), position 6 (value 6), position 7 (value 9). Each value is bigger than the last one, confirming a genuine length-4 increasing subsequence.

\[ 2 < 3 < 6 < 9\ \checkmark \]

35. Work backwards from the answer: Worked example: LIS - read back the actual…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Verify the reconstructed subsequence is strictly increasing

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The dp array tells us the best length is 4, but not which numbers form it. Trace backward from a position that achieved the max.

36. LIS answers aren't always unique

Concept

The dp table only guarantees the correct length. Several different subsequences can share that same longest length, and the table doesn't prefer one over another.

In our array, position 8 (value 7) also has dp[8] = 4, giving a second valid answer: 2, 3, 6, 7. Both this and 2, 3, 6, 9 are correct longest increasing subsequences.

37. Something is wrong here: LIS elements do not need to be contiguous

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student sees "increasing subsequence" and only checks consecutive runs of the array, sliding a window and testing whether each window is increasing.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This checks contiguous stretches only - a fundamentally different, smaller search space than all subsequences.

Elements of an increasing subsequence can skip over any number of positions, as long as their values still increase left to right.

Why: This checks contiguous stretches only - a fundamentally different, smaller search space than all subsequences.

38. Trap: LIS elements do not need to be contiguous

Trap

The trap

A student sees "increasing subsequence" and only checks consecutive runs of the array, sliding a window and testing whether each window is increasing.

Scan for the longest run of adjacent increasing numbers

Why: This checks contiguous stretches only - a fundamentally different, smaller search space than all subsequences.

Report length 2 for 5, 2, 8, 6, 3, 6, 9, 7

Why: The longest adjacent increasing run is just 2 numbers long (for example 2 then 8, or 3 then 6, or 6 then 9) - this badly undercounts the true answer.

The fix

Elements of an increasing subsequence can skip over any number of positions, as long as their values still increase left to right.

Allow skipped positions between chosen elements

Why: 2 (position 2), 3 (position 5), 6 (position 6), 9 (position 7) skips positions 1, 3, and 4 entirely, and is still perfectly valid.

Report the true length: 4

Why: Matching the dp table's answer confirms that dropping the 'adjacent' requirement is exactly what the problem calls for.

39. Break it on purpose: LIS elements do not need to be contiguous

Break the constraint

Discussion prompt

The rule this trap just fixed:

Matching the dp table's answer confirms that dropping the 'adjacent' requirement is exactly what the problem calls for.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This checks contiguous stretches only - a fundamentally different, smaller search space than all subsequences.

40. LIS time complexity

Concept

Filling dp[i] requires scanning every earlier position j to look for smaller values. Doing this for every position i gives a nested loop over pairs of positions.

\[ O(n^{2}) \]

A cleverer method using binary search can bring this down to n log n, but the quadratic version is the one to master first, since it directly matches the recurrence.

41. Cost of the sequence DPs

Picture it

Animation

Shows: Cost of the sequence DPs — a rendered Manim animation.

Rendered with Manim.

Takeaway: Quadratic in the input lengths, and for long DNA that matters.

42. The DP setup checklist

Pattern

1. Define the state in plain words first

Why: Say exactly what smaller question each table entry answers, before writing any formula - for example, 'the longest increasing subsequence ending exactly at position i.'

2. Write the match/mismatch or extend/don't-extend recurrence

Why: Identify the handful of ways a bigger answer can be built from smaller ones, and write one formula per case.

3. Nail down the base row, column, or case

Why: Decide the smallest inputs directly - don't assume they're zero. Some bases really are zero (LCS); others count something real (Edit Distance).

4. Know how you will trace back the actual answer

Why: A number alone (a length or a cost) isn't the full answer. Decide, before you code, which direction you'll walk through the table to recover the real sequence.

43. Where this shows up: Dynamic Programming III: LIS, LCS & Edit Distance

Real world

Discussion prompt

Outside this lesson: where does Dynamic Programming III: LIS, LCS & Edit Distance actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of The DP setup checklist is doing the work in it.

Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.

Answer:

Setting up three classic sequence DPs from scratch: Longest Increasing Subsequence, Longest Common Subsequence, and Edit Distance. Emphasizes defining the state in words, choosing the correct base row/column, writing the match/mismatch recurrence, and tracing the grid to read back the actual answer.

44. Rule out three: Check yourself: computing an LIS entry

Elimination

Eliminate the wrong options

What is dp[5], the length of the longest increasing subsequence ending at position 5 (value 5)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1
  • B. 2
  • C. 3
  • D. 5

Survives elimination: C

Why: dp[5] extends from the best earlier position whose value is smaller than 5. All four earlier positions (3, 1, 4, 1) qualify, with dp values 1, 1, 2, 1. The largest is dp[3]=2 (the chain 3 then 4, or 1 then 4), so dp[5] = 1 + 2 = 3, giving a chain like 1, 4, 5 or 3, 4, 5.

45. Check yourself: computing an LIS entry

Check

Consider the array below.

\[ a:\quad 3,\ 1,\ 4,\ 1,\ 5 \qquad (\text{positions } 1 \text{ to } 5) \]

Check your understanding

What is dp[5], the length of the longest increasing subsequence ending at position 5 (value 5)?

  • A. 1
  • B. 2
  • C. 3 (correct)
  • D. 5

Answer: C

Why: dp[5] extends from the best earlier position whose value is smaller than 5. All four earlier positions (3, 1, 4, 1) qualify, with dp values 1, 1, 2, 1. The largest is dp[3]=2 (the chain 3 then 4, or 1 then 4), so dp[5] = 1 + 2 = 3, giving a chain like 1, 4, 5 or 3, 4, 5.

Why A tempts people
This assumes dp[5] resets to 1 because the immediately preceding element (value 1) is smaller than 5 but doesn't extend anything useful - it ignores the other three earlier positions that also qualify and have a better dp value.
Why B tempts people
This extends from a valid smaller earlier element (such as position 1, value 3, with dp = 1) but doesn't take the maximum dp value among ALL qualifying earlier positions - position 3 (value 4) has the best dp value, 2, not 1.
Why D tempts people
This confuses the array value at position 5 (which is 5) with the subsequence length dp[5]. The recurrence tracks how many elements are chained together, not the numeric value stored at that position.

46. Answer it before you see the options: Check yourself: the LIS recurrence…

Prediction

Predict first

Which statement correctly completes the LIS recurrence: dp[i] equals...

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 1 plus the maximum dp[j] among all j less than i with value at j less than value at i (or 1 if no such j exists).

Why: The recurrence must scan every earlier position (not just the immediate neighbor), require a strictly smaller value to preserve the increasing property, and add 1 to account for including position i itself. Any looser or tighter version breaks on some input.

47. Check yourself: the LIS recurrence itself

Check

Recall the setup: dp[i] is the length of the longest increasing subsequence ending exactly at position i.

Check your understanding

Which statement correctly completes the LIS recurrence: dp[i] equals...

  • A. 1 plus the maximum dp[j] among all j less than i with value at j less than value at i (or 1 if no such j exists). (correct)
  • B. 1 plus dp[i-1], provided the value at i-1 is less than the value at i.
  • C. 1 plus the maximum dp[j] among all j less than i with value at j less than or equal to value at i.
  • D. The maximum dp[j] among all j less than i with value at j less than value at i, without adding 1.

Answer: A

Why: The recurrence must scan every earlier position (not just the immediate neighbor), require a strictly smaller value to preserve the increasing property, and add 1 to account for including position i itself. Any looser or tighter version breaks on some input.

Why B tempts people
This only checks the single position right before i, missing cases where the best extension comes from farther back - like extending from position 5 to position 7 while skipping position 6.
Why C tempts people
Using less-than-or-equal-to would let equal values chain together, producing a non-decreasing subsequence instead of a strictly increasing one.
Why D tempts people
Forgetting the plus 1 would report the length of the subsequence BEFORE adding the current element, undercounting every answer by one.

48. The Longest Common Subsequence problem

Concept

You're given two sequences (think of them as strings). You want the length - and eventually the actual letters - of the longest subsequence that appears in BOTH of them, in order, in each.

\[ X = x_1 x_2 \cdots x_m, \qquad Y = y_1 y_2 \cdots y_n \]

Symbols: X and Y are the two strings; m and n are their lengths. We compare a prefix of X against a prefix of Y and grow both prefixes.

49. Build up from smaller prefixes of both strings

Intuition

Instead of comparing the whole strings at once, ask a smaller question: what is the longest common subsequence of just the first i characters of X and the first j characters of Y?

Grow i and j one step at a time. Once you know the answer for every smaller pair of prefixes, the answer for the full strings falls out of the same rule applied one more time.

50. LCS state: dp[i][j] in words

Concept

Let i range over prefix lengths of X (from 0 to m) and j range over prefix lengths of Y (from 0 to n). Define dp[i][j] as the length of the longest common subsequence between the first i characters of X and the first j characters of Y.

\[ \text{dp}[i][j] = \text{LCS length of } x_1 \cdots x_i \text{ and } y_1 \cdots y_j \]

Notice this state takes two indices, one per string - unlike LIS, which only needed one, because here we're tracking progress through two sequences at once.

51. LCS is what diff computes

Picture it

Animation

Shows: LCS is what diff computes — a rendered Manim animation.

Rendered with Manim.

Takeaway: Which is why a good diff highlights so little.

52. Table size: (m+1) rows by (n+1) columns

Concept

Because i and j both range starting from 0 (the empty prefix) up through the full length, the table needs one extra row and one extra column beyond the string lengths.

\[ \text{table size: } (m+1) \times (n+1) \]

Forgetting that extra row and column is a common setup mistake - row 0 and column 0 represent comparing against an empty string, and they must exist before you can fill in row 1 or column 1.

53. LCS base case: comparing against an empty string

Concept

Row 0 represents using zero characters of X; column 0 represents using zero characters of Y. The longest common subsequence with an empty string is always empty.

\[ \text{dp}[0][j] = 0 \ \text{ for all } j, \qquad \text{dp}[i][0] = 0 \ \text{ for all } i \]

So the entire top row and entire left column of the table are zero - there is nothing to match yet.

54. LCS recurrence: when the characters match

Concept

If the i-th character of X equals the j-th character of Y, that shared character can be the new last character of a common subsequence - extend the best answer from one row up and one column left.

\[ x_i = y_j \ \Rightarrow \ \text{dp}[i][j] = \text{dp}[i-1][j-1] + 1 \]

This is the only case where the answer grows: matching characters are the entire reason LCS length increases.

55. The two cases of LCS

Picture it

Animation

Shows: The two cases of LCS — a rendered Manim animation.

Rendered with Manim.

Takeaway: Match walks diagonally; mismatch drops one character from one string.

56. Decision point: the last two characters do not match

Intuition

What move should we make next?

Comparing prefixes of two strings, and the characters at the ends differ:

\[ dp[i][j] \;\text{ with }\; X[i] \ne Y[j] \]

They cannot both be in the common subsequence, since the subsequence would have to end with both at once.

If at least one of them has to go, how many cases is that, and what do you do with the results?

_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.

57. LCS recurrence: when the characters don't match

Concept

If the two characters differ, this pair of positions itself contributes nothing new. The best you can do is fall back to whichever neighbor - dropping the current character of X, or dropping the current character of Y - already has the better answer.

\[ x_i \neq y_j \ \Rightarrow \ \text{dp}[i][j] = \max\big(\text{dp}[i-1][j],\ \text{dp}[i][j-1]\big) \]

No character is added in this case - you're just copying forward the better of two smaller answers you already computed.

58. Fill order: why we sweep row by row

Concept

Every entry dp[i][j] depends only on entries with a smaller row (i-1) or, within the same row, an earlier column (j-1). So if you fill row 0 first, then row 1 left to right, then row 2 left to right, and so on, every value you need is already sitting in the table when you need it.

This fill order isn't just convenient bookkeeping - it's forced by the recurrence itself: you can never compute a cell before the cells it depends on exist.

59. Longest common subsequence in pseudocode

Concept

One table, one comparison, two branches. Every cell looks at three neighbours and never anywhere else, which is what makes the row-by-row sweep safe.

LCS(X, Y)
  for i = 0 to X.length
    dp[i][0] = 0
  for j = 0 to Y.length
    dp[0][j] = 0
  for i = 1 to X.length
    for j = 1 to Y.length
      if X[i] == Y[j]
        dp[i][j] = dp[i-1][j-1] + 1
      else
        dp[i][j] = max(dp[i-1][j], dp[i][j-1])
  return dp[X.length][Y.length]

On a match you move diagonally and add one. On a mismatch you do not add anything — you just take the better of dropping one character from either string. The diagonal is the only direction that ever grows the answer.

60. Reading LCS line by line

Notation

Every line of LCS says one thing. Read the line, then read what it does — not the other way round.

Annotate

  • The zero row and zero column: matching anything against an empty string gives an empty subsequence.
  • The one comparison. Everything else in this procedure is bookkeeping around this single test.
  • Match: the diagonal neighbour is the answer for both strings one character shorter. Add the character you just matched.
  • Mismatch: try dropping the last character of X, or of Y, and keep whichever did better. No plus one — nothing was matched.
  • Every cell reads only up, left, and up-left. All three are already filled when the sweep reaches it, which is the whole reason the loop order works.
  • The bottom-right corner: both strings fully considered.

61. Step LCS yourself

Invariant

Each cell holds the length of the longest common subsequence of the first i characters of X and the first j characters of Y. Say that before reading any number.

Step through it

For each cell, say whether it came from the diagonal or from a neighbour, and why.

  1. Line 3: against an empty string, always zero
  2. Line 5: and the same the other way
  3. Line 8: A against B: no match
  4. Line 11: best neighbour is 0
  5. Line 8: A against A: match
  6. Line 9: diagonal plus one
  7. Line 8: B against B: match
  8. Line 9: diagonal plus one again
  9. Line 11: B against A: no match, so take the better neighbour
  10. Line 12: the longest common subsequence has length 1

62. Diagonal on a match, sideways on a mismatch

Picture it

Animation

Shows: LCS executing: the current line of pseudocode is highlighted while the data it touches changes.

Rendered with Manim.

Takeaway: Match moves diagonally and adds one; mismatch takes the better neighbour and adds nothing.

63. Guess the shape of the answer: Worked example: LCS - build the table for…

Estimation

Predict first

Find the longest common subsequence of X = "AC" (m=2 characters) and Y = "GAC" (n=3 characters). Compare prefixes of each, growing the table one row at a time.

Commit before you compute: what does Worked example: LCS - build the table for X="AC", Y="GAC" come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Check the bottom-right entry against the rules used

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. dp[2][3] = 2 was built purely from the match/mismatch rules above with no shortcuts, so the table is internally consistent before we read off the answer.

64. Worked example: LCS - build the table for X="AC", Y="GAC"

Worked example

Find the longest common subsequence of X = "AC" (m=2 characters) and Y = "GAC" (n=3 characters). Compare prefixes of each, growing the table one row at a time.

Row 0 and column 0 are all zero

Why: Comparing anything against an empty prefix always gives a common subsequence of length 0.

\[ \text{dp}[0][j] = 0, \quad \text{dp}[i][0] = 0 \]

Fill row 1 (X prefix = "A")

Why: Compare A against each of G, A, C: mismatch with G gives max(0,0)=0; match with A gives dp[0][1]+1 = 0+1 = 1; mismatch with C gives max(dp[0][3], dp[1][2]) = max(0,1) = 1.

\[ \text{dp}[1][1]=0,\ \text{dp}[1][2]=1,\ \text{dp}[1][3]=1 \]

Fill row 2 (X prefix = "AC")

Why: Compare C against G, A, C: mismatch with G gives max(dp[1][1], dp[2][0]) = max(0,0)=0; mismatch with A gives max(dp[1][2], dp[2][1]) = max(1,0)=1; match with C gives dp[1][2]+1 = 1+1 = 2.

\[ \text{dp}[2][1]=0,\ \text{dp}[2][2]=1,\ \text{dp}[2][3]=2 \]

The completed table:

"" (j=0)G (j=1)A (j=2)C (j=3)
"" (i=0)0000
A (i=1)0011
AC (i=2)0012

Check the bottom-right entry against the rules used

Why: dp[2][3] = 2 was built purely from the match/mismatch rules above with no shortcuts, so the table is internally consistent before we read off the answer.

65. The LCS table

Picture it

Animation

Shows: The LCS table — a rendered Manim animation.

Rendered with Manim.

Takeaway: A match takes the diagonal plus one. A mismatch takes the better neighbour.

66. Plan first: Worked example: LCS - read the answer and trace back the…

Step zero

Discussion prompt

Worked example: LCS - read the answer and trace back the subsequence — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Read the final answer

Answer:

  1. Read the final answer
  2. Trace back from the bottom-right corner
  3. Continue the trace from dp[1][2]
  4. Verify the recovered subsequence

67. Worked example: LCS - read the answer and trace back the subsequence

Worked example

The bottom-right entry, dp[2][3], is the length of the longest common subsequence of the full strings.

Read the final answer

Why: dp[2][3] = 2, so X="AC" and Y="GAC" share a common subsequence of length 2 - but the table alone doesn't say which two characters.

\[ \text{dp}[2][3] = 2 \]

Trace back from the bottom-right corner

Why: X's 2nd character is C and Y's 3rd character is C - they match, so this character belongs to the subsequence. A match always steps diagonally, to dp[1][2].

Continue the trace from dp[1][2]

Why: X's 1st character is A and Y's 2nd character is A - they match, so this character also belongs to the subsequence. Step diagonally again, to dp[0][1], which is a base-case zero - the trace stops.

Verify the recovered subsequence

Why: Reading the matches in the order they were found (C, then A) and reversing gives "AC" - which is exactly X itself, and it does appear in order inside Y="GAC" (at positions 2 and 3).

\[ \text{LCS}(\text{"AC"},\ \text{"GAC"}) = \text{"AC"}\ \checkmark \]

68. LCS - read the answer and trace back the subsequence — line by line

Picture it

Animation

Shows: Each line of the worked example "LCS - read the answer and trace back the subsequence", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Reading the matches in the order they were found (C, then A) and reversing gives "AC" - which is exactly X itself, and it does appear in order inside Y="GAC" (at positions 2 and 3).

69. Something is wrong here: tracing back in the wrong direction, or from the wrong…

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student wants the LCS length for the full strings X="AC", Y="GAC", and reads the value out of dp[0][0] - the corner where the table started - instead of the corner where it finished.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: dp[0][0] = 0 by the base case - it represents comparing two empty prefixes, not the full strings.

The answer to the full problem always sits in the cell that has compared the FULL length of both strings - the bottom-right corner, dp[m][n].

Why: dp[0][0] = 0 by the base case - it represents comparing two empty prefixes, not the full strings.

70. Trap: tracing back in the wrong direction, or from the wrong cell

Trap

The trap

A student wants the LCS length for the full strings X="AC", Y="GAC", and reads the value out of dp[0][0] - the corner where the table started - instead of the corner where it finished.

Read dp[0][0]

Why: dp[0][0] = 0 by the base case - it represents comparing two empty prefixes, not the full strings.

Conclude the strings share no common subsequence

Why: This wrongly reports length 0, when the correct table clearly builds up to a shared subsequence of length 2 by the time both full strings have been compared.

The fix

The answer to the full problem always sits in the cell that has compared the FULL length of both strings - the bottom-right corner, dp[m][n].

Read dp[2][3] (m=2, n=3)

Why: This cell has compared all of X against all of Y, so it holds the true answer: length 2.

Trace back starting from that same bottom-right cell

Why: Reconstruction must also start where the answer lives and walk backward toward dp[0][0], never the other way around - walking forward from the start does not follow the choices the recurrence actually made.

71. LCS time and space complexity

Concept

The table has one entry per pair of prefixes, and each entry takes constant work to fill once its neighbors are known.

\[ O(mn) \text{ time}, \quad O(mn) \text{ space} \]

Both string lengths matter here, unlike LIS - doubling the length of either string roughly doubles the total work.

72. The length is easy; the sequence needs the table

Picture it

Animation

Shows: The length is easy; the sequence needs the table — a rendered Manim animation.

Rendered with Manim.

Takeaway: Space optimisation costs you the traceback. Decide which you need first.

73. State the rule before it runs: Worked example: LCS - a second trace with…

Hypothesis

Predict first

Worked example: LCS - a second trace with more mismatches is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Base row and column are zero

Why: Same rule as before: comparing against an empty prefix always gives length 0.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

74. Worked example: LCS - a second trace with more mismatches

Worked example

Find the longest common subsequence of X = "ABCB" (m=4) and Y = "BDCB" (n=4).

Base row and column are zero

Why: Same rule as before: comparing against an empty prefix always gives length 0.

\[ \text{dp}[0][j] = 0, \quad \text{dp}[i][0] = 0 \]

Fill row 1 (X prefix = "A")

Why: A matches none of B, D, C, B, so every entry in row 1 falls back to its neighbor - which is always 0 here.

\[ \text{dp}[1][1..4] = 0,\ 0,\ 0,\ 0 \]

Fill row 2 (X prefix = "AB")

Why: The new character is B. It matches Y's 1st character (B) and Y's 4th character (B). dp[2][1] = dp[1][0]+1 = 1. The mismatches with D and C carry the best neighbor forward, giving dp[2][2]=dp[2][3]=1, and the second match gives dp[2][4] = dp[1][3]+1 = 1.

\[ \text{dp}[2][1..4] = 1,\ 1,\ 1,\ 1 \]

Fill row 3 (X prefix = "ABC")

Why: The new character is C, matching Y's 3rd character: dp[3][3] = dp[2][2]+1 = 2. The mismatches before and after take the best neighbor: dp[3][1]=1, dp[3][2]=1, and dp[3][4]=max(dp[2][4], dp[3][3])=max(1,2)=2.

\[ \text{dp}[3][1..4] = 1,\ 1,\ 2,\ 2 \]

Check the completed table with row 4 filled in

Why: Row 4 (X prefix "ABCB") is filled the same way: the two matches with B give dp[4][1]=dp[3][0]+1=1 and dp[4][4]=dp[3][3]+1=3, while the two mismatches falling back to their best neighbor give dp[4][2]=1 and dp[4][3]=2 - all consistent with the earlier rows.

"" (j=0)B (j=1)D (j=2)C (j=3)B (j=4)
"" (i=0)00000
A (i=1)00000
AB (i=2)01111
ABC (i=3)01122
ABCB (i=4)01123

75. LCS - a second trace with more mismatches — line by line

Picture it

Animation

Shows: Each line of the worked example "LCS - a second trace with more mismatches", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Row 4 (X prefix "ABCB") is filled the same way: the two matches with B give dp[4][1]=dp[3][0]+1=1 and dp[4][4]=dp[3][3]+1=3, while the two mismatches falling back to their best neighbor give dp[4][2]=1 and dp[4][3]=2 - all consistent with the earlier rows.

76. What has to happen first: Worked example: LCS - trace back the second example

Ranking

Put in order

Put the moves of Worked example: LCS - trace back the second example into the order they have to happen.

  1. Trace back from dp[4][4]
  2. Continue from dp[3][3]
  3. Continue from dp[2][2]
  4. Finish at dp[2][1]
  5. Verify the recovered subsequence

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. X's 4th character is B and Y's 4th character is B - they match, so B belongs to the subsequence.

77. Worked example: LCS - trace back the second example

Worked example

The bottom-right entry, dp[4][4] = 3, says X="ABCB" and Y="BDCB" share a common subsequence of length 3.

Trace back from dp[4][4]

Why: X's 4th character is B and Y's 4th character is B - they match, so B belongs to the subsequence. Step diagonally to dp[3][3].

Continue from dp[3][3]

Why: X's 3rd character is C and Y's 3rd character is C - they match, so C belongs to the subsequence too. Step diagonally to dp[2][2].

Continue from dp[2][2]

Why: X's 2nd character is B and Y's 2nd character is D - they mismatch. dp[2][2]=1 came from the better neighbor, dp[2][1]=1 (not dp[1][2]=0), so step left to dp[2][1] - no character is added here.

Finish at dp[2][1]

Why: X's 2nd character is B and Y's 1st character is B - they match, so a second B belongs to the subsequence. Step diagonally to dp[1][0], a base case - the trace stops.

Verify the recovered subsequence

Why: Reading the matches in the order found (B, C, B) and reversing gives "BCB": in X="ABCB" the letters B, C, B appear at positions 2, 3, 4 in order; in Y="BDCB" they appear at positions 1, 3, 4 in order.

\[ \text{LCS}(\text{"ABCB"},\ \text{"BDCB"}) = \text{"BCB"}\ \checkmark \]

78. Work backwards from the answer: Worked example: LCS - trace back the second…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Verify the recovered subsequence

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The bottom-right entry, dp[4][4] = 3, says X="ABCB" and Y="BDCB" share a common subsequence of length 3.

79. Check yourself: LCS base case

Check

Suppose X has 4 characters and Y has 6 characters.

Check your understanding

What value does dp[0][5] hold in the LCS table, and why?

  • A. 0, because comparing an empty prefix of X against any prefix of Y always gives a common subsequence of length 0. (correct)
  • B. 5, because it counts the 5 characters used so far from Y.
  • C. 4, because it counts the length of X.
  • D. It cannot be determined without knowing the actual characters in Y.

Answer: A

Why: Row 0 represents using zero characters of X. No matter how many characters of Y you compare against an empty string, there is nothing to match, so every entry in row 0, including dp[0][5], is 0 by the base case.

Why B tempts people
This confuses the column index (how far into Y we've compared) with the dp VALUE. The index 5 counts characters consumed from Y, but the base case still forces the LCS length itself to be 0 whenever i=0.
Why C tempts people
This substitutes a different length (X's) that has nothing to do with row 0. Row 0 is defined by having compared zero characters of X, not m of them.
Why D tempts people
The base case for row 0 never depends on which characters Y actually contains - comparing against an empty prefix is always 0 regardless of content.

80. How sure are you: Check yourself: LCS mismatch case

Commit first

Predict first

Which value should dp[i][j] take in this mismatch case?

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: The maximum of dp[i-1][j] and dp[i][j-1].

Why: A mismatch adds no new shared character, so dp[i][j] can't grow past what's already known. The best available answer is whichever neighbor - dropping the current character of X, or dropping the current character of Y - already has the larger LCS length, hence the maximum of dp[i-1][j] and dp[i][j-1].

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

81. Check yourself: LCS mismatch case

Check

Suppose X's i-th character and Y's j-th character are different, so this is a mismatch cell.

\[ x_i \neq y_j \]

Check your understanding

Which value should dp[i][j] take in this mismatch case?

  • A. The maximum of dp[i-1][j] and dp[i][j-1]. (correct)
  • B. dp[i-1][j-1] + 1, the same as the match case.
  • C. The minimum of dp[i-1][j] and dp[i][j-1].
  • D. 0, since no character can be added when the characters disagree.

Answer: A

Why: A mismatch adds no new shared character, so dp[i][j] can't grow past what's already known. The best available answer is whichever neighbor - dropping the current character of X, or dropping the current character of Y - already has the larger LCS length, hence the maximum of dp[i-1][j] and dp[i][j-1].

Why B tempts people
This applies the match case's plus-one rule to a mismatch, incorrectly crediting a shared character that doesn't actually exist at this pair of positions.
Why C tempts people
Taking the minimum instead of the maximum would throw away useful earlier work, deliberately picking the worse of two known partial answers instead of the better one.
Why D tempts people
Resetting to 0 discards every match found earlier in the strings. A mismatch simply fails to add a new character; it does not erase previously found ones.

82. The Edit Distance problem

Concept

You're given two strings. You want the minimum number of single-character edits - insertions, deletions, or replacements - needed to turn the first string into the second.

\[ X = x_1 \cdots x_m \ \longrightarrow \ Y = y_1 \cdots y_n \]

This is also called Levenshtein distance. Unlike LCS, which only asks what characters two strings share, edit distance asks how much work it takes to transform one into the other.

83. Edit distance, filling in

Picture it

Animation

Shows: Edit distance, filling in — a rendered Manim animation.

Rendered with Manim.

Takeaway: Three moves compete in every cell: insert, delete, substitute.

84. Picture aligning two strings, letter by letter

Intuition

Line the two strings up character by character. At each position you have three moves available: delete a character from X, insert a character to match Y, or replace a character of X with the one Y needs. Matching characters need no move at all.

The DP's job is to find the cheapest combination of these moves that turns all of X into all of Y.

85. Why you can't just count mismatched positions

Intuition

If the two strings were the same length, you might be tempted to just line them up position by position and count how many positions differ. But real strings can differ in length, and a single insertion or deletion shifts every character after it.

That's why edit distance needs a full grid of subproblems, not a single left-to-right scan: it has to consider every way the two strings could be lined up, including ones with extra characters inserted or removed.

86. Edit Distance state: dp[i][j] in words

Concept

Let dp[i][j] be the minimum number of edits needed to turn the first i characters of X into the first j characters of Y.

\[ \text{dp}[i][j] = \text{min edits to turn } x_1 \cdots x_i \text{ into } y_1 \cdots y_j \]

This looks just like the LCS state - two indices, one per string - but the number it stores means something different: a cost to minimize, not a length to maximize.

87. Where edit distance actually gets used

Picture it

Animation

Shows: Where edit distance actually gets used — a rendered Manim animation.

Rendered with Manim.

Takeaway: Changing the cost of each operation changes the application, not the algorithm.

88. Same shape as LCS, opposite goal

Intuition

Edit Distance's table looks identical to LCS's - same two indices, same grid, same fill order. But where LCS's recurrence rewards agreement (matches add to a growing length), Edit Distance's recurrence penalizes disagreement (mismatches add to a growing cost).

Keeping these two goals straight - maximizing shared characters versus minimizing edit cost - is what keeps the two recurrences, and their very different base cases, from blurring together.

89. Tracing back through the table

Picture it

Animation

Shows: Tracing back through the table — a rendered Manim animation.

Rendered with Manim.

Takeaway: Diagonal moves are the matches — following them backwards spells the answer.

90. Edit Distance base case: pure insertions or deletions

Concept

Turning i characters of X into an empty string takes exactly i deletions - delete everything. Turning an empty string into j characters of Y takes exactly j insertions - insert everything.

\[ \text{dp}[i][0] = i, \qquad \text{dp}[0][j] = j \]

This is the opposite of LCS's base case. LCS's base row and column are zero because there's nothing to MATCH against an empty string. Edit distance's base row and column count real work, because turning into or out of an empty string still costs one edit per character.

91. Something is wrong here: the base row and column are NOT zero

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student copies the LCS base case out of habit and initializes dp[i][0] = 0 and dp[0][j] = 0 for the Edit Distance table.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Copying the LCS pattern treats an empty target as free - as if turning "CAT" into nothing costs nothing.

Base cells must count the real work of pure insertions or pure deletions - they are not free.

Why: Copying the LCS pattern treats an empty target as free - as if turning "CAT" into nothing costs nothing.

92. Trap: the base row and column are NOT zero

Trap

The trap

A student copies the LCS base case out of habit and initializes dp[i][0] = 0 and dp[0][j] = 0 for the Edit Distance table.

Set dp[3][0] = 0 for X="CAT", Y=""

Why: Copying the LCS pattern treats an empty target as free - as if turning "CAT" into nothing costs nothing.

Get a nonsensical answer

Why: Turning "CAT" into the empty string clearly takes 3 deletions, not 0. Every downstream cell that depends on this wrong base inherits the error.

The fix

Base cells must count the real work of pure insertions or pure deletions - they are not free.

Set dp[3][0] = 3 for X="CAT", Y=""

Why: Turning "CAT" into the empty string requires deleting all 3 characters, one edit each.

Set dp[0][j] = j the same way

Why: Turning the empty string into a target of length j requires inserting all j characters. The base row and column are literally i and j, counted from the corner outward.

93. Which of these survive contact with Dynamic Programming III: LIS, LCS & Edit…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
You have named 13 reusable moves so far. Say as many as you can out loud, by number, from memory.; If you accidentally think in terms of contiguous runs, the recurrences will look wrong and the trace tables won't match what the DP is actually counting.; Longest Increasing Subsequence: finding the longest stretch of improving values buried in noisy data, like the best run of increasing days in a stock's price history.
Breaks
A student is asked whether "ACE" is a subsequence of "ABCDE", and instead checks whether "ACE" appears as a contiguous block of letters inside "ABCDE".; A student defines dp[i] as "the length of the longest increasing subsequence using any of the first i numbers," and expects dp[i] to always be at least as large as dp[i-1].
sound
These are stated as this lesson states them — each one survives the edge cases Dynamic Programming III: LIS, LCS & Edit Distance puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

94. Why is this step legal: It gives a wrong answer immediately

Explain it to yourself

Discussion prompt

In Process: guessing the edit-distance base row this move is made:

It gives a wrong answer immediately

Why is that legal? Name the rule or definition it rests on before you read on.

Hint: If you can only say "because that is what you do", the rule is the thing to go and find.

Answer:

Turning the empty string into a string of length 3 takes three insertions, not zero. A zero there claims the work is free, and every entry that depends on it inherits the error.

95. Process: guessing the edit-distance base row

Intuition

Watch me not know the answer. This is what the first two minutes actually look like.

Setting up the edit distance table, where the entry is the fewest edits to turn the first i characters of X into the first j characters of Y.

Try zeros along the top row and left column, the way LCS does it

Why: LCS starts at zero when either string is empty, and this table looks identical in shape. Copy it.

It gives a wrong answer immediately

Why: Turning the empty string into a string of length 3 takes three insertions, not zero. A zero there claims the work is free, and every entry that depends on it inherits the error.

Dead end. Not a mistake — a move that was worth trying and did not pay off. This happens in most proofs.

Back up. Read the sentence, not the shape

Why: The entry counts edits. With one string empty, the only way across is to insert every character of the other — so the base row is 0, 1, 2, 3 and the base column likewise.

Same table shape, opposite base cases, because LCS maximizes something you can have none of, while edit distance minimizes something you must pay for. The shape of the table never tells you the base case. The sentence does.

The expert does not see the whole path in advance. The expert tries something, reads the result, and adjusts. That is the skill.

96. The base row and column

Picture it

Animation

Shows: The base row and column — a rendered Manim animation.

Rendered with Manim.

Takeaway: Two similar tables with different borders — get them wrong and everything shifts.

97. Edit Distance recurrence: when characters match

Concept

If the i-th character of X already equals the j-th character of Y, that position needs no edit at all - just carry forward the cost of matching everything before it.

\[ x_i = y_j \ \Rightarrow \ \text{dp}[i][j] = \text{dp}[i-1][j-1] \]

No plus-one here - a real match is free, unlike LCS, where a match adds one to the running length.

98. An alignment, read off the table

Picture it

Animation

Shows: An alignment, read off the table — a rendered Manim animation.

Rendered with Manim.

Takeaway: The path through the table IS the edit script.

99. Edit Distance recurrence: when characters don't match

Concept

If the characters differ, you must pay for one edit, and you get to choose the cheapest of three options: delete the current character of X, insert the character Y needs, or replace one character for the other.

\[ x_i \neq y_j \ \Rightarrow \ \text{dp}[i][j] = 1 + \min\big(\text{dp}[i-1][j],\ \text{dp}[i][j-1],\ \text{dp}[i-1][j-1]\big) \]

Each of the three neighbors corresponds to one operation: the cell above is a deletion, the cell to the left is an insertion, and the diagonal cell is a replacement.

100. Plan first: Worked example: Edit Distance - set up X="SEA", Y="EAT"

Step zero

Discussion prompt

Worked example: Edit Distance - set up X="SEA", Y="EAT" — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Fill the base row and column

Answer:

  1. Fill the base row and column
  2. Fill row 1 (X prefix = "S")
  3. Fill row 2 (X prefix = "SE")
  4. Check row 3 (X prefix = "SEA") and the completed table

101. Worked example: Edit Distance - set up X="SEA", Y="EAT"

Worked example

Find the edit distance between X = "SEA" (m=3) and Y = "EAT" (n=3).

Fill the base row and column

Why: Turning i characters of X into nothing costs i deletions; turning nothing into j characters of Y costs j insertions.

\[ \text{dp}[0][0..3] = 0,\ 1,\ 2,\ 3 \qquad \text{dp}[0..3][0] = 0,\ 1,\ 2,\ 3 \]

Fill row 1 (X prefix = "S")

Why: S mismatches every character of Y (E, A, T), so each cell costs 1 plus the minimum of its three neighbors: dp[1][1]=1+min(0,1,1)=1; dp[1][2]=1+min(1,2,1)=2; dp[1][3]=1+min(2,3,2)=3.

\[ \text{dp}[1][1..3] = 1,\ 2,\ 3 \]

Fill row 2 (X prefix = "SE")

Why: E matches Y's 1st character: dp[2][1]=dp[1][0]=1 (no added cost). The rest mismatch: dp[2][2]=1+min(dp[1][1],dp[1][2],dp[2][1])=1+min(1,2,1)=2; dp[2][3]=1+min(dp[1][2],dp[1][3],dp[2][2])=1+min(2,3,2)=3.

\[ \text{dp}[2][1..3] = 1,\ 2,\ 3 \]

Check row 3 (X prefix = "SEA") and the completed table

Why: A matches Y's 2nd character: dp[3][2]=dp[2][1]=1 (no added cost). The mismatches cost one plus the best neighbor: dp[3][1]=1+min(dp[2][0],dp[2][1],dp[3][0])=1+min(2,1,3)=2; dp[3][3]=1+min(dp[2][2],dp[2][3],dp[3][2])=1+min(2,3,1)=2.

"" (j=0)E (j=1)A (j=2)T (j=3)
"" (i=0)0123
S (i=1)1123
E (i=2)2123
A (i=3)3212

102. Weighting the operations

Picture it

Animation

Shows: Weighting the operations — a rendered Manim animation.

Rendered with Manim.

Takeaway: The algorithm does not change; only the numbers in the min do.

103. Guess the shape of the answer: Worked example: Edit Distance - read the…

Estimation

Predict first

The bottom-right entry, dp[3][3], is the minimum number of edits to turn the full string X into the full string Y.

Commit before you compute: what does Worked example: Edit Distance - read the answer and trace… come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Verify the total cost matches

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Reading the moves in order (delete S, keep E, keep A, insert T) turns "SEA" into "EA" and then into "EAT" using exactly 1 deletion plus 1 insertion, or 2 edits total, matching dp[3][3].

104. Worked example: Edit Distance - read the answer and trace back the edits

Worked example

The bottom-right entry, dp[3][3], is the minimum number of edits to turn the full string X into the full string Y.

Read the final answer

Why: dp[3][3] = 2, so "SEA" can become "EAT" in as few as 2 edits.

\[ \text{dp}[3][3] = 2 \]

Trace back from dp[3][3]

Why: X's 3rd character is A and Y's 3rd character is T - they mismatch. Its value, 2, equals 1 + dp[3][2] (the left neighbor, value 1) - the smallest of the three candidates - so this step is an INSERT of the character T.

Continue from dp[3][2]

Why: X's 3rd character is A and Y's 2nd character is A - they match, so no edit here. Step diagonally to dp[2][1] for free.

Continue from dp[2][1]

Why: X's 2nd character is E and Y's 1st character is E - they match, so no edit here either. Step diagonally to dp[1][0] for free.

Finish at dp[1][0]

Why: This is a base-case cell (j=0) with value 1: turning the single character S into nothing costs one DELETION. The trace stops here.

Verify the total cost matches

Why: Reading the moves in order (delete S, keep E, keep A, insert T) turns "SEA" into "EA" and then into "EAT" using exactly 1 deletion plus 1 insertion, or 2 edits total, matching dp[3][3].

\[ \text{"SEA"} \xrightarrow{\text{delete S}} \text{"EA"} \xrightarrow{\text{insert T}} \text{"EAT"}\ \checkmark \]

105. Edit Distance - read the answer and trace back the… — line by line

Picture it

Animation

Shows: Each line of the worked example "Edit Distance - read the answer and trace back the edits", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: Reading the moves in order (delete S, keep E, keep A, insert T) turns "SEA" into "EA" and then into "EAT" using exactly 1 deletion plus 1 insertion, or 2 edits total, matching dp[3][3].

106. Reading operations off the direction you traced

Concept

Each direction you step during traceback corresponds to exactly one kind of edit.

\[ \text{diagonal} = \text{match (free) or replace (cost 1)}, \quad \text{up} = \text{delete}, \quad \text{left} = \text{insert} \]

A diagonal step is free when the characters already match, and costs one edit when they don't (a replacement). Knowing this turns the bare numbers in the table into an actual list of edits.

107. Edit Distance time and space complexity

Concept

Just like LCS, every one of the (m+1) by (n+1) cells takes constant work once its neighbors are known.

\[ O(mn) \text{ time}, \quad O(mn) \text{ space} \]

The only difference from LCS is what each cell computes - a minimum cost instead of a maximum length - the shape of the computation is identical.

108. Three edits, three neighbours

Picture it

Animation

Shows: Three edits, three neighbours — a rendered Manim animation.

Rendered with Manim.

Takeaway: The geometry of the table is the algorithm.

109. Plan first: Worked example: Edit Distance - a second example with two…

Step zero

Discussion prompt

Worked example: Edit Distance - a second example with two replacements — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Fill the base row and column

Answer:

  1. Fill the base row and column
  2. Fill row 1 (X prefix = "A")
  3. Check row 2 (X prefix = "AB") and the completed table

110. Worked example: Edit Distance - a second example with two replacements

Worked example

Find the edit distance between X = "AB" (m=2) and Y = "BA" (n=2).

Fill the base row and column

Why: dp[i][0]=i counts pure deletions; dp[0][j]=j counts pure insertions.

\[ \text{dp}[0][0..2] = 0,\ 1,\ 2 \qquad \text{dp}[0..2][0] = 0,\ 1,\ 2 \]

Fill row 1 (X prefix = "A")

Why: A mismatches Y's 1st character (B): dp[1][1]=1+min(dp[0][0],dp[0][1],dp[1][0])=1+min(0,1,1)=1. A matches Y's 2nd character (A): dp[1][2]=dp[0][1]=1 (no added cost).

\[ \text{dp}[1][1..2] = 1,\ 1 \]

Check row 2 (X prefix = "AB") and the completed table

Why: B matches Y's 1st character (B): dp[2][1]=dp[1][0]=1 (no added cost). B mismatches Y's 2nd character (A): dp[2][2]=1+min(dp[1][1],dp[1][2],dp[2][1])=1+min(1,1,1)=2.

"" (j=0)B (j=1)A (j=2)
"" (i=0)012
A (i=1)111
B (i=2)212

111. Edit Distance - a second example with two… — line by line

Picture it

Animation

Shows: Each line of the worked example "Edit Distance - a second example with two replacements", appearing one at a time.

The same working the example does, in the order a tutor would write it.

Takeaway: B matches Y's 1st character (B): dp[2][1]=dp[1][0]=1 (no added cost). B mismatches Y's 2nd character (A): dp[2][2]=1+min(dp[1][1],dp[1][2],dp[2][1])=1+min(1,1,1)=2.

112. What has to happen first: Worked example: Edit Distance - trace back the second…

Ranking

Put in order

Put the moves of Worked example: Edit Distance - trace back the second example into the order they have to happen.

  1. Trace back from dp[2][2]
  2. Continue from dp[1][1]
  3. Verify the total cost matches

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. X's 2nd character is B and Y's 2nd character is A - they mismatch.

113. Worked example: Edit Distance - trace back the second example

Worked example

The bottom-right entry, dp[2][2] = 2, says "AB" needs at least 2 edits to become "BA".

Trace back from dp[2][2]

Why: X's 2nd character is B and Y's 2nd character is A - they mismatch. Its value, 2, equals 1 + dp[1][1] (the diagonal neighbor, value 1), so this step is a REPLACE of B with A.

Continue from dp[1][1]

Why: X's 1st character is A and Y's 1st character is B - they mismatch. Its value, 1, equals 1 + dp[0][0] (the diagonal neighbor, value 0), so this step is also a REPLACE, of A with B.

Verify the total cost matches

Why: Replacing A with B and B with A turns "AB" directly into "BA" using exactly 2 replacements, matching dp[2][2] = 2. No cheaper combination of insert, delete, or replace exists here, since both characters must change.

\[ \text{"AB"} \xrightarrow{\text{replace both}} \text{"BA"}\ \checkmark \]

114. Decode the notation: Worked example: Edit Distance - trace back the second…

Notation

Annotate

From Worked example: Edit Distance - trace back the second… — read this one piece at a time. What is each part doing?

On: \( \text{"AB"} \xrightarrow{\text{replace both}} \text{"BA"}\ \checkmark \)

  • X's 2nd character is B and Y's 2nd character is A - they mismatch. Its value, 2, equals 1 + dp[1][1] (the diagonal neighbor, value 1), so this step is a REPLACE of B with A.
  • X's 1st character is A and Y's 1st character is B - they mismatch. Its value, 1, equals 1 + dp[0][0] (the diagonal neighbor, value 0), so this step is also a REPLACE, of A with B.
  • Replacing A with B and B with A turns "AB" directly into "BA" using exactly 2 replacements, matching dp[2][2] = 2. No cheaper combination of insert, delete, or replace exists here, since both characters must change.

115. Answer it before you see the options: Check yourself: Edit Distance base case

Prediction

Predict first

What is dp[5][0], and what does it represent?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 5, because turning 5 characters into nothing requires 5 deletions.

Why: dp[i][0] is defined as the cost to turn i characters of X into the empty string. Every character must be deleted separately, one edit each, so dp[5][0] = 5 - five real deletions, not a free base case.

116. Check yourself: Edit Distance base case

Check

Suppose X has 5 characters and Y is the empty string.

Check your understanding

What is dp[5][0], and what does it represent?

  • A. 5, because turning 5 characters into nothing requires 5 deletions. (correct)
  • B. 0, because comparing against an empty string always costs nothing.
  • C. 5, because it just copies the length of X with no real meaning.
  • D. 1, because a single deletion clears the whole string at once.

Answer: A

Why: dp[i][0] is defined as the cost to turn i characters of X into the empty string. Every character must be deleted separately, one edit each, so dp[5][0] = 5 - five real deletions, not a free base case.

Why B tempts people
This copies the LCS base case, where matching against empty is free (0). Edit distance's base case counts real work - pure insertions or deletions - so it is the opposite of LCS's zero base.
Why C tempts people
This treats the 5 as an arbitrary label rather than an actual edit count. dp[5][0] = 5 specifically because 5 separate delete operations are required, not just because 5 happens to equal m.
Why D tempts people
Deletions in this DP remove one character at a time; there is no single edit that clears an entire string, so 5 characters require 5 separate deletions.

117. Decision point: three sequence DPs, from memory

Intuition

What move should we make next?

LIS, LCS and edit distance are all done. Close the deck.

For each: the subproblem sentence, the last decision, and the base case. Nine answers.

This is the same drill as last lesson, and it will be the same drill next lesson. The recurrences are worth nothing to you if the procedure that produces them is not automatic.

_Look at your toolkit. Say a move number out loud before this slide advances._ A wrong guess is useful. A silent guess is not.

118. Teach it back: Decision point: three sequence DPs, from memory

Explain it

Discussion prompt

Explain Decision point: three sequence DPs, from memory to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

LIS, LCS and edit distance are all done. Close the deck.

119. Two sequences means two indices

Picture it

Animation

Shows: Two sequences means two indices — a rendered Manim animation.

Rendered with Manim.

Takeaway: One index per sequence — the table's shape follows from the state.

120. The three setups, side by side

Concept

All three DPs share the same shape: define a state in words, write a recurrence with a matching case and a non-matching case, nail down a base, and know how to trace back the real answer. Only the details change.

LISLCSEdit Distance
Statedp[i]: longest run ending at idp[i][j]: LCS length of two prefixesdp[i][j]: min edits between two prefixes
Basedp[i]=1 (alone)dp[0][j]=dp[i][0]=0 (empty match)dp[i][0]=i, dp[0][j]=j (pure edits)
Match/extend caseextend from a smaller earlier value+1 from the diagonal0 cost from the diagonal
Non-match caseskip that earlier positionmax of the two neighbors1 + min of three neighbors

121. Fill in: LIS for The three setups, side by side

Comparison

Comparison matrix

From The three setups, side by side: refill the LIS column from what you know. The rest of the table is as it appeared.

LISLCSEdit Distance
Statedp[i]: longest run ending at idp[i][j]: LCS length of two prefixesdp[i][j]: min edits between two prefixes
Basedp[i]=1 (alone)dp[0][j]=dp[i][0]=0 (empty match)dp[i][0]=i, dp[0][j]=j (pure edits)
Match/extend caseextend from a smaller earlier value+1 from the diagonal0 cost from the diagonal
Non-match caseskip that earlier positionmax of the two neighbors1 + min of three neighbors

122. How to recognize which DP a problem calls for

Concept

One sequence, looking for the best run by some ordering property (like increasing values): that's an LIS-shaped problem, with one index into one sequence.

Two sequences, asking what they share in order: that's an LCS-shaped problem, with two indices, one per sequence, and a recurrence that rewards matches.

Two sequences, asking the cost to transform one into the other: that's an Edit-Distance-shaped problem, with two indices and a recurrence that penalizes mismatches instead of rewarding matches.

123. By analogy: How to recognize which DP a problem calls for

Analogy

Discussion prompt

Explain How to recognize which DP a problem calls for by analogy to something with no CS3000 Algorithms in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

One sequence, looking for the best run by some ordering property (like increasing values): that's an LIS-shaped problem, with one index into one sequence.

124. Why all three run in polynomial time despite an exponential search space

Intuition

Brute force for any of these three problems means examining an exponential number of candidate subsequences or alignments. But every candidate's evaluation reuses the same small set of prefix comparisons over and over.

By storing each prefix comparison's answer exactly once - in a table with only a polynomial number of cells - we replace exponentially many repeated computations with one polynomial-sized table, filled once, left to right, top to bottom.

125. Two rows are often enough

Picture it

Animation

Shows: Two rows are often enough — a rendered Manim animation.

Rendered with Manim.

Takeaway: You lose traceback in exchange, so keep the full table when you need the path.

126. Rule out three: Check yourself: matching the DP to the problem

Elimination

Eliminate the wrong options

Which of today's three DPs directly fits this problem, and why?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Longest Common Subsequence, because it compares two sequences and looks for the longest subsequence shared by both.
  • B. Longest Increasing Subsequence, because DNA bases can be put in alphabetical order and treated as numbers.
  • C. Edit Distance, because it also uses a two-index table over two strings.
  • D. None of them, because DNA sequences are too long for any table-based approach.

Survives elimination: A

Why: The problem asks for the longest subsequence common to two given sequences - exactly the Longest Common Subsequence setup: two indices, one per strand, and a recurrence that rewards shared bases. Its length is the direct answer.

127. Check yourself: matching the DP to the problem

Check

Consider this problem: given two DNA strands (sequences of the letters A, C, G, T), find the length of the longest sequence of bases that appears in both strands, in order, allowing gaps.

Check your understanding

Which of today's three DPs directly fits this problem, and why?

  • A. Longest Common Subsequence, because it compares two sequences and looks for the longest subsequence shared by both. (correct)
  • B. Longest Increasing Subsequence, because DNA bases can be put in alphabetical order and treated as numbers.
  • C. Edit Distance, because it also uses a two-index table over two strings.
  • D. None of them, because DNA sequences are too long for any table-based approach.

Answer: A

Why: The problem asks for the longest subsequence common to two given sequences - exactly the Longest Common Subsequence setup: two indices, one per strand, and a recurrence that rewards shared bases. Its length is the direct answer.

Why B tempts people
LIS operates on a single sequence and looks for internal ordering (like increasing numbers), not agreement between two separate sequences. Alphabetizing bases doesn't turn a two-sequence comparison into a one-sequence problem.
Why C tempts people
Edit Distance also uses a two-index table, but it measures the cost to transform one string into another, not the length of a subsequence they share - a different question with a different, minimizing rather than maximizing, recurrence.
Why D tempts people
A polynomial-time table, sized proportional to the product of the two lengths, handles even long DNA strands efficiently - this is in fact one of LCS's most common real-world uses.

128. Toolkit update

Concept

Moves added today: none.

That is a result, not a gap. Everything in this lesson was proved with moves you already owned.

Moves you reused today:

Fourth lesson in a row with no new move. If the skeleton still feels like something you are reading rather than something you own, that is the signal to drill it, not to learn something new.

Full toolkit so far: #1 through #13.

Next session opens with you naming every one of these from memory, before any new material.

129. Break it if you can: Toolkit update

Counterexample

Discussion prompt

Fourth lesson in a row with no new move. If the skeleton still feels like something you are reading rather than something you own, that is the signal to drill it, not to learn something new.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Next session opens with you naming every one of these from memory, before any new material.

130. Connect it up: Dynamic Programming III: LIS, LCS & Edit Distance

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The DP setup checklist · Toolkit check-in: name them before you look · Recap: what makes a problem dynamic programming · What a subsequence is · Why this distinction matters for today's three DPs. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

131. What you can do now

Recap

You now have the setup recipe for three sequence DPs that show up constantly in this course and beyond.

StepWhat to do
1. StateSay what each table entry means, in words, before writing a formula
2. RecurrenceWrite the matching case and the non-matching case separately
3. BaseWork out the smallest cases by hand - never assume they're zero
4. TracebackWalk from the answer cell backward, one real move at a time

Sources

  1. Cormen, Leiserson, Rivest & Stein, Introduction to Algorithms, 3rd/4th ed., Ch. 15 (Dynamic Programming) - Longest Common Subsequence and Edit Distance — MIT Press.
  2. Kleinberg & Tardos, Algorithm Design, Ch. 6 (Dynamic Programming) - Sequence Alignment and related subsequence problems — Addison-Wesley, 2005.
  3. Every recurrence, dp table, and traceback in this deck was recomputed by hand for the exact inputs used (LIS array 5,2,8,6,3,6,9,7; LCS pairs "AC"/"GAC" and "ABCB"/"BDCB"; Edit Distance pairs "SEA"/"EAT" and "AB"/"BA"), cell by cell, before being written up. — Verified 2026-07-18.
  4. Northeastern University CS 3000, Algorithms and Data (Summer 2026) — course page and syllabus — course.ccs.neu.edu/cs3000su26. Sets Cormen, Leiserson, Rivest and Stein, Introduction to Algorithms (3rd ed.) as the textbook; listings follow its conventions.
  5. CS 3000 course notes and midterm references circulated by students — github.com/vigneshsaravanakumar404/CS-3000-Algorithms-Data. Notes are typeset with the algpseudocode package, which is the style the listings in this deck follow.

Want this taught 1-on-1? Alexander tutors CS3000 Algorithms — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108