This lesson generates random words weighted by frequency, builds a Markov model mapping prefixes to suffixes, works through the data structure choices that model forces, and closes with the five strategies for debugging a hard problem.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 13 — Case study: data structure selection
§13.7-13.10, pp. 130-133
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-133 — the pages these objectives are drawn from
Warm-up
You can now pick words weighted by how often they appear. Try it.
Discussion prompt
Sampling ten words from Emma by frequency gives: this the small regard harriet which knightley's it most things. The word choice is right and the result is not a sentence. What is missing?
Hint: What does the usually come before?
Answer:
Any relationship between successive words. In a real sentence you would expect an article like the to be followed by an adjective or a noun, and probably not a verb or adverb.
The frequency model knows how often each word appears and nothing whatever about what follows it. Each pick is independent of the last.
So the fix is to make each choice depend on what came before — which is what Markov analysis does, and it is the rest of this lesson.
Concept
A series of random words seldom makes sense because there is no relationship between successive words. One way to measure these kinds of relationships is Markov analysis, which characterises, for a given sequence of words, the probability of the words that might come next.
Markov analysis — A way of characterising, for a sequence of words, the probability of the words that might come next.
The result is a mapping from each prefix to all possible suffixes. Given the mapping, you can generate random text by starting with any prefix, choosing at random from its possible suffixes, and repeating.
Figure (svg): Two columns contrasting independent word choice with choice conditioned on the preceding words
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-131
Section
Section 1
Concept
To choose a random word from the histogram, the simplest algorithm is to build a list with multiple copies of each word, according to the observed frequency, and then choose from the list.
def random_word(h):
t = []
for word, freq in h.items():
t.extend([word] * freq)
return random.choice(t)| Expression | What it does | Note |
|---|---|---|
| [word] * freq | freq copies of the string | list repetition |
| t.extend(...) | adds every element | unlike append |
| random.choice(t) | uniform over positions | so weighted by frequency |
The expression [word] * freq creates a list with freq copies of the string word, and the extend method is similar to append except that the argument is a sequence. This algorithm works, but it is not very efficient: each time you choose a random word, it rebuilds the list, which is as big as the original book.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-130
Picture it
Every call reconstructs a list the size of the whole text.
Figure (svg): A growth chart comparing rebuilding the list per call against building it once or using cumulative sums
The first is fine for one word and unusable for a thousand. That difference is the chapter's subject.
Worked example
The two are easy to confuse and do different things.
>>> t = ['a']
>>> t.append(['b', 'b'])
>>> t
['a', ['b', 'b']]
>>> t = ['a']
>>> t.extend(['b', 'b'])
>>> t
['a', 'b', 'b']| Method | What it adds | Result |
|---|---|---|
| append | adds ONE element | which happens to be a list |
| extend | adds each element of the argument | flattens one level |
| for random_word | extend is wanted | the copies must be separate |
Recall what append does.
Why: It adds a single element, so appending a list nests it — lesson 10b's four wrong appends.
Compare with extend.
Why: The extend method is similar to append except that the argument is a sequence, and each of its elements is added separately.
See which random_word needs.
Why: The copies must be separate elements for choice to pick between them, so extend is the right one.
Figure (svg): The state of the program after each line of Worked example extend against append, drawn as a ladder with one rung per traced line
append nests and extend flattens. Using append here would produce a list of lists, and choice would return a list rather than a word.
Verify: Check what the bug would look like.
Why: The function would return something like ['the', 'the', 'the'] rather than 'the' — a legal value of the wrong shape, which is a shape error in the sense of lesson 12c. It would not raise until the result was used as a string, several steps later.
Prediction
A list times an integer.
word = 'bee'
freq = 3
print(['bee'] * freq)| Part | What it does | Result |
|---|---|---|
| ['bee'] | a one-element list | a singleton |
| * 3 | list repetition | three copies |
| the result | three separate elements | not a nested list |
Predict first
What does this print?
Correct: ['bee', 'bee', 'bee'] — the expression creates a list with freq copies of the string.
Why: List repetition copies the elements, not the list itself, so three separate strings end up at three positions. Note that 'bee' * 3 without the brackets would give option B, a repeated string — one pair of brackets is the difference between three list elements and one long word.
Worked example
The book's four-step improvement, which keeps the histogram.
# 1. the words
words = list(h.keys())
# 2. the running totals; the last is n, the book's length
cum = []
total = 0
for w in words:
total = total + h[w]
cum.append(total)
# 3. pick 1..n and bisect for the insertion point
# 4. that index gives the word| Step | What it produces | Note |
|---|---|---|
| the cumulative list | running totals of frequencies | increasing |
| the last item | the total number of words | n |
| a random 1..n | a uniform position | in the whole text |
| bisection | finds where it falls | in about log n steps |
Build the running totals once.
Why: Each word occupies a stretch of the cumulative list as wide as its frequency — the range picture from lesson 13a, made concrete.
Draw a position uniformly.
Why: Choose a random number from 1 to n, where n is the last item and the total number of words in the book.
Find it by bisection.
Why: Use a bisection search to find the index where the random number would be inserted, then use the index to find the corresponding word.
Figure (svg): A ladder showing cumulative frequency totals and the ranges each word occupies
The same probabilities from a structure the size of the vocabulary rather than of the text, with each choice costing about log n steps rather than a full rebuild.
Verify: Check why bisection is possible at all.
Why: Because the cumulative sums are increasing by construction — every frequency is at least 1, so each total exceeds the last. A bisection search needs a sorted sequence, and this one is sorted for free, which is what makes the fast lookup available.
Trap
A generator calls random_word a thousand times to produce a passage.
Call the function you have
Why: It returns a correctly weighted random word, which is what is needed.
Each call rebuilds a list of 161,080 strings, so a thousand words means building a hundred and sixty million list entries. The program is correct and takes minutes.
Build the expensive structure once.
An obvious improvement is to build the list once and then make multiple selections
Why: Which is the book's own first suggestion, and it is a large win for very little work.
Or use the cumulative sums, which are smaller as well as reusable
Why: Vocabulary-sized rather than text-sized.
This is the memo pattern from chapter 11 in another guise: work that does not change between calls should not be repeated in every call. Recognising that shape is worth more than either specific fix.
Discrimination
Ask whether the argument is one element or several.
Sort into buckets
For each goal, which method is right?
Faded example
Each copy must be its own element.
Fill in the blanks
for word, freq in h.items():
t.extend([word] * freq)
Why: extend adds each element of its argument separately, so the copies become separate list entries that choice can pick between. append would add the whole list as one nested element, and choice would then return a list of identical words rather than a word — a shape error that raises only when the result is used.
Socratic
The function is correct either way.
Discussion prompt
random_word gives the right answer every time. Why is it worth replacing, and when would it not be?
Hint: How many times will it be called?
Answer:
For one word it is perfect: simple, obviously correct, and the cost is a single pass over the histogram.
For a thousand words the cost is a thousand passes, each building a list as big as the book. The work is proportional to calls times text length, and both are large.
So the answer depends entirely on the use. The book's own advice in §13.9 is to choose the structure that is easiest to implement and see whether it is fast enough — and only if not, to improve it. Replacing this function before knowing how often it will be called would be optimising without evidence.
Section
Section 2
Concept
The result of Markov analysis is a mapping from each prefix to all possible suffixes. Given this mapping, you can generate a random text by starting with any prefix and choosing at random from the possible suffixes.
# from 'Eric, the Half a Bee':
# 'half the' -> always followed by 'bee'
# 'the bee' -> might be followed by 'has' or 'is'
# 'a bee' -> 'philosophically', 'be' or 'due'| Part | What it is | Note |
|---|---|---|
| a prefix | a short sequence of words | two, in this example |
| its suffixes | every word that followed it | with repeats allowed |
| generation | pick a suffix, shift the prefix | and repeat |
In the example the length of the prefix is always two, but you can do Markov analysis with any prefix length — and the exercise asks you to write the program in a way that makes it easy to try others.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 131-131
Picture it
Each prefix, and everything observed after it.
Figure (svg): A dictionary mapping two-word prefixes to the lists of words observed after them
Notice the first prefix has only one possible suffix, so generation from it is forced. Variety comes from prefixes with several.
Worked example
One pass over the words, keeping a sliding window.
def build(words, n=2):
d = {}
prefix = tuple(words[:n])
for word in words[n:]:
d.setdefault(prefix, []).append(word)
prefix = shift(prefix, word)
return d| Part | What it does | Note |
|---|---|---|
| prefix | the first n words | as a tuple |
| for each later word | it is a suffix of the current prefix | record it |
| shift | drop the first, add this word | the next prefix |
Start with the first n words.
Why: The prefix is a tuple, for reasons idea 3 sets out — it has to be usable as a dictionary key.
Record each following word as a suffix.
Why: setdefault creates an empty list on a first sighting and returns the existing one otherwise, which is invert_dict's singleton pattern in one call.
Slide the window forward.
Why: The end of the prefix and the new suffix combine to form the next prefix, and repeat.
Figure (svg): A pipeline showing a sliding window over words producing prefix-suffix pairs
A dictionary from prefixes to lists of observed suffixes, built in one pass with a window sliding along the text.
Verify: Check the count of entries.
Why: Every word after the first n contributes exactly one suffix, so the total number of suffixes across all the lists equals len(words) - n. Summing the list lengths and comparing is a consistency check of the kind lesson 11c described, and it catches a window that skips or repeats a position.
Prediction
One prefix appears only once in the text.
# in 'Eric, the Half a Bee':
# the prefix ('half', 'a') appears once,
# followed by 'bee'| Part | What is true | Result |
|---|---|---|
| ('half', 'a') | one occurrence | one suffix recorded |
| its suffix list | ['bee'] | a single element |
| random.choice | no alternative | 'bee' |
Predict first
Starting from the prefix ('half', 'a'), what does generation produce next?
Correct: 'bee', necessarily — the prefix only appears once in the text, so there is only one possible suffix.
Why: The book makes this point directly: if you start with the prefix Half a, then the next word has to be bee. Variety in the generated text comes from prefixes with several recorded suffixes, like ('a', 'bee'), which might be followed by philosophically, be or due.
Worked example
Pick a suffix, shift, repeat.
def generate(d, prefix, count):
for i in range(count):
suffixes = d[prefix]
word = random.choice(suffixes)
print(word, end=' ')
prefix = shift(prefix, word)| Step | What it does | Note |
|---|---|---|
| d[prefix] | the possible next words | a list |
| random.choice | one of them | weighted by repeats |
| shift | the new prefix | end of old plus new word |
Look up the prefix.
Why: The mapping gives every word observed after this prefix, with repeats — so a word that followed it often is in the list often.
Choose one at random.
Why: choice is uniform over positions, and the repeats do the weighting — exactly the mechanism from lesson 13a.
Form the next prefix.
Why: Combine the end of the prefix and the new suffix, and repeat.
Figure (svg): The state of the program after each line of Worked example generating text, drawn as a ladder with one rung per traced line
Text that follows the source's local patterns. The book's Emma sample is almost syntactically correct, but not quite; semantically, it almost makes sense, but not quite.
Verify: Ask what happens at a prefix with no entry.
Why: d[prefix] raises KeyError — which happens if generation reaches the very last prefix in the text, since no word followed it. A real implementation has to handle that, either by restarting from a random prefix or by using get with a fallback, and noticing it before it happens is the useful skill.
Trap
A program stores each prefix's suffixes without duplicates, to save space.
Avoid storing the same word repeatedly
Why: It looks wasteful to keep ten copies of the.
The duplicates were carrying the probabilities. Without them every possible continuation becomes equally likely, and a word that followed once competes evenly with one that followed a hundred times — so the generated text stops resembling the source.
Keep the repeats, or keep counts instead.
A list with duplicates
Why: The simplest option, and choice weights it automatically.
Or a histogram of suffixes
Why: One entry per distinct word with its count — same information, less space, harder to sample from.
That trade is exactly the one §13.9 discusses. Either is defensible; discarding the frequencies altogether is not, because they are the model.
Invariant
Follow the prefix as it shifts.
Step through it
What stays the same about the prefix across all four frames?
Its length. The window is always two words wide — it slides rather than growing, which is what makes the number of distinct prefixes manageable and the model a fixed size.
Faded example
A first sighting needs a list to append to.
Fill in the blanks
d.setdefault(prefix, []).append(word)
Why: setdefault returns the existing list for a known prefix and installs a new empty one for a first sighting, so append always has a list to work with — the singleton pattern from invert_dict, in a single call. Using get instead would return the default without storing it, so nothing would ever accumulate.
Real world
Predicting the next thing from the last few is a very general idea.
Discussion prompt
Where have you seen something predict what comes next from what came just before? What are its prefixes and suffixes?
Hint: Check your phone.
Answer:
Predictive text and autocomplete are the obvious ones: the prefix is what you have typed and the suffixes are the words that usually follow it.
So are music recommendation by what you just played, route prediction from the last few turns, and the sequence models underlying much larger systems.
What they share is the assumption that the recent past is enough — that you need not remember the whole history to predict the next step. That assumption is what makes the model small and what makes its output almost make sense but not quite, since the sentence's beginning is forgotten by the time its end is generated.
Section
Section 3
Concept
In your solution you had to choose how to represent the prefixes, how to represent the collection of possible suffixes, and how to represent the mapping. The last one is easy: a dictionary is the obvious choice for a mapping from keys to corresponding values.
The first step is to think about the operations you will need to implement for each. For the prefixes, we need to be able to remove words from the beginning and add to the end — and we also need to use them as keys, which is what settles it.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 132-132
Picture it
Two operations are needed, and only one structure supports both.
Figure (svg): Two columns comparing a list and a tuple against the two requirements for a prefix
So the easier structure loses to a requirement that has nothing to do with ease. That is what data structure selection usually looks like.
Worked example
Building the next prefix without modifying anything.
def shift(prefix, word):
return prefix[1:] + (word,)| Part | What it produces | Note |
|---|---|---|
| prefix[1:] | everything but the first word | a new tuple |
| (word,) | a one-element tuple | note the comma |
| the concatenation | a new tuple | nothing was modified |
Drop the first word.
Why: Slicing from index 1 gives a new tuple containing everything after the beginning.
Make the new word into a tuple.
Why: The trailing comma is required — without it, ('bee') is a string and the concatenation raises a TypeError.
Join them.
Why: shift takes a tuple of words and a string, and forms a new tuple that has all the words in prefix except the first, with word added to the end.
Figure (svg): A prefix tuple with its first element dropped and a new word appended
A new prefix, built rather than modified. With tuples you can't append or remove, but you can use the addition operator to form a new tuple.
Verify: Check that the length is preserved.
Why: The slice removes one element and the concatenation adds one, so the prefix stays exactly n words wide however many times shift is called. That invariant is what keeps the model a fixed size, and it would be worth an assertion in a longer program.
Prediction
A slice and a concatenation.
def shift(prefix, word):
return prefix[1:] + (word,)
print(shift(('half', 'a'), 'bee'))| Part | Value | Note |
|---|---|---|
| prefix[1:] | ('a',) | the first is dropped |
| (word,) | ('bee',) | a singleton tuple |
| the sum | ('a', 'bee') | still two words |
Predict first
What does this print?
Correct: ('a', 'bee') — the first word is dropped and the new one added, leaving the prefix two words wide.
Why: The slice removes one element and the concatenation adds one, so the length is preserved — which is what keeps the window sliding rather than growing. Option D is what would happen if the comma were omitted: ('bee') is a string, and adding a string to a tuple raises TypeError.
Worked example
The third option, and it nearly works.
# a string prefix
prefix = 'half a'
# shifting it means splitting and rejoining:
words = prefix.split()
prefix = ' '.join(words[1:] + [word])| Aspect | What is true | Note |
|---|---|---|
| a string prefix | hashable, so usable as a key | the requirement is met |
| shifting | split, slice, join | three operations |
| the risk | a word containing a space | would split wrongly |
Check the key requirement.
Why: Strings are immutable and hashable, so a string prefix could be a dictionary key — the requirement that ruled out lists does not rule this out.
Check the shifting operation.
Why: It needs splitting into words and rejoining, which is more work than slicing a tuple and less direct.
Notice the separator problem.
Why: The words are joined by a space, so the structure depends on no word containing one — the compound-key hazard from lesson 12c.
Figure (svg): The state of the program after each line of Worked example why not a string , drawn as a ladder with one rung per traced line
A string works and is worse: more work to shift, and it flattens the word boundaries into a separator that could in principle collide.
Verify: Compare what each structure preserves.
Why: The tuple keeps the words as separate elements, so prefix[0] is a word and shifting is a slice. The string has to reconstruct the boundaries every time, and it can only do so by convention. Keeping structure rather than encoding it into text is the general lesson, and it is the same reason lesson 12c preferred tuple keys to joined strings.
Trap
A student writes prefix.pop(0) and prefix.append(word) to slide the window.
Use the operations that fit the job
Why: Removing from the front and adding to the back is exactly what a list does well.
Those methods do not exist on a tuple, so it raises AttributeError — and switching to a list to get them makes the prefix unusable as a dictionary key, which is why the tuple was chosen.
Build a new prefix instead of changing one.
return prefix[1:] + (word,)
Why: A slice and a concatenation, both of which create rather than modify.
Assign the result
Why: prefix = shift(prefix, word), since nothing was changed in place.
The requirement chain is worth stating once: the prefix must be a key, keys must be hashable, hashable means immutable, and immutable means building rather than modifying. Each step follows from the one before.
Ranking
Four steps, in the order the argument runs.
Put in order
Why: The list is the first thought, because the operations needed are the ones lists do well. Then the key requirement appears, which forces hashability, which forces immutability — and the tuple is what is left. The conclusion is reached by elimination rather than by the operations, which is why the easiest structure is not the chosen one.
Faded example
The comma is doing essential work.
Fill in the blanks
def shift(prefix, word):
return prefix[1:] + (word,)
Why: Without the comma, (word) is just the string and adding a string to a tuple raises TypeError: can only concatenate tuple to tuple. The singleton comma from lesson 12a is load-bearing here, and its absence produces an error rather than silence — which for once is a mercy.
Explain it
The obvious structure is the wrong one, for a reason worth stating.
Discussion prompt
A classmate points out that a list handles adding and removing much more naturally than a tuple. Explain why the program uses tuples anyway.
Hint: What is the prefix used for besides shifting?
Answer:
Shifting is not the only thing a prefix does. It is also the key of the dictionary, and that is the requirement that decides everything.
Keys must be hashable, and a list is not — chapter 11's argument, that a mutable key could move after it was filed. So the list is ruled out however convenient its methods are.
What you lose is small: with tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is what shift does in one line. Ease of one operation loses to a hard requirement on another, which is what data structure selection usually comes down to.
Section
Section 4
Concept
So far we have been talking mostly about ease of implementation, but there are other factors to consider in choosing data structures.
benchmarking — The process of choosing between data structures by implementing alternatives and testing them on a sample of the possible inputs.
A practical alternative is to choose the data structure that is easiest to implement, and then see if it is fast enough for the intended application. If so, there is no need to go on.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 132-133
Picture it
Same information, different shapes, different costs.
Figure (svg): Two columns comparing a list of suffixes with a histogram of suffixes
Neither is right in the abstract. The book's advice is to take the easier one and measure before replacing it.
Worked example
Two structures, four considerations.
# list of suffixes
d[prefix].append(word)
random.choice(d[prefix])
# histogram of suffixes
h = d[prefix]
h[word] = h.get(word, 0) + 1
choose_from_hist(h) # needs the walk from 13a| Operation | Which is easier | Note |
|---|---|---|
| adding | one line either way | equally easy |
| choosing | one call, or a whole function | the list wins |
| space | one entry per occurrence, or per word | the histogram wins |
Compare adding.
Why: Adding a new suffix, or increasing the frequency of an existing one, is equally easy for the list implementation or the histogram.
Compare choosing.
Why: Choosing a random element from a list is easy; choosing from a histogram is harder to do efficiently, as exercise 13.7 showed.
Compare space.
Why: Using a histogram might take less space, because you only have to store each word once, no matter how many times it appears in the text.
Figure (svg): The state of the program after each line of Worked example the case for each, drawn as a ladder with one rung per traced line
The list is easier to sample and larger; the histogram is smaller and harder to sample. Which matters depends on the text's size and how much generation you do.
Verify: Ask when space would decide it.
Why: The book says that in some cases saving space can also make your program run faster, and in the extreme, your program might not run at all if you run out of memory — but that for many applications, space is a secondary consideration after run time. So the histogram wins only when the text is large enough for memory to bite.
Sorting
Four considerations, split between the two.
Sort into buckets
For each consideration, which representation is better?
Worked example
Two ways to find out, and one way to avoid needing to.
# benchmarking: implement both and compare
# on a sample of the possible inputs
# or: use the profile module to find
# where the program actually spends its time
# or, practically: write the easy one and
# see whether it is fast enough| Approach | What it involves | Note |
|---|---|---|
| benchmarking | implement both, measure | definitive and expensive |
| profiling | find the slow part first | before optimising anything |
| the practical route | easiest first, measure after | the book's advice |
Consider benchmarking.
Why: One option is to implement both and see which is better — definitive, and it costs writing the program twice.
Consider profiling.
Why: There are tools, like the profile module, that can identify the places in a program that take the most time — which tells you where optimising would even help.
Take the practical route first.
Why: Choose the structure that is easiest to implement, and then see if it is fast enough for the intended application. If so, there is no need to go on.
Figure (svg): A decision flowchart for choosing and then improving a data structure
Three approaches in increasing order of effort, and the book recommends starting with the cheapest. Optimising before measuring is work spent on a problem you have not confirmed you have.
Verify: Ask what profiling protects you from.
Why: Optimising the wrong thing. A program that spends ninety per cent of its time reading the file will not get noticeably faster however cleverly the suffixes are stored — and without measuring, there is no way to know that before doing the work.
Trap
A student implements the histogram version of the suffixes first, because it uses less memory.
Choose the more efficient structure from the start
Why: Why write something you know you will replace?
The harder version takes longer to write, is harder to get right, and may make no measurable difference — and until the program runs, there is no evidence that memory was ever the constraint.
Write the easy one and measure.
Choose the structure that is easiest to implement
Why: The list, here, and choice does the weighting for free.
Then see if it is fast enough for the intended application
Why: If it is, there is no need to go on.
The book's closing thought is worth keeping too: since analysis and generation are separate phases, you could use one structure for analysis and convert to another for generation — a net win if the time saved during generation exceeded the time spent converting.
Prediction
Two reasonable-sounding strategies.
# strategy A: pick the fastest structure up front
# strategy B: pick the easiest, then measure| Strategy | What it assumes | Note |
|---|---|---|
| up front | requires knowing the answer | often you do not |
| easiest first | and see if it is fast enough | the book's advice |
| if not | profile, then benchmark | in that order |
Predict first
What does the book recommend as the practical approach?
Correct: Choose the structure that is easiest to implement, then see if it is fast enough — and if so, there is no need to go on.
Why: Benchmarking is named as an option and it means writing the program twice, so it is the expensive route rather than the default. The book is explicit that often you don't know ahead of time which implementation will be faster, which is exactly why measuring the easy version first is the practical order.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C overstates a hedged claim. The book says that in SOME cases saving space can also make your program run faster, and that in the extreme a program might not run at all if it runs out of memory — but also that for many applications, space is a secondary consideration after run time. Sometimes and always are very different claims.
Real world
The advice generalises well beyond data structures.
Discussion prompt
Think of a time effort went into making something efficient before anyone knew whether it needed to be. What was the cost?
Hint: What else could that effort have gone into?
Answer:
Elaborate systems built for a scale that never arrived; optimisations that made code harder to change; time spent on a bottleneck that turned out not to be one.
The cost is rarely the wasted effort alone — it is that the complicated version is harder to modify when the actual requirement shows up somewhere else entirely.
Which is why the order matters: build the simple version, find out where it hurts, then fix that. The profile module exists precisely because the slow part is usually not where people guess it is.
Section
Section 5
Concept
When you are debugging a program, and especially if you are working on a hard bug, there are five things to try.
The list is ordered roughly by how much you have already tried. Retreating is last because it discards work, and it is on the list because sometimes discarding work is the fastest route.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 133-133
Picture it
Each works on a different kind of stuck.
Figure (svg): Two columns pairing each debugging strategy with the situation it suits
Recognising which kind you are in is most of the value. Doing more of what has already failed is the commonest debugging mistake.
Worked example
Thinking is a technique with specific questions attached.
# What kind of error is it?
# syntax, runtime, or semantic
# What can the error message tell me?
# What kind of error could cause this?
# What did I change last, before the problem appeared?| Question | What it narrows | Note |
|---|---|---|
| the kind | narrows where to look | syntax is at the point named |
| the message | names a type or a variable | already a strong clue |
| what changed last | the single best question | and easy to forget |
Classify the error.
Why: Syntax, runtime or semantic — the three kinds from chapter 1, and each has a different search strategy.
Mine the message.
Why: What information can you get from the error messages, or from the output of the program? A TypeError naming NoneType is a different investigation from a ValueError about unpacking.
Ask what changed.
Why: What did you change last, before the problem appeared? A program that worked an hour ago has a bug in the difference.
Figure (svg): The state of the program after each line of Worked example ruminating, with the book's questions, drawn as a ladder with one rung per traced line
Four specific questions rather than a vague instruction to think. The last one is the most powerful and the easiest to skip when you are frustrated.
Verify: Check the last question against version control.
Why: If the change since the last working version is small, the bug is in it — which turns an open-ended search into reading a handful of lines. That is why committing often is a debugging technique as much as a bookkeeping one.
Sorting
Each kind of stuck has its own move.
Sort into buckets
For each situation, which strategy is most likely to help?
Worked example
The listener is optional.
# 'So this function takes the prefix,
# looks up the suffixes, picks one at random,
# and then shifts the prefix - except it
# doesn't, because I never assigned the result...'| Stage | What happens | Note |
|---|---|---|
| explaining | forces a complete account | no steps assumed |
| the gap | appears mid-sentence | before the question is finished |
| the listener | not required | a rubber duck will do |
Explain the problem out loud.
Why: If you explain the problem to someone else, you sometimes find the answer before you finish asking the question.
Notice why it works.
Why: Explaining forces you to say every step, including the ones you have been assuming — and the assumption is usually where the bug is.
Notice the listener is optional.
Why: Often you don't need the other person; you could just talk to a rubber duck. That is the origin of the well-known strategy called rubber duck debugging.
Figure (svg): A flowchart showing explaining a program aloud surfacing an unstated assumption
A technique that works because of what it does to the explainer, not to the listener. The book is emphatic that it is not making this up.
Verify: Ask why reading silently is not the same.
Why: Reading lets you skim the parts you believe you understand, which are exactly the parts hiding the bug. Explaining aloud forces every step to be stated, and a step you cannot state is one you have not checked — which is what makes the technique work when re-reading has already failed.
Trap
A program is wrong, so changes are made and it is run again — repeatedly, for an hour.
Experiment until something works
Why: Running is on the list, and each attempt feels like progress.
Without a hypothesis, each run tests nothing, and the program drifts further from the version that was understood. After an hour there is neither a fix nor a program anyone can explain.
Change strategy when one stops producing information.
Read, or ruminate, or explain it to someone
Why: Each uses different evidence, so switching gives you something new.
And if the program is no longer understood, retreat
Why: Back off, undoing recent changes, until you get back to a program that works and that you understand. Then start rebuilding.
The five are alternatives rather than an escalation. The signal to switch is that the last several attempts taught you nothing, which is a much better trigger than frustration.
Prediction
Four questions from the Ruminating list.
# The program worked this morning.
# Now it raises a TypeError.| Fact | What it gives you | Note |
|---|---|---|
| it worked before | the bug is in the difference | a bounded search |
| the error kind | narrows the family | useful but broader |
| what changed last | the strongest clue here | and easy to skip |
Predict first
Which question narrows the search the most in this situation?
Correct: What did I change last, before the problem appeared? — a program that worked this morning has its bug in the difference.
Why: The other three are all useful and all search the whole program. This one bounds the search to whatever changed since the last working version, which may be a handful of lines. It is also the question people skip when frustrated, which is why the book lists it explicitly rather than leaving it to take some time to think.
Explain it to yourself
It discards work, which feels like the opposite of progress.
Discussion prompt
Undoing changes throws away effort. Why is it a debugging strategy rather than an admission of defeat?
Hint: What is the alternative costing you?
Answer:
Because the changes may be the problem. A program that has drifted through twenty edits contains twenty opportunities for a new bug, and the original one may already be fixed underneath them.
And because debugging requires understanding. Once you no longer know what the program does, every further change is a guess, and guesses are what got you here.
So the point of retreating is to get back to a program that works and that you understand, and then rebuild deliberately. The work is not wasted — what you learned about the bug survives; only the flailing edits are discarded.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C misreads what retreating is for. Backing off until you reach a program that works and that you understand is a step towards the fix, not away from it — because debugging requires understanding, and once that is gone every further change is a guess. The book's own phrasing is that then you can start rebuilding.
Comparison
Fill the blanks. One is obvious, one is forced, and one is a genuine trade.
Comparison matrix
| What to represent | The choice | Why |
|---|---|---|
| the mapping | a dictionary | it is a mapping from keys to values |
| the prefixes | a tuple of strings | they must be keys, so they must be hashable |
| the suffixes | a list, or a histogram | a genuine trade: easy sampling against less space |
| how to decide the last one | write the easiest, then measure | you often cannot tell in advance |
Only the third row is an open question, and the book's answer is to resolve it by measurement rather than by argument.
Pattern
Six steps. The third is where the data structure decision lands.
Step 4 must build a new tuple rather than modify one, because a tuple cannot be modified — and it must be a tuple because step 3 uses it as a key.
Python documentation — random — Generate pseudo-random numbers random — Generate pseudo-random numbers
Check
One requirement decides it.
Check your understanding
Why must a Markov prefix be a tuple rather than a list?
Answer: A
Why: Your first choice might be a list, since it is easy to add and remove elements — but the prefixes also have to be keys in a dictionary, and that rules out lists. With tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is what shift does.
Check
One element out, one in.
def shift(prefix, word):
return prefix[1:] + (word,)
print(shift(('the', 'bee'), 'is'))| Part | Value | Note |
|---|---|---|
| prefix[1:] | ('bee',) | the first is dropped |
| + ('is',) | concatenation | a new tuple |
| the result | ('bee', 'is') | still two words |
Check your understanding
What does this print?
Answer: A
Why: The slice removes the first word and the concatenation appends the new one, so the prefix stays two words wide. That length invariant is what makes the window slide rather than grow, and it keeps the number of distinct prefixes bounded.
Check
Twenty changes deep, and nothing works.
Check your understanding
You have made many changes trying to fix a bug and the program is now worse and no longer understood. Which strategy does the book recommend?
Answer: A
Why: At some point the best thing to do is back off, undoing recent changes, until you get back to a program that works and that you understand — then you can start rebuilding. Debugging requires understanding, and once that is gone, further changes are guesses.
Real world
Choosing how to store something before you know what you will ask of it.
Discussion prompt
Think of a time information was organised one way and the question that later mattered needed it another way. What did the conversion cost?
Hint: Filed by date, and you needed it by person.
Answer:
Records filed by date when you need them by customer; photographs in folders by year when you want them by who is in them; notes organised by source when you want them by topic.
The conversion costs a full pass over everything, and it has to be repeated whenever new material arrives in the original organisation.
Which is why the book's closing thought is worth keeping: it would be possible to use one structure for analysis and convert to another for generation, and that is a net win if the time saved exceeds the conversion. Converting is not a failure — it is a decision with a cost you can weigh.
Commit first
Answer, then rate your confidence.
Predict first
Why does the Markov program represent prefixes as tuples rather than lists, when lists are easier to add to and remove from?
Correct: Because prefixes are used as dictionary keys, and lists are unhashable.
Why: The book walks through this reasoning explicitly, and the shape of the argument is worth keeping. Your first choice might be a list, since the operations needed — remove from the beginning, add to the end — are exactly what lists do well. But the prefixes must also be dictionary keys, keys must be hashable, and hashable means immutable, which rules lists out. What you give up is small: with tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is the whole of shift — prefix[1:] + (word,), one slice and one concatenation. The easiest structure lost to a hard requirement on a different operation, and that is what data structure selection usually looks like.
Explain it
Random words gave nonsense; Markov analysis nearly gives sentences.
Discussion prompt
A classmate's frequency-based generator produces the right vocabulary and no grammar. Explain what Markov analysis adds and why it helps.
Hint: What does each choice depend on?
Answer:
Their generator picks each word independently, so it knows how often the appears and nothing about what usually follows it. A series of random words seldom makes sense because there is no relationship between successive words.
Markov analysis records, for each short prefix, every word that followed it in the source — so half the is always followed by bee, and the bee might be followed by has or is.
Generating from that mapping makes each choice depend on the last few words, which is enough for local grammar. The result is almost syntactically correct, but not quite — because a two-word window forgets the start of the sentence by the time it reaches the end.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The efficiency point is really the memo pattern again: work that does not change between calls should not be repeated in each one. Markov analysis is the most fun and the most mechanical once the picture is clear. The data-structure section is the chapter's actual subject and the easiest to skim, because it is almost all prose — but the requirement chain from key to hashable to immutable to build-don't-modify is worth being able to recite. And the five R's are the sort of advice that only becomes useful once you are genuinely stuck, which is exactly when it is hardest to remember.
Connect it up
One page, from memory.
Draw it
Draw the Markov mapping as a dictionary with two tuple keys and their suffix lists, and beside it write shift in one line, marking the slice and the singleton comma. Underneath, write the three representation choices — mapping, prefixes, suffixes — and beside each the reason it was decided, noting which of the three is still an open trade. Finally list the five R's with one word each on when to use them.
Recap
Four pages, and chapter 13 is finished: the case study the whole book has been building towards.
| If you remember one thing | It is this |
|---|---|
| From random words | Work that does not change between calls should not be inside the call. |
| From Markov analysis | Condition each choice on the last few words and local grammar appears. |
| From the prefix choice | Key means hashable means immutable means build, don't modify. |
| From the trade-offs | Write the easiest version and measure. Often you cannot tell in advance. |
| From the five R's | When several attempts have taught you nothing, change strategy rather than repeating. |
The next chapter turns to files and persistence: reading and writing, format strings, filenames and paths, catching exceptions with try and except, and pickling — the tools for a program whose data survive after it stops running.
Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-133 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.