13c Random Words, Markov Analysis, and Choosing a Data Structure

This lesson generates random words weighted by frequency, builds a Markov model mapping prefixes to suffixes, works through the data structure choices that model forces, and closes with the five strategies for debugging a hard problem.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 13c Random Words, Markov Analysis, and Choosing a Data Structure

Title

Python · Chapter 13 — Case study: data structure selection

§13.7-13.10, pp. 130-133

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-133 — the pages these objectives are drawn from

3. Before we start: why random words are not sentences

Warm-up

You can now pick words weighted by how often they appear. Try it.

Discussion prompt

Sampling ten words from Emma by frequency gives: this the small regard harriet which knightley's it most things. The word choice is right and the result is not a sentence. What is missing?

Hint: What does the usually come before?

Answer:

Any relationship between successive words. In a real sentence you would expect an article like the to be followed by an adjective or a noun, and probably not a verb or adverb.

The frequency model knows how often each word appears and nothing whatever about what follows it. Each pick is independent of the last.

So the fix is to make each choice depend on what came before — which is what Markov analysis does, and it is the rest of this lesson.

4. The one idea behind this lesson: condition on what came before

Concept

A series of random words seldom makes sense because there is no relationship between successive words. One way to measure these kinds of relationships is Markov analysis, which characterises, for a given sequence of words, the probability of the words that might come next.

Markov analysis — A way of characterising, for a sequence of words, the probability of the words that might come next.

The result is a mapping from each prefix to all possible suffixes. Given the mapping, you can generate random text by starting with any prefix, choosing at random from its possible suffixes, and repeating.

Figure (svg): Two columns contrasting independent word choice with choice conditioned on the preceding words

The same words, the same frequencies. Only the conditioning differs.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-131

5. Choosing a random word efficiently

Section

Section 1

6. The simplest algorithm, and its cost

Concept

To choose a random word from the histogram, the simplest algorithm is to build a list with multiple copies of each word, according to the observed frequency, and then choose from the list.

def random_word(h):
    t = []
    for word, freq in h.items():
        t.extend([word] * freq)
    return random.choice(t)
ExpressionWhat it doesNote
[word] * freqfreq copies of the stringlist repetition
t.extend(...)adds every elementunlike append
random.choice(t)uniform over positionsso weighted by frequency

The expression [word] * freq creates a list with freq copies of the string word, and the extend method is similar to append except that the argument is a sequence. This algorithm works, but it is not very efficient: each time you choose a random word, it rebuilds the list, which is as big as the original book.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-130

7. Picture it: the cost of rebuilding

Picture it

Every call reconstructs a list the size of the whole text.

Figure (svg): A growth chart comparing rebuilding the list per call against building it once or using cumulative sums

Three algorithms with the same output and very different costs.

The first is fine for one word and unusable for a thousand. That difference is the chapter's subject.

8. Worked example: extend against append

Worked example

The two are easy to confuse and do different things.

>>> t = ['a']
>>> t.append(['b', 'b'])
>>> t
['a', ['b', 'b']]
>>> t = ['a']
>>> t.extend(['b', 'b'])
>>> t
['a', 'b', 'b']
MethodWhat it addsResult
appendadds ONE elementwhich happens to be a list
extendadds each element of the argumentflattens one level
for random_wordextend is wantedthe copies must be separate

Recall what append does.

Why: It adds a single element, so appending a list nests it — lesson 10b's four wrong appends.

Compare with extend.

Why: The extend method is similar to append except that the argument is a sequence, and each of its elements is added separately.

See which random_word needs.

Why: The copies must be separate elements for choice to pick between them, so extend is the right one.

Figure (svg): The state of the program after each line of Worked example extend against append, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

append nests and extend flattens. Using append here would produce a list of lists, and choice would return a list rather than a word.

Verify: Check what the bug would look like.

Why: The function would return something like ['the', 'the', 'the'] rather than 'the' — a legal value of the wrong shape, which is a shape error in the sense of lesson 12c. It would not raise until the result was used as a string, several steps later.

9. Predict: what does the list repetition produce?

Prediction

A list times an integer.

word = 'bee'
freq = 3
print(['bee'] * freq)
PartWhat it doesResult
['bee']a one-element lista singleton
* 3list repetitionthree copies
the resultthree separate elementsnot a nested list

Predict first

What does this print?

  • ['bee', 'bee', 'bee']
  • 'beebeebee'
  • [['bee'], ['bee'], ['bee']]
  • ['beebeebee']

Correct: ['bee', 'bee', 'bee'] — the expression creates a list with freq copies of the string.

Why: List repetition copies the elements, not the list itself, so three separate strings end up at three positions. Note that 'bee' * 3 without the brackets would give option B, a repeated string — one pair of brackets is the difference between three list elements and one long word.

10. Worked example: the cumulative sum alternative

Worked example

The book's four-step improvement, which keeps the histogram.

# 1. the words
words = list(h.keys())
# 2. the running totals; the last is n, the book's length
cum = []
total = 0
for w in words:
    total = total + h[w]
    cum.append(total)
# 3. pick 1..n and bisect for the insertion point
# 4. that index gives the word
StepWhat it producesNote
the cumulative listrunning totals of frequenciesincreasing
the last itemthe total number of wordsn
a random 1..na uniform positionin the whole text
bisectionfinds where it fallsin about log n steps

Build the running totals once.

Why: Each word occupies a stretch of the cumulative list as wide as its frequency — the range picture from lesson 13a, made concrete.

Draw a position uniformly.

Why: Choose a random number from 1 to n, where n is the last item and the total number of words in the book.

Find it by bisection.

Why: Use a bisection search to find the index where the random number would be inserted, then use the index to find the corresponding word.

Figure (svg): A ladder showing cumulative frequency totals and the ranges each word occupies

Each word's range is as wide as its frequency, so a uniform draw is weighted correctly.

The same probabilities from a structure the size of the vocabulary rather than of the text, with each choice costing about log n steps rather than a full rebuild.

Verify: Check why bisection is possible at all.

Why: Because the cumulative sums are increasing by construction — every frequency is at least 1, so each total exceeds the last. A bisection search needs a sorted sequence, and this one is sorted for free, which is what makes the fast lookup available.

11. Trap: rebuilding the list inside a loop

Trap

The trap

A generator calls random_word a thousand times to produce a passage.

Call the function you have

Why: It returns a correctly weighted random word, which is what is needed.

Each call rebuilds a list of 161,080 strings, so a thousand words means building a hundred and sixty million list entries. The program is correct and takes minutes.

The fix

Build the expensive structure once.

An obvious improvement is to build the list once and then make multiple selections

Why: Which is the book's own first suggestion, and it is a large win for very little work.

Or use the cumulative sums, which are smaller as well as reusable

Why: Vocabulary-sized rather than text-sized.

This is the memo pattern from chapter 11 in another guise: work that does not change between calls should not be repeated in every call. Recognising that shape is worth more than either specific fix.

12. Discriminate: append or extend?

Discrimination

Ask whether the argument is one element or several.

Sort into buckets

For each goal, which method is right?

append
add one word to the list; add a list as a single nested element; add one tuple as one element
extend
add several copies of a word; add a whole other list's elements; add each character of a string
app
Each adds exactly one element, whatever its type. Appending a list or a tuple nests it, which is sometimes precisely what you want.
ext
Each adds the elements of a sequence separately. extend is similar to append except that the argument is a sequence, so a string is added one character at a time.

13. Complete it: add the copies

Faded example

Each copy must be its own element.

Fill in the blanks

for word, freq in h.items():
t.extend([word] * freq)

Why: extend adds each element of its argument separately, so the copies become separate list entries that choice can pick between. append would add the whole list as one nested element, and choice would then return a list of identical words rather than a word — a shape error that raises only when the result is used.

14. Think it through: why does rebuilding matter so much?

Socratic

The function is correct either way.

Discussion prompt

random_word gives the right answer every time. Why is it worth replacing, and when would it not be?

Hint: How many times will it be called?

Answer:

For one word it is perfect: simple, obviously correct, and the cost is a single pass over the histogram.

For a thousand words the cost is a thousand passes, each building a list as big as the book. The work is proportional to calls times text length, and both are large.

So the answer depends entirely on the use. The book's own advice in §13.9 is to choose the structure that is easiest to implement and see whether it is fast enough — and only if not, to improve it. Replacing this function before knowing how often it will be called would be optimising without evidence.

15. Markov analysis

Section

Section 2

16. A mapping from prefixes to possible suffixes

Concept

The result of Markov analysis is a mapping from each prefix to all possible suffixes. Given this mapping, you can generate a random text by starting with any prefix and choosing at random from the possible suffixes.

# from 'Eric, the Half a Bee':
#   'half the'  ->  always followed by 'bee'
#   'the bee'   ->  might be followed by 'has' or 'is'
#   'a bee'     ->  'philosophically', 'be' or 'due'
PartWhat it isNote
a prefixa short sequence of wordstwo, in this example
its suffixesevery word that followed itwith repeats allowed
generationpick a suffix, shift the prefixand repeat

In the example the length of the prefix is always two, but you can do Markov analysis with any prefix length — and the exercise asks you to write the program in a way that makes it easy to try others.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 131-131

17. Picture it: the mapping the analysis produces

Picture it

Each prefix, and everything observed after it.

Figure (svg): A dictionary mapping two-word prefixes to the lists of words observed after them

Tuple keys, list values. Both choices are argued for in idea 3.

Notice the first prefix has only one possible suffix, so generation from it is forced. Variety comes from prefixes with several.

18. Worked example: building the map

Worked example

One pass over the words, keeping a sliding window.

def build(words, n=2):
    d = {}
    prefix = tuple(words[:n])
    for word in words[n:]:
        d.setdefault(prefix, []).append(word)
        prefix = shift(prefix, word)
    return d
PartWhat it doesNote
prefixthe first n wordsas a tuple
for each later wordit is a suffix of the current prefixrecord it
shiftdrop the first, add this wordthe next prefix

Start with the first n words.

Why: The prefix is a tuple, for reasons idea 3 sets out — it has to be usable as a dictionary key.

Record each following word as a suffix.

Why: setdefault creates an empty list on a first sighting and returns the existing one otherwise, which is invert_dict's singleton pattern in one call.

Slide the window forward.

Why: The end of the prefix and the new suffix combine to form the next prefix, and repeat.

Figure (svg): A pipeline showing a sliding window over words producing prefix-suffix pairs

The window advances one word at a time, and each position contributes one pair.

A dictionary from prefixes to lists of observed suffixes, built in one pass with a window sliding along the text.

Verify: Check the count of entries.

Why: Every word after the first n contributes exactly one suffix, so the total number of suffixes across all the lists equals len(words) - n. Summing the list lengths and comparing is a consistency check of the kind lesson 11c described, and it catches a window that skips or repeats a position.

19. Predict: which word must come next?

Prediction

One prefix appears only once in the text.

# in 'Eric, the Half a Bee':
# the prefix ('half', 'a') appears once,
# followed by 'bee'
PartWhat is trueResult
('half', 'a')one occurrenceone suffix recorded
its suffix list['bee']a single element
random.choiceno alternative'bee'

Predict first

Starting from the prefix ('half', 'a'), what does generation produce next?

  • 'bee', necessarily — it is the only recorded suffix
  • Any word from the text, chosen at random
  • Nothing — the prefix has no suffixes
  • The commonest word in the text

Correct: 'bee', necessarily — the prefix only appears once in the text, so there is only one possible suffix.

Why: The book makes this point directly: if you start with the prefix Half a, then the next word has to be bee. Variety in the generated text comes from prefixes with several recorded suffixes, like ('a', 'bee'), which might be followed by philosophically, be or due.

20. Worked example: generating text

Worked example

Pick a suffix, shift, repeat.

def generate(d, prefix, count):
    for i in range(count):
        suffixes = d[prefix]
        word = random.choice(suffixes)
        print(word, end=' ')
        prefix = shift(prefix, word)
StepWhat it doesNote
d[prefix]the possible next wordsa list
random.choiceone of themweighted by repeats
shiftthe new prefixend of old plus new word

Look up the prefix.

Why: The mapping gives every word observed after this prefix, with repeats — so a word that followed it often is in the list often.

Choose one at random.

Why: choice is uniform over positions, and the repeats do the weighting — exactly the mechanism from lesson 13a.

Form the next prefix.

Why: Combine the end of the prefix and the new suffix, and repeat.

Figure (svg): The state of the program after each line of Worked example generating text, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Text that follows the source's local patterns. The book's Emma sample is almost syntactically correct, but not quite; semantically, it almost makes sense, but not quite.

Verify: Ask what happens at a prefix with no entry.

Why: d[prefix] raises KeyError — which happens if generation reaches the very last prefix in the text, since no word followed it. A real implementation has to handle that, either by restarting from a random prefix or by using get with a fallback, and noticing it before it happens is the useful skill.

21. Trap: collapsing the suffix list into a set of distinct words

Trap

The trap

A program stores each prefix's suffixes without duplicates, to save space.

Avoid storing the same word repeatedly

Why: It looks wasteful to keep ten copies of the.

The duplicates were carrying the probabilities. Without them every possible continuation becomes equally likely, and a word that followed once competes evenly with one that followed a hundred times — so the generated text stops resembling the source.

The fix

Keep the repeats, or keep counts instead.

A list with duplicates

Why: The simplest option, and choice weights it automatically.

Or a histogram of suffixes

Why: One entry per distinct word with its count — same information, less space, harder to sample from.

That trade is exactly the one §13.9 discusses. Either is defensible; discarding the frequencies altogether is not, because they are the model.

22. Watch the window: two steps of generation

Invariant

Follow the prefix as it shifts.

Step through it

What stays the same about the prefix across all four frames?

  1. The starting prefix, two words long as in the book's example.
  2. Its suffix list holds one word, so the choice is forced.
  3. The new prefix drops the first word and adds the chosen one: the end of the old prefix plus the new suffix.
  4. This prefix has three recorded suffixes, so here the generated text can branch.

Its length. The window is always two words wide — it slides rather than growing, which is what makes the number of distinct prefixes manageable and the model a fixed size.

23. Complete it: record a suffix

Faded example

A first sighting needs a list to append to.

Fill in the blanks

d.setdefault(prefix, []).append(word)

Why: setdefault returns the existing list for a known prefix and installs a new empty one for a first sighting, so append always has a list to work with — the singleton pattern from invert_dict, in a single call. Using get instead would return the default without storing it, so nothing would ever accumulate.

24. Where Markov models are used

Real world

Predicting the next thing from the last few is a very general idea.

Discussion prompt

Where have you seen something predict what comes next from what came just before? What are its prefixes and suffixes?

Hint: Check your phone.

Answer:

Predictive text and autocomplete are the obvious ones: the prefix is what you have typed and the suffixes are the words that usually follow it.

So are music recommendation by what you just played, route prediction from the last few turns, and the sequence models underlying much larger systems.

What they share is the assumption that the recent past is enough — that you need not remember the whole history to predict the next step. That assumption is what makes the model small and what makes its output almost make sense but not quite, since the sentence's beginning is forgotten by the time its end is generated.

25. Choosing how to represent a prefix

Section

Section 3

26. The choice is decided by one requirement

Concept

In your solution you had to choose how to represent the prefixes, how to represent the collection of possible suffixes, and how to represent the mapping. The last one is easy: a dictionary is the obvious choice for a mapping from keys to corresponding values.

The first step is to think about the operations you will need to implement for each. For the prefixes, we need to be able to remove words from the beginning and add to the end — and we also need to use them as keys, which is what settles it.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 132-132

27. Picture it: the requirement that rules out a list

Picture it

Two operations are needed, and only one structure supports both.

Figure (svg): Two columns comparing a list and a tuple against the two requirements for a prefix

Your first choice might be a list, since it is easy to add and remove elements — but they cannot be keys.

So the easier structure loses to a requirement that has nothing to do with ease. That is what data structure selection usually looks like.

28. Worked example: the shift function

Worked example

Building the next prefix without modifying anything.

def shift(prefix, word):
    return prefix[1:] + (word,)
PartWhat it producesNote
prefix[1:]everything but the first worda new tuple
(word,)a one-element tuplenote the comma
the concatenationa new tuplenothing was modified

Drop the first word.

Why: Slicing from index 1 gives a new tuple containing everything after the beginning.

Make the new word into a tuple.

Why: The trailing comma is required — without it, ('bee') is a string and the concatenation raises a TypeError.

Join them.

Why: shift takes a tuple of words and a string, and forms a new tuple that has all the words in prefix except the first, with word added to the end.

Figure (svg): A prefix tuple with its first element dropped and a new word appended

One element out at the front, one in at the back, and the width never changes.

A new prefix, built rather than modified. With tuples you can't append or remove, but you can use the addition operator to form a new tuple.

Verify: Check that the length is preserved.

Why: The slice removes one element and the concatenation adds one, so the prefix stays exactly n words wide however many times shift is called. That invariant is what keeps the model a fixed size, and it would be worth an assertion in a longer program.

29. Predict: what does shift return?

Prediction

A slice and a concatenation.

def shift(prefix, word):
    return prefix[1:] + (word,)

print(shift(('half', 'a'), 'bee'))
PartValueNote
prefix[1:]('a',)the first is dropped
(word,)('bee',)a singleton tuple
the sum('a', 'bee')still two words

Predict first

What does this print?

  • ('a', 'bee')
  • ('half', 'a', 'bee')
  • ('half', 'bee')
  • A TypeError about concatenating a string and a tuple

Correct: ('a', 'bee') — the first word is dropped and the new one added, leaving the prefix two words wide.

Why: The slice removes one element and the concatenation adds one, so the length is preserved — which is what keeps the window sliding rather than growing. Option D is what would happen if the comma were omitted: ('bee') is a string, and adding a string to a tuple raises TypeError.

30. Worked example: why not a string?

Worked example

The third option, and it nearly works.

# a string prefix
prefix = 'half a'
# shifting it means splitting and rejoining:
words = prefix.split()
prefix = ' '.join(words[1:] + [word])
AspectWhat is trueNote
a string prefixhashable, so usable as a keythe requirement is met
shiftingsplit, slice, jointhree operations
the riska word containing a spacewould split wrongly

Check the key requirement.

Why: Strings are immutable and hashable, so a string prefix could be a dictionary key — the requirement that ruled out lists does not rule this out.

Check the shifting operation.

Why: It needs splitting into words and rejoining, which is more work than slicing a tuple and less direct.

Notice the separator problem.

Why: The words are joined by a space, so the structure depends on no word containing one — the compound-key hazard from lesson 12c.

Figure (svg): The state of the program after each line of Worked example why not a string , drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A string works and is worse: more work to shift, and it flattens the word boundaries into a separator that could in principle collide.

Verify: Compare what each structure preserves.

Why: The tuple keeps the words as separate elements, so prefix[0] is a word and shifting is a slice. The string has to reconstruct the boundaries every time, and it can only do so by convention. Keeping structure rather than encoding it into text is the general lesson, and it is the same reason lesson 12c preferred tuple keys to joined strings.

31. Trap: trying to modify the prefix in place

Trap

The trap

A student writes prefix.pop(0) and prefix.append(word) to slide the window.

Use the operations that fit the job

Why: Removing from the front and adding to the back is exactly what a list does well.

Those methods do not exist on a tuple, so it raises AttributeError — and switching to a list to get them makes the prefix unusable as a dictionary key, which is why the tuple was chosen.

The fix

Build a new prefix instead of changing one.

return prefix[1:] + (word,)

Why: A slice and a concatenation, both of which create rather than modify.

Assign the result

Why: prefix = shift(prefix, word), since nothing was changed in place.

The requirement chain is worth stating once: the prefix must be a key, keys must be hashable, hashable means immutable, and immutable means building rather than modifying. Each step follows from the one before.

32. Rank: the reasoning that picks the prefix type

Ranking

Four steps, in the order the argument runs.

Put in order

  1. a list is easiest for adding and removing
  2. the prefix has to be a dictionary key
  3. keys must be hashable, so the prefix must be immutable
  4. a tuple, with a new one built by slicing and concatenating

Why: The list is the first thought, because the operations needed are the ones lists do well. Then the key requirement appears, which forces hashability, which forces immutability — and the tuple is what is left. The conclusion is reached by elimination rather than by the operations, which is why the easiest structure is not the chosen one.

33. Complete it: form the next prefix

Faded example

The comma is doing essential work.

Fill in the blanks

def shift(prefix, word):
return prefix[1:] + (word,)

Why: Without the comma, (word) is just the string and adding a string to a tuple raises TypeError: can only concatenate tuple to tuple. The singleton comma from lesson 12a is load-bearing here, and its absence produces an error rather than silence — which for once is a mercy.

34. Explain it: why not use a list, which is easier?

Explain it

The obvious structure is the wrong one, for a reason worth stating.

Discussion prompt

A classmate points out that a list handles adding and removing much more naturally than a tuple. Explain why the program uses tuples anyway.

Hint: What is the prefix used for besides shifting?

Answer:

Shifting is not the only thing a prefix does. It is also the key of the dictionary, and that is the requirement that decides everything.

Keys must be hashable, and a list is not — chapter 11's argument, that a mutable key could move after it was filed. So the list is ruled out however convenient its methods are.

What you lose is small: with tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is what shift does in one line. Ease of one operation loses to a hard requirement on another, which is what data structure selection usually comes down to.

35. Weighing the options: time, space, and effort

Section

Section 4

36. Three factors, and one piece of practical advice

Concept

So far we have been talking mostly about ease of implementation, but there are other factors to consider in choosing data structures.

benchmarking — The process of choosing between data structures by implementing alternatives and testing them on a sample of the possible inputs.

A practical alternative is to choose the data structure that is easiest to implement, and then see if it is fast enough for the intended application. If so, there is no need to go on.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 132-133

37. Picture it: the two suffix representations

Picture it

Same information, different shapes, different costs.

Figure (svg): Two columns comparing a list of suffixes with a histogram of suffixes

Adding is equally easy. Choosing and storing are where they differ.

Neither is right in the abstract. The book's advice is to take the easier one and measure before replacing it.

38. Worked example: the case for each

Worked example

Two structures, four considerations.

# list of suffixes
d[prefix].append(word)
random.choice(d[prefix])

# histogram of suffixes
h = d[prefix]
h[word] = h.get(word, 0) + 1
choose_from_hist(h)      # needs the walk from 13a
OperationWhich is easierNote
addingone line either wayequally easy
choosingone call, or a whole functionthe list wins
spaceone entry per occurrence, or per wordthe histogram wins

Compare adding.

Why: Adding a new suffix, or increasing the frequency of an existing one, is equally easy for the list implementation or the histogram.

Compare choosing.

Why: Choosing a random element from a list is easy; choosing from a histogram is harder to do efficiently, as exercise 13.7 showed.

Compare space.

Why: Using a histogram might take less space, because you only have to store each word once, no matter how many times it appears in the text.

Figure (svg): The state of the program after each line of Worked example the case for each, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The list is easier to sample and larger; the histogram is smaller and harder to sample. Which matters depends on the text's size and how much generation you do.

Verify: Ask when space would decide it.

Why: The book says that in some cases saving space can also make your program run faster, and in the extreme, your program might not run at all if you run out of memory — but that for many applications, space is a secondary consideration after run time. So the histogram wins only when the text is large enough for memory to bite.

39. Sort: which suffix structure wins on this?

Sorting

Four considerations, split between the two.

Sort into buckets

For each consideration, which representation is better?

a list of suffixes
choosing a random suffix; simplicity of the code; getting the weighting for free
a histogram of suffixes
storage space for a large text; storing each distinct word only once; counting how often a word followed
list
Choosing a random element from a list is easy, and the repeats supply the weighting with no extra code. Simplicity is the list's whole case.
hist
Each distinct word is stored once with its count, which saves space on a large text and makes the frequency directly readable. The cost is that sampling from it is harder to do efficiently.

40. Worked example: measuring rather than guessing

Worked example

Two ways to find out, and one way to avoid needing to.

# benchmarking: implement both and compare
# on a sample of the possible inputs

# or: use the profile module to find
# where the program actually spends its time

# or, practically: write the easy one and
# see whether it is fast enough
ApproachWhat it involvesNote
benchmarkingimplement both, measuredefinitive and expensive
profilingfind the slow part firstbefore optimising anything
the practical routeeasiest first, measure afterthe book's advice

Consider benchmarking.

Why: One option is to implement both and see which is better — definitive, and it costs writing the program twice.

Consider profiling.

Why: There are tools, like the profile module, that can identify the places in a program that take the most time — which tells you where optimising would even help.

Take the practical route first.

Why: Choose the structure that is easiest to implement, and then see if it is fast enough for the intended application. If so, there is no need to go on.

Figure (svg): A decision flowchart for choosing and then improving a data structure

Measure before optimising, and optimise only what the measurement points at.

Three approaches in increasing order of effort, and the book recommends starting with the cheapest. Optimising before measuring is work spent on a problem you have not confirmed you have.

Verify: Ask what profiling protects you from.

Why: Optimising the wrong thing. A program that spends ninety per cent of its time reading the file will not get noticeably faster however cleverly the suffixes are stored — and without measuring, there is no way to know that before doing the work.

41. Trap: optimising before measuring

Trap

The trap

A student implements the histogram version of the suffixes first, because it uses less memory.

Choose the more efficient structure from the start

Why: Why write something you know you will replace?

The harder version takes longer to write, is harder to get right, and may make no measurable difference — and until the program runs, there is no evidence that memory was ever the constraint.

The fix

Write the easy one and measure.

Choose the structure that is easiest to implement

Why: The list, here, and choice does the weighting for free.

Then see if it is fast enough for the intended application

Why: If it is, there is no need to go on.

The book's closing thought is worth keeping too: since analysis and generation are separate phases, you could use one structure for analysis and convert to another for generation — a net win if the time saved during generation exceeded the time spent converting.

42. Predict: which advice does the book give?

Prediction

Two reasonable-sounding strategies.

# strategy A: pick the fastest structure up front
# strategy B: pick the easiest, then measure
StrategyWhat it assumesNote
up frontrequires knowing the answeroften you do not
easiest firstand see if it is fast enoughthe book's advice
if notprofile, then benchmarkin that order

Predict first

What does the book recommend as the practical approach?

  • Choose the structure that is easiest to implement, then see if it is fast enough
  • Always benchmark both implementations before choosing
  • Always prefer the structure that uses less memory
  • Always prefer a dictionary, since in is faster on it

Correct: Choose the structure that is easiest to implement, then see if it is fast enough — and if so, there is no need to go on.

Why: Benchmarking is named as an option and it means writing the program twice, so it is the expensive route rather than the default. The book is explicit that often you don't know ahead of time which implementation will be faster, which is exactly why measuring the easy version first is the practical order.

43. Two truths and a lie: choosing structures

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. The in operator is faster for dictionaries than for lists, at least when there are many elements
  • B. You could use one structure for analysis and convert to another for generation
  • C. Saving storage space always makes a program run faster

Survives elimination: C

Why: C overstates a hedged claim. The book says that in SOME cases saving space can also make your program run faster, and that in the extreme a program might not run at all if it runs out of memory — but also that for many applications, space is a secondary consideration after run time. Sometimes and always are very different claims.

44. Where *easiest first* is the right order

Real world

The advice generalises well beyond data structures.

Discussion prompt

Think of a time effort went into making something efficient before anyone knew whether it needed to be. What was the cost?

Hint: What else could that effort have gone into?

Answer:

Elaborate systems built for a scale that never arrived; optimisations that made code harder to change; time spent on a bottleneck that turned out not to be one.

The cost is rarely the wasted effort alone — it is that the complicated version is harder to modify when the actual requirement shows up somewhere else entirely.

Which is why the order matters: build the simple version, find out where it hurts, then fix that. The profile module exists precisely because the slow part is usually not where people guess it is.

45. The five R's of debugging

Section

Section 5

46. Five things to try on a hard bug

Concept

When you are debugging a program, and especially if you are working on a hard bug, there are five things to try.

The list is ordered roughly by how much you have already tried. Retreating is last because it discards work, and it is on the list because sometimes discarding work is the fastest route.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 133-133

47. Picture it: five approaches, and what each assumes

Picture it

Each works on a different kind of stuck.

Figure (svg): Two columns pairing each debugging strategy with the situation it suits

Five different kinds of stuck, and a different move for each.

Recognising which kind you are in is most of the value. Doing more of what has already failed is the commonest debugging mistake.

48. Worked example: ruminating, with the book's questions

Worked example

Thinking is a technique with specific questions attached.

# What kind of error is it?
#   syntax, runtime, or semantic
# What can the error message tell me?
# What kind of error could cause this?
# What did I change last, before the problem appeared?
QuestionWhat it narrowsNote
the kindnarrows where to looksyntax is at the point named
the messagenames a type or a variablealready a strong clue
what changed lastthe single best questionand easy to forget

Classify the error.

Why: Syntax, runtime or semantic — the three kinds from chapter 1, and each has a different search strategy.

Mine the message.

Why: What information can you get from the error messages, or from the output of the program? A TypeError naming NoneType is a different investigation from a ValueError about unpacking.

Ask what changed.

Why: What did you change last, before the problem appeared? A program that worked an hour ago has a bug in the difference.

Figure (svg): The state of the program after each line of Worked example ruminating, with the book's questions, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Four specific questions rather than a vague instruction to think. The last one is the most powerful and the easiest to skip when you are frustrated.

Verify: Check the last question against version control.

Why: If the change since the last working version is small, the bug is in it — which turns an open-ended search into reading a handful of lines. That is why committing often is a debugging technique as much as a bookkeeping one.

49. Sort: which R fits this situation?

Sorting

Each kind of stuck has its own move.

Sort into buckets

For each situation, which strategy is most likely to help?

Running
you cannot see what a variable holds at that point
Reading
you are sure the code is right but it is not
Ruminating
you have an error message you have not really read; the program worked an hour ago
Rubberducking
you cannot describe what the function is supposed to do
Retreating
you have made twenty changes and nothing works any more
run
You need to see inside the program at a particular point, which means displaying the right thing at the right place — or building scaffolding to do it.
read
Being sure the code is right while it is not is exactly the case for reading it back to yourself and checking that it says what you meant to say.
rum
Both have information already available and unused: an unread error message, and the knowledge of what changed last before the problem appeared.
duck
If you cannot describe what it should do, explaining it aloud will surface that — and you often find the answer before you finish asking.
ret
Twenty changes deep with nothing working means backing off, undoing until you reach a program that works and that you understand.

50. Worked example: rubberducking, and why it works

Worked example

The listener is optional.

# 'So this function takes the prefix,
#  looks up the suffixes, picks one at random,
#  and then shifts the prefix - except it
#  doesn't, because I never assigned the result...'
StageWhat happensNote
explainingforces a complete accountno steps assumed
the gapappears mid-sentencebefore the question is finished
the listenernot requireda rubber duck will do

Explain the problem out loud.

Why: If you explain the problem to someone else, you sometimes find the answer before you finish asking the question.

Notice why it works.

Why: Explaining forces you to say every step, including the ones you have been assuming — and the assumption is usually where the bug is.

Notice the listener is optional.

Why: Often you don't need the other person; you could just talk to a rubber duck. That is the origin of the well-known strategy called rubber duck debugging.

Figure (svg): A flowchart showing explaining a program aloud surfacing an unstated assumption

The step you cannot articulate is the step you never verified.

A technique that works because of what it does to the explainer, not to the listener. The book is emphatic that it is not making this up.

Verify: Ask why reading silently is not the same.

Why: Reading lets you skim the parts you believe you understand, which are exactly the parts hiding the bug. Explaining aloud forces every step to be stated, and a step you cannot state is one you have not checked — which is what makes the technique work when re-reading has already failed.

51. Trap: running more versions when reading would do

Trap

The trap

A program is wrong, so changes are made and it is run again — repeatedly, for an hour.

Experiment until something works

Why: Running is on the list, and each attempt feels like progress.

Without a hypothesis, each run tests nothing, and the program drifts further from the version that was understood. After an hour there is neither a fix nor a program anyone can explain.

The fix

Change strategy when one stops producing information.

Read, or ruminate, or explain it to someone

Why: Each uses different evidence, so switching gives you something new.

And if the program is no longer understood, retreat

Why: Back off, undoing recent changes, until you get back to a program that works and that you understand. Then start rebuilding.

The five are alternatives rather than an escalation. The signal to switch is that the last several attempts taught you nothing, which is a much better trigger than frustration.

52. Predict: which question is the most useful?

Prediction

Four questions from the Ruminating list.

# The program worked this morning.
# Now it raises a TypeError.
FactWhat it gives youNote
it worked beforethe bug is in the differencea bounded search
the error kindnarrows the familyuseful but broader
what changed lastthe strongest clue hereand easy to skip

Predict first

Which question narrows the search the most in this situation?

  • What did I change last, before the problem appeared?
  • What kind of error is it?
  • What does the output look like?
  • Which line does the traceback name?

Correct: What did I change last, before the problem appeared? — a program that worked this morning has its bug in the difference.

Why: The other three are all useful and all search the whole program. This one bounds the search to whatever changed since the last working version, which may be a handful of lines. It is also the question people skip when frustrated, which is why the book lists it explicitly rather than leaving it to take some time to think.

53. Explain it yourself: why is retreating on the list?

Explain it to yourself

It discards work, which feels like the opposite of progress.

Discussion prompt

Undoing changes throws away effort. Why is it a debugging strategy rather than an admission of defeat?

Hint: What is the alternative costing you?

Answer:

Because the changes may be the problem. A program that has drifted through twenty edits contains twenty opportunities for a new bug, and the original one may already be fixed underneath them.

And because debugging requires understanding. Once you no longer know what the program does, every further change is a guess, and guesses are what got you here.

So the point of retreating is to get back to a program that works and that you understand, and then rebuild deliberately. The work is not wasted — what you learned about the bug survives; only the flailing edits are discarded.

54. Two truths and a lie: the five R's

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. Rubberducking works even with no listener present
  • B. Ruminating includes asking what you changed last, before the problem appeared
  • C. Retreating is a last resort that means giving up on the bug

Survives elimination: C

Why: C misreads what retreating is for. Backing off until you reach a program that works and that you understand is a step towards the fix, not away from it — because debugging requires understanding, and once that is gone every further change is a guess. The book's own phrasing is that then you can start rebuilding.

55. Compare: the three choices Markov analysis forces

Comparison

Fill the blanks. One is obvious, one is forced, and one is a genuine trade.

Comparison matrix

What to representThe choiceWhy
the mappinga dictionaryit is a mapping from keys to values
the prefixesa tuple of stringsthey must be keys, so they must be hashable
the suffixesa list, or a histograma genuine trade: easy sampling against less space
how to decide the last onewrite the easiest, then measureyou often cannot tell in advance

Only the third row is an open question, and the book's answer is to resolve it by measurement rather than by argument.

56. The procedure: building and using a Markov model

Pattern

Six steps. The third is where the data structure decision lands.

  1. Read the text into a list of words, deciding first whether to keep punctuation attached.
  2. Take the first n words as the starting prefix, as a tuple.
  3. For each remaining word: record it as a suffix of the current prefix, in a dictionary keyed by that tuple.
  4. Shift the prefix — drop the first word, append the new one — with prefix[1:] + (word,).
  5. To generate, look up the current prefix, choose a suffix at random, print it, and shift.
  6. Handle the prefix with no recorded suffixes, which happens at the end of the source text.

Step 4 must build a new tuple rather than modify one, because a tuple cannot be modified — and it must be a tuple because step 3 uses it as a key.

Python documentation — random — Generate pseudo-random numbers random — Generate pseudo-random numbers

57. Check yourself 1 of 3: the prefix type

Check

One requirement decides it.

Check your understanding

Why must a Markov prefix be a tuple rather than a list?

  • A. Because it is used as a dictionary key, and keys must be hashable (correct)
  • B. Because tuples are faster to slice than lists
  • C. Because a list cannot hold strings
  • D. Because tuples use less memory

Answer: A

Why: Your first choice might be a list, since it is easy to add and remove elements — but the prefixes also have to be keys in a dictionary, and that rules out lists. With tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is what shift does.

Why B tempts people
Speed is not the argument. The list's operations are easier, which is why it is the first thought.
Why C tempts people
Lists hold values of any type, strings included. Nothing about the contents rules a list out.
Why D tempts people
Memory is a secondary consideration and not the reason given. Hashability is a hard requirement rather than a preference.

58. Check yourself 2 of 3: shift

Check

One element out, one in.

def shift(prefix, word):
    return prefix[1:] + (word,)

print(shift(('the', 'bee'), 'is'))
PartValueNote
prefix[1:]('bee',)the first is dropped
+ ('is',)concatenationa new tuple
the result('bee', 'is')still two words

Check your understanding

What does this print?

  • A. ('bee', 'is') (correct)
  • B. ('the', 'bee', 'is')
  • C. ('the', 'is')
  • D. A TypeError

Answer: A

Why: The slice removes the first word and the concatenation appends the new one, so the prefix stays two words wide. That length invariant is what makes the window slide rather than grow, and it keeps the number of distinct prefixes bounded.

Why B tempts people
This would be the result of appending without slicing, and the prefix would grow with every step.
Why C tempts people
This drops the wrong word. prefix[1:] keeps everything except the first.
Why D tempts people
The comma makes (word,) a tuple, so tuple plus tuple is fine. Omitting it would give the TypeError.

59. Check yourself 3 of 3: debugging

Check

Twenty changes deep, and nothing works.

Check your understanding

You have made many changes trying to fix a bug and the program is now worse and no longer understood. Which strategy does the book recommend?

  • A. Retreating — back off, undoing changes until you reach a program that works and that you understand (correct)
  • B. Running — keep making changes and testing them
  • C. Reading — go through the current version line by line
  • D. Ruminating — think harder about the original error message

Answer: A

Why: At some point the best thing to do is back off, undoing recent changes, until you get back to a program that works and that you understand — then you can start rebuilding. Debugging requires understanding, and once that is gone, further changes are guesses.

Why B tempts people
More of what has already failed. Twenty changes without progress is the signal to stop, not to continue.
Why C tempts people
Reading is useful, but the current version now contains twenty edits' worth of new opportunities for bugs.
Why D tempts people
The original message may no longer even apply, since the program has changed substantially since it appeared.

60. Where this shows up outside this course

Real world

Choosing how to store something before you know what you will ask of it.

Discussion prompt

Think of a time information was organised one way and the question that later mattered needed it another way. What did the conversion cost?

Hint: Filed by date, and you needed it by person.

Answer:

Records filed by date when you need them by customer; photographs in folders by year when you want them by who is in them; notes organised by source when you want them by topic.

The conversion costs a full pass over everything, and it has to be repeated whenever new material arrives in the original organisation.

Which is why the book's closing thought is worth keeping: it would be possible to use one structure for analysis and convert to another for generation, and that is a net win if the time saved exceeds the conversion. Converting is not a failure — it is a decision with a cost you can weigh.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

Why does the Markov program represent prefixes as tuples rather than lists, when lists are easier to add to and remove from?

  • Because prefixes are used as dictionary keys, and lists are unhashable
  • Because tuples are faster to slice
  • Because lists cannot be concatenated
  • Because tuples take less memory

Correct: Because prefixes are used as dictionary keys, and lists are unhashable.

Why: The book walks through this reasoning explicitly, and the shape of the argument is worth keeping. Your first choice might be a list, since the operations needed — remove from the beginning, add to the end — are exactly what lists do well. But the prefixes must also be dictionary keys, keys must be hashable, and hashable means immutable, which rules lists out. What you give up is small: with tuples you cannot append or remove, but you can use the addition operator to form a new tuple, which is the whole of shift — prefix[1:] + (word,), one slice and one concatenation. The easiest structure lost to a hard requirement on a different operation, and that is what data structure selection usually looks like.

62. Explain it to someone else

Explain it

Random words gave nonsense; Markov analysis nearly gives sentences.

Discussion prompt

A classmate's frequency-based generator produces the right vocabulary and no grammar. Explain what Markov analysis adds and why it helps.

Hint: What does each choice depend on?

Answer:

Their generator picks each word independently, so it knows how often the appears and nothing about what usually follows it. A series of random words seldom makes sense because there is no relationship between successive words.

Markov analysis records, for each short prefix, every word that followed it in the source — so half the is always followed by bee, and the bee might be followed by has or is.

Generating from that mapping makes each choice depend on the last few words, which is enough for local grammar. The result is almost syntactically correct, but not quite — because a two-word window forgets the start of the sentence by the time it reaches the end.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • Choosing a random word efficiently, and why rebuilding is costly
  • Markov analysis: prefixes, suffixes, and the shifting window
  • The data structure choices, and why prefixes must be tuples
  • The five debugging strategies

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The efficiency point is really the memo pattern again: work that does not change between calls should not be repeated in each one. Markov analysis is the most fun and the most mechanical once the picture is clear. The data-structure section is the chapter's actual subject and the easiest to skim, because it is almost all prose — but the requirement chain from key to hashable to immutable to build-don't-modify is worth being able to recite. And the five R's are the sort of advice that only becomes useful once you are genuinely stuck, which is exactly when it is hardest to remember.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the Markov mapping as a dictionary with two tuple keys and their suffix lists, and beside it write shift in one line, marking the slice and the singleton comma. Underneath, write the three representation choices — mapping, prefixes, suffixes — and beside each the reason it was decided, noting which of the three is still an open trade. Finally list the five R's with one word each on when to use them.

65. What you can do now

Recap

Four pages, and chapter 13 is finished: the case study the whole book has been building towards.

If you remember one thingIt is this
From random wordsWork that does not change between calls should not be inside the call.
From Markov analysisCondition each choice on the last few words and local grammar appears.
From the prefix choiceKey means hashable means immutable means build, don't modify.
From the trade-offsWrite the easiest version and measure. Often you cannot tell in advance.
From the five R'sWhen several attempts have taught you nothing, change strategy rather than repeating.

The next chapter turns to files and persistence: reading and writing, format strings, filenames and paths, catching exceptions with try and except, and pickling — the tools for a program whose data survive after it stops running.

Think Python, 2nd edition — Allen B. Downey §13.7-13.10, pp. 130-133 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §13.7-13.10, pp. 130-133
  2. Python documentation — random — Generate pseudo-random numbers
  3. Python documentation — Data Structures

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108