This lesson gives the book's solution to the word-frequency exercises: reading a file into a histogram, counting it two ways, finding the commonest words by sorting tuples, writing functions with optional parameters, and subtracting one dictionary from another.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 13 — Case study: data structure selection
§13.3-13.6, pp. 127-129
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-129 — the pages these objectives are drawn from
Warm-up
You have a dictionary of word counts and want the ten commonest.
Discussion prompt
sorted(hist.items()) gives you the items in order — but in order of what? What would you have to change to order them by frequency instead?
Hint: Which element of each tuple comes first?
Answer:
Each item is a (word, count) tuple, and tuples compare element by element from the left — so it sorts alphabetically by word, which is not what you asked for.
To sort by frequency, the frequency has to be at position 0. Which means building the tuples the other way round: (count, word).
That is exactly what the book's most_common does, and it is why chapter 12's comparison rule was worth learning: the order of fields inside a tuple decides what a plain sort does.
Concept
To find the most common words, we can make a list of tuples, where each tuple contains a word and its frequency, and sort it. In each tuple, the frequency appears first, so the resulting list is sorted by frequency.
def most_common(hist):
t = []
for key, value in hist.items():
t.append((value, key)) # frequency FIRST
t.sort(reverse=True)
return t| Line | What it does | Note |
|---|---|---|
| hist.items() | (word, frequency) pairs | the dictionary's order |
| (value, key) | swapped: (frequency, word) | the sort field first |
| t.sort(reverse=True) | highest frequency first | one call, no options |
Nothing tells sort what to order by. The ordering comes from the position of the frequency inside each tuple, which is a decision made two lines earlier.
Figure (svg): Two columns contrasting tuples built word-first with tuples built frequency-first
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-128
Section
Section 1
Concept
Here is a program that reads a file and builds a histogram of the words in it. process_file loops through the lines of the file, passing them one at a time to process_line, and the histogram is being used as an accumulator.
import string
def process_file(filename):
hist = dict()
fp = open(filename)
for line in fp:
process_line(line, hist)
return hist| Line | What it does | Note |
|---|---|---|
| hist = dict() | the accumulator | created once |
| for line in fp | one line per pass | a file is iterable |
| process_line(line, hist) | the SAME dictionary is passed | modified in place |
| return hist | everything accumulated | one dictionary |
The histogram is passed to process_line and modified there, which works because a dictionary is mutable and a function receives a reference to it — chapter 10's list-arguments point, in the place it is genuinely useful.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-127
Picture it
Every line adds to the same accumulator.
Figure (svg): A call diagram showing process_file passing one histogram to many calls of process_line
Notice process_line returns nothing. It modifies the accumulator, which is a design decision the caller has to know about — and lesson 10c's rule says it should therefore return None, which it does.
Worked example
The steps lesson 13a described, as the book writes them.
def process_line(line, hist):
line = line.replace('-', ' ')
for word in line.split():
word = word.strip(string.punctuation + string.whitespace)
word = word.lower()
hist[word] = hist.get(word, 0) + 1| Step | What it does | Note |
|---|---|---|
| replace('-', ' ') | hyphens become spaces | before splitting |
| split() | a list of words | on whitespace |
| strip(...) | punctuation off both ends | a new string |
| lower() | case normalised | another new string |
| hist.get(word, 0) + 1 | count it | lesson 11a's idiom |
Replace hyphens before splitting.
Why: process_line uses replace to replace hyphens with spaces before using split to break the line into a list of strings.
Clean each word.
Why: It traverses the list of words and uses strip and lower to remove punctuation and convert to lower case.
Count it.
Why: Finally, process_line updates the histogram by creating a new item or incrementing an existing one — the get idiom from lesson 11a.
Figure (svg): The state of the program after each line of Worked example the cleaning, in three lines, drawn as a ladder with one rung per traced line
Six lines that turn a raw line of a book into counted words. Every line of the file passes through this.
Verify: Check the reassignments.
Why: Each of the three cleaning calls assigns its result back to word, which it must: it is a shorthand to say that strings are converted, since strings are immutable and methods like strip and lower return new strings. Omitting any one assignment leaves that step undone, silently.
Prediction
process_line modifies the dictionary it was passed.
def add_word(word, hist):
hist[word] = hist.get(word, 0) + 1
h = {}
add_word('cat', h)
print(h)| Step | What happens | Result |
|---|---|---|
| the call | passes a reference | not a copy |
| hist[word] = ... | modifies the object | in place |
| print(h) | the same object | {'cat': 1} |
Predict first
What does this print?
Correct: {'cat': 1} — the function modified the dictionary the caller passed, because both names refer to one object.
Why: A function receives a reference to a mutable argument, so modifying it is visible to the caller — which is exactly what makes the accumulator pattern work. Note that the function assigns to an item rather than to the name hist; hist = {} inside the function would repoint the parameter and leave the caller's dictionary empty.
Worked example
Both are one line, and both come free from the histogram.
def total_words(hist):
return sum(hist.values())
def different_words(hist):
return len(hist)
# Total number of words: 161080
# Number of different words: 7214| Measure | What it counts | Emma |
|---|---|---|
| sum of values | every occurrence | 161,080 |
| number of items | every distinct word | 7,214 |
| the ratio | about 22 uses per word | on average |
Add the frequencies for the length.
Why: Every occurrence contributed 1 to some counter, so summing the counters recovers the total.
Count the items for the vocabulary.
Why: Each distinct word created exactly one item, so the number of items is the number of distinct words.
Look at the numbers.
Why: Emma is 161,080 words long and uses 7,214 different ones — so the average word is used about twenty-two times.
Figure (svg): A pipeline from a file through per-line processing to a histogram and its two totals
Two one-line functions and two numbers that mean quite different things. The histogram answered both without a second pass over the text.
Verify: Sanity-check the two against each other.
Why: The total must be at least the vocabulary, since every distinct word occurs at least once — and 161,080 against 7,214 is comfortably consistent. If they were close to equal, almost every word would be used once, which would suggest the cleaning had failed and near-duplicates were being counted separately.
Trap
A student moves hist = dict() into process_line, so each function owns its own data.
Keep a function's data inside it
Why: Which is normally good practice and avoids passing things around.
Every line then starts with an empty histogram and the counts are thrown away at the end of each call. The program runs, returns an empty dictionary, and reports a vocabulary of zero.
Create the accumulator once and pass it in.
hist = dict() in process_file, before the loop
Why: So that one dictionary survives across every line.
process_line modifies what it was given
Why: Which works because a function receives a reference to a mutable object.
This is the same lifetime question as chapter 11's memo: the accumulator has to outlive the call that adds to it, and here that is achieved by passing rather than by a global.
Ranking
Five steps, in order.
Put in order
Why: The hyphen replacement must precede the split, because it is what causes the split to happen at those points. Everything after operates on one word at a time: strip, lower, then count. Stripping and lowering could be swapped without changing the result, since neither affects what the other removes.
Faded example
The idiom from lesson 11a, in its natural home.
Fill in the blanks
word = word.lower()
hist[word] = hist.get(word, 0) + 1
Why: The default of 0 is the count of a word not yet seen, so adding one gives 1 on a first sighting and the correct increment thereafter. This is what removes the need for a conditional distinguishing the first occurrence from the rest — and it is why the whole counting step fits on one line.
Explain it to yourself
process_line could have returned an updated dictionary.
Discussion prompt
Why does process_line take the histogram as an argument and modify it, rather than returning a new one for process_file to merge?
Hint: How many lines are there in a book?
Answer:
Emma has tens of thousands of lines. Returning a new dictionary per line would mean merging tens of thousands of dictionaries, which is far more work than adding to one.
Modifying in place costs nothing extra: the function already has a reference to the caller's dictionary, so each word goes straight into the structure that will be the answer.
The cost is that process_line has a side effect the caller must know about — which is exactly why it returns None and why the parameter name says what it is. Chapter 10's advice was to pick one contract and document it, and this function has picked the modifying one for a good reason.
Section
Section 2
Concept
To find the most common words, we can make a list of tuples, where each tuple contains a word and its frequency, and sort it. The following function takes a histogram and returns a list of word-frequency tuples.
def most_common(hist):
t = []
for key, value in hist.items():
t.append((value, key))
t.sort(reverse=True)
return t| Line | What it does | Note |
|---|---|---|
| hist.items() | (word, frequency) | the dictionary's own order |
| append((value, key)) | swapped to (frequency, word) | sort field first |
| sort(reverse=True) | descending | commonest first |
| return t | a list of tuples | ordered |
In each tuple, the frequency appears first, so the resulting list is sorted by frequency. Nothing tells sort what to compare — the tuple comparison rule from lesson 12a does it, using position 0.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-128
Picture it
items gives one order and the function wants the other.
Figure (svg): A pipeline showing dictionary items swapped into frequency-first tuples and then sorted
And the tie-break comes free: two words with the same frequency are ordered by the word itself, because that is position 1.
Worked example
A slice, a loop, and one keyword argument.
t = most_common(hist)
print('The most common words are:')
for freq, word in t[:10]:
print(word, freq, sep='\t')
# to 5242
# the 5205
# and 4897| Part | What it does | Note |
|---|---|---|
| t[:10] | the first ten tuples | already sorted |
| for freq, word in ... | tuple assignment | unpacked per pass |
| sep='\t' | a tab between the columns | so they line up |
Slice off the first ten.
Why: The list is already sorted in descending order, so the commonest words are at the front.
Unpack each tuple.
Why: for freq, word in ... binds the two fields in the order they appear in the tuple — frequency first, because that is how they were built.
Use the sep keyword argument.
Why: I use the keyword argument sep to tell print to use a tab character as a separator, rather than a space, so the second column is lined up.
Figure (svg): The state of the program after each line of Worked example printing the top ten, drawn as a ladder with one rung per traced line
The ten commonest words in Emma, in two aligned columns. The unpacking order must match the tuple order, or the columns swap.
Verify: Check the results against expectation.
Why: The top ten are all function words — to, the, and, of, i, a, it, her, was, she — which is what any English text produces. If a content word appeared in the top ten, that would suggest the header was not skipped or the cleaning had failed, so the plausibility of the list is itself a check.
Prediction
The tuples were built frequency-first.
t = [(3, 'cat'), (10, 'the'), (1, 'zoo')]
t.sort(reverse=True)
print(t[0])| Part | What happens | Result |
|---|---|---|
| comparison | by position 0 | the frequency |
| reverse=True | descending | largest first |
| t[0] | the largest frequency | (10, 'the') |
Predict first
What does this print?
Correct: (10, 'the') — sorting descending by position 0 puts the highest frequency first.
Why: Tuples compare element by element from the left, so position 0 — the frequency — decides the order, and reverse=True makes it descending. Had the tuples been built word-first, position 0 would be the word and the answer would be ('zoo', 1), sorted alphabetically in reverse.
Worked example
Build the tuples the other way and the sort does something else.
# frequency first: sorts by frequency
t.append((value, key))
# [(5242, 'to'), (5205, 'the'), ...]
# word first: sorts alphabetically
t.append((key, value))
# [('a', 3130), ('and', 4897), ...]| Tuple order | What position 0 holds | Sorted by |
|---|---|---|
| (frequency, word) | position 0 is the count | sorts by count |
| (word, frequency) | position 0 is the word | sorts alphabetically |
| the sort call | identical in both | the tuples differ |
Recall the comparison rule.
Why: Python compares the first element from each sequence and moves on only if they are equal — so position 0 dominates completely.
See what that means here.
Why: Whichever field is at position 0 becomes the primary ordering, and the other becomes the tie-break.
Note that sort takes no hint.
Why: The two versions call sort identically. The difference lives entirely in how the tuples were built.
Figure (svg): Two lists of the same data built with fields in opposite orders and their resulting sorts
Frequency-first sorts by frequency and word-first sorts alphabetically, from the same data and the same sort call. The ordering is decided at construction.
Verify: Check the tie-break behaviour.
Why: With (frequency, word), two words of equal frequency are ordered by the word — reversed, since sort was given reverse=True, so z comes before a among ties. That is a small oddity worth noticing: reverse applies to the whole comparison, not just the first field.
Trap
The printing loop is written as for word, freq in t[:10], reading the names in the order they appear on the page.
Name them in the order you think of them
Why: Word then frequency is how you would say it aloud.
The tuples hold (frequency, word), so word binds to a number and freq to a string. The output prints the count first with no error, and the columns are silently swapped.
Match the order the tuples were built in.
for freq, word in t[:10]
Why: Because most_common appended (value, key).
Check one line of output
Why: A word where a number should be is immediately visible.
Tuple assignment binds strictly by position and cannot know what you meant — lesson 12a's silent failure, showing up here in the one place the tuples are deliberately built in an unusual order.
Faded example
Which field must come first?
Fill in the blanks
for key, value in hist.items():
t.append((value, key))
t.sort(reverse=True)
Why: The value is the frequency, and putting it at position 0 is what makes the sort order by frequency — tuples compare from the left, and nothing else tells sort what to look at. Appending (key, value) instead would sort the words alphabetically, using the same sort call on the same data.
Discrimination
Read position 0 of each tuple.
Sort into buckets
For each list of tuples, what does sorting it order by?
Explain it
Nothing in the sort call mentions it.
Discussion prompt
A classmate cannot see how t.sort() knows to order by frequency when the call takes no arguments about it. Explain.
Hint: What is being sorted?
Answer:
It does not know anything about frequencies. It is sorting tuples, and tuples compare element by element from the left — so position 0 decides the order whatever happens to be there.
The decision was made two lines earlier, when the tuples were built as (value, key) rather than (key, value). Putting the frequency at position 0 is what makes it the sort field.
Which means you can sort by any field you like without a single option: put it at position 0. The book mentions that sort also has a key parameter for doing this without rearranging the data — but the tuple trick works with what you already know.
Section
Section 3
Concept
We have seen built-in functions and methods that take optional arguments. It is possible to write programmer-defined functions with optional arguments too.
def print_most_common(hist, num=10):
t = most_common(hist)
print('The most common words are:')
for freq, word in t[:num]:
print(word, freq, sep='\t')
print_most_common(hist) # num gets 10
print_most_common(hist, 20) # num gets 20| Parameter | Required? | Note |
|---|---|---|
| hist | required | no default |
| num=10 | optional | the default value |
| one argument | num is 10 | the default applies |
| two arguments | num is 20 | the argument overrides |
The first parameter is required; the second is optional, with a default value of 10. If you only provide one argument, num gets the default value; if you provide two, the optional argument overrides the default.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 129-129
Picture it
The default is used only when no argument arrives.
Figure (svg): A flowchart showing a parameter taking either the supplied argument or its default
The body cannot tell which happened, and does not need to — by the time it runs, num holds a number either way.
Worked example
The ordering rule, and what happens when you break it.
# legal: required first, then optional
def print_most_common(hist, num=10):
...
# illegal
def bad(num=10, hist):
...
# SyntaxError: non-default argument follows default argument| Definition | Legal? | Note |
|---|---|---|
| required then optional | legal | the required one is unambiguous |
| optional then required | SyntaxError | caught at definition |
| the reason | which argument is which? | position would be ambiguous |
State the rule.
Why: If a function has both required and optional parameters, all the required parameters have to come first, followed by the optional ones.
See why it must be so.
Why: Arguments are matched by position, so if the optional one came first, a single argument would be ambiguous — is it the optional one or the required one?
Note when the error appears.
Why: At definition time, not at the call — which is unusually early and unusually helpful.
Figure (svg): The state of the program after each line of Worked example required before optional, drawn as a ladder with one rung per traced line
Required parameters first. The rule follows from positional matching, and Python enforces it when the function is defined rather than when it is called.
Verify: Check what a single argument would have to mean.
Why: In bad(num=10, hist), calling bad(5) could plausibly mean num=5 with hist missing, or hist=5 with num defaulted. There is no rule that would settle it, which is why the definition is rejected rather than the call.
Prediction
Only one argument is supplied.
def show(hist, num=10):
return num
print(show({}))| Part | What happens | Result |
|---|---|---|
| one argument | matched to hist | positionally |
| num | no argument given | the default applies |
| the result | 10 | the default value |
Predict first
What does this print?
Correct: 10 — no second argument was given, so num gets its default value.
Why: The first parameter is required and receives the dictionary; the second is optional and falls back to its default. Providing a second argument would override it — the optional argument overrides the default, in the book's words. A TypeError would only occur if a required parameter were missing.
Worked example
The function is more useful for having one.
print_most_common(hist) # the usual case: ten
print_most_common(hist, 20) # the exercise asked for twenty
print_most_common(hist, 1) # just the commonest
print_most_common(hist, 100) # a longer look| Call | What it does | Note |
|---|---|---|
| no second argument | the common case | stays short |
| a second argument | the unusual case | still available |
| without a default | every call must say 10 | noise at every call site |
Identify the usual case.
Why: Ten is what you want almost every time, so making it the default removes an argument from almost every call.
Keep the general case available.
Why: Exercise 13.3 asks for twenty, and the same function serves both without modification.
Note what a default is not for.
Why: It should be the value that is right most of the time, not merely a value that avoids an error.
Figure (svg): Two columns comparing a function with a default parameter against one without
One function covering every case, with the common one requiring no argument. That is what an optional parameter buys.
Verify: Ask what happens with a num larger than the vocabulary.
Why: t[:100000] gives the whole list rather than raising, because slicing beyond the end is legal — lesson 10a's rule. So the function degrades gracefully on a bad argument, which is worth knowing before someone adds a length check that was never needed.
Trap
A function is written as def collect(item, results=[]) so the caller need not supply a list.
Give the optional parameter an empty list
Why: It is the obvious default for something that accumulates.
The default is created once, when the function is defined, and shared by every call that omits the argument. Results accumulate across calls that were meant to be independent, which looks like data appearing from nowhere.
Default to None and create the list inside.
def collect(item, results=None)
Why: None is immutable and cannot accumulate anything.
Then: if results is None: results = []
Why: A fresh list per call, which is what was meant.
The rule is that a default value should be immutable. num=10 is safe for exactly this reason, and it is why the book's example never runs into the problem.
Error analysis
Mark each and say whether it is legal.
Annotate
Line 3 is the one to watch. It produces no error and results from earlier calls appear in later ones.
Faded example
Ten unless the caller says otherwise.
Fill in the blanks
def print_most_common(hist, num=10):
t = most_common(hist)
for freq, word in t[:num]:
print(word, freq, sep='\t')
Why: An equals sign and a value in the parameter list make the parameter optional with that default. Without it, num would be required and every call site would have to supply a number — including the many that just want the usual ten.
Socratic
The rule looks arbitrary until you try to break it.
Discussion prompt
Suppose def f(num=10, hist) were allowed. What would the call f(5) mean?
Hint: There are two readings and no rule to choose between them.
Answer:
It could mean num=5 with hist not supplied — which would be an error, since hist is required. Or it could mean hist=5 with num defaulted, matching by skipping the optional one.
Nothing decides between those readings. Arguments are matched by position, and position cannot express skip this one.
So the language rules out the definition rather than trying to resolve the call, and it does so at definition time — before the ambiguous call has even been written. That is the earliest and cheapest place to catch it.
Section
Section 4
Concept
Finding the words from the book that are not in the word list is a problem you might recognise as set subtraction: we want to find all the words from one set — the words in the book — that are not in the other.
def subtract(d1, d2):
res = dict()
for key in d1:
if key not in d2:
res[key] = None
return res| Line | What it does | Note |
|---|---|---|
| for key in d1 | every key in the first | the book's words |
| if key not in d2 | a fast membership test | the word list |
| res[key] = None | keep the key | the value is unused |
subtract takes dictionaries d1 and d2 and returns a new dictionary containing all the keys from d1 that are not in d2. Since we don't really care about the values, we set them all to None.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 129-129
Picture it
Everything in the book, minus everything in the word list.
Figure (svg): Two columns showing which words survive the subtraction and which are removed
Which is the exercise's real question: how many are typos, how many are common words missing from the list, and how many are genuinely obscure?
Worked example
The dictionary is being used for its keys alone.
res[key] = None
# the result is used only for its keys:
for word in diff:
print(word, end=' ')| Part | What it is for | Note |
|---|---|---|
| the value | None | never read |
| the key | the word | all that matters |
| why a dictionary at all | fast membership, no duplicates | not for the values |
Note what the structure is for.
Why: Since we don't really care about the values, we set them all to None — the dictionary is holding a collection of keys.
Note why that is reasonable.
Why: A dictionary gives fast membership testing and automatic removal of duplicates, both of which are wanted here.
Note how the result is used.
Why: Looping over the result yields the keys, and nothing ever looks at a value.
Figure (svg): A dictionary whose values are all None, used purely as a collection of keys
A dictionary used as a set of words. The values exist because a dictionary requires them, not because the program wants them.
Verify: Ask what Python offers for this directly.
Why: A set, which is exactly a collection of unique hashable things with fast membership and no values at all. The book has not introduced it, so a dictionary with None values is the standing substitute — and recognising that this is what the code is doing makes the eventual introduction of sets read as a simplification rather than a new idea.
Prediction
Keys in the first and not the second.
d1 = {'a': 1, 'b': 2, 'c': 3}
d2 = {'b': 99}
print(sorted(subtract(d1, d2)))| Key | Test | Result |
|---|---|---|
| 'a' | not in d2 | kept |
| 'b' | in d2 | dropped |
| 'c' | not in d2 | kept |
Predict first
What does this print?
Correct: ['a', 'c'] — the keys of d1 that do not appear as keys of d2.
Why: The function keeps a key when it is absent from the second dictionary, so 'b' is dropped and the other two survive. Note that d2's value of 99 is never consulted: only membership matters, which is why the word list's counts are irrelevant and why its values in the result are set to None.
Worked example
The membership test runs once per word in the book.
words = process_file('words.txt') # a dictionary
diff = subtract(hist, words)
# the test inside: key not in d2
# 7,214 tests, each about constant time
# if words were a LIST of 100,000 entries:
# 7,214 tests x up to 100,000 comparisons each| Structure | Cost per test | Note |
|---|---|---|
| words as a dictionary | each test about constant | fast |
| words as a list | each test scans | up to 100,000 comparisons |
| the total | hundreds of millions | against a few thousand |
Count the tests.
Why: One per distinct word in the book — 7,214 for Emma.
Cost each test.
Why: For a dictionary, in takes about the same amount of time no matter how many items there are; for a list, it searches in order.
Multiply.
Why: The dictionary version does a few thousand fast lookups; the list version does up to hundreds of millions of comparisons.
Figure (svg): The state of the program after each line of Worked example why the second structure must be a dictionary, drawn as a ladder with one rung per traced line
The same algorithm, made practical by the structure of its second argument. This is precisely the data-structure selection the chapter is named for.
Verify: Check that the book builds the word list the same way.
Why: It calls process_file on words.txt, producing a dictionary rather than a list — which is a deliberate reuse and a deliberate structure choice. The counts in that histogram are meaningless, and the fast membership is the whole reason for it.
Trap
A student writes if d1[key] not in d2, comparing counts rather than words.
Use the value, since that is what the dictionary holds
Why: The key is the index, so the value feels like the content.
That asks whether a frequency appears as a word in the word list, which is nearly always false — so almost every word survives and the result is meaningless.
Test the keys.
if key not in d2
Why: The words are the keys, in both dictionaries.
Remember what in checks
Why: It checks keys, so the test is already asking the right question about d2.
This is lesson 11a's asymmetry again: a dictionary is built to be asked about its keys, and here both dictionaries are keyed by word for exactly that reason.
Faded example
The values do not matter.
Fill in the blanks
def subtract(d1, d2):
res = dict()
for key in d1:
if key not in d2:
res[key] = None
return res
Why: The function keeps the keys of d1 that are absent from d2, so the test is for absence. Using in instead would compute the intersection — the words the book and the list have in common — which is a perfectly good function and not the one the exercise asks for.
Sorting
Ask whether the structure is searched repeatedly.
Sort into buckets
For each use, which structure is right?
Real world
What is in one and not the other is a very common question.
Discussion prompt
Think of a situation where you compared two collections to find what was missing from one. What made the comparison slow or fast?
Hint: How did you look each item up?
Answer:
Checking a guest list against arrivals, reconciling two records, finding which files were not backed up — all the same shape: for everything in A, is it in B?
What decides the speed is how B is organised. Scanning an unsorted list for every item in A is the slow way, and it is what people do by hand.
Indexing B first — alphabetising, or building a dictionary — turns each check into a direct lookup. It costs one pass to build and saves a scan on every one of the checks, which is the same trade as the memo in chapter 11.
Section
Section 5
Concept
The program produces four kinds of result, and each one answers a question the exercises actually asked.
The last one is the most interesting, because the exercise asks you to sort them out: how many are typos, how many are common words that should be in the word list, and how many are really obscure?
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-129
Picture it
Four outputs, four different questions about one text.
Figure (svg): Four outputs of the program paired with the question each one answers
Notice that only the fourth needed data from outside the book. The others are all read off a structure built in a single pass.
Worked example
The result is the same for almost any English text.
# The most common words are:
# to 5242
# the 5205
# and 4897
# of 4295
# i 3191
# a 3130| Observation | What is true | Note |
|---|---|---|
| the top ten | articles, prepositions, pronouns | no content words |
| their share | about a fifth of the book | from ten words |
| what it says about Emma | almost nothing | any novel looks like this |
Look at what is there.
Why: Every one is a function word — an article, preposition, conjunction or pronoun — and not one is about the story.
Add up their counts.
Why: The top ten alone account for roughly thirty-five thousand of the book's 161,080 words, about a fifth of the text.
Draw the conclusion.
Why: The commonest words tell you the text is English, and nothing else. Distinguishing texts needs a different measure.
Figure (svg): A ladder showing the cumulative share of the text taken by the commonest words
A ranking dominated by structural words, which is true of nearly every English text. The result is correct and not very informative, which is itself worth noticing.
Verify: Ask what would be informative instead.
Why: The words unusually common in this book compared with English generally — which is exactly what subtracting a word list starts to approximate, since it removes everything ordinary. That progression, from a correct-but-dull result to a more useful one, is the shape of most data analysis.
Prediction
The commonest words in a novel.
# The most common words are:
# to 5242
# the 5205
# and 4897| Observation | What is true | Note |
|---|---|---|
| the words | articles, prepositions | structural |
| content words | absent from the top | far rarer |
| any English text | the same pattern | not specific to Emma |
Predict first
What do the ten commonest words in Emma tell you about the book?
Correct: Almost nothing — they are function words that dominate any English text.
Why: Articles, prepositions and pronouns are the commonest words in essentially all English prose, so the ranking identifies the language rather than the book. That is why the exercise moves on to subtracting a word list: removing the ordinary words is what leaves something characteristic of this text.
Worked example
The words not in the word list fall into three kinds.
words = process_file('words.txt')
diff = subtract(hist, words)
for word in diff:
print(word, end=' ')
# proper names, archaisms, and typos - mixed| Category | Example | Note |
|---|---|---|
| proper names | Woodhouse, Highbury | absent from any word list |
| archaic forms | old spellings and contractions | correct in 1815 |
| typos | genuine errors in the file | what the exercise hunts |
Expect proper names to dominate.
Why: A general word list contains no names, so every character and place in the novel appears in the difference.
Expect period spellings.
Why: A book from 1815 uses forms a modern list omits, and they are not errors.
Look for the genuine typos among them.
Why: The exercise asks how many are typos, how many are common words that should be in the word list, and how many are really obscure.
Figure (svg): The state of the program after each line of Worked example interpreting the leftovers, drawn as a ladder with one rung per traced line
A mixed list requiring human judgement to sort out. The program narrows the whole vocabulary down to a few hundred candidates, and cannot do the last step.
Verify: Ask what the result says about the word list itself.
Why: As much as it says about the book. A common word appearing in the difference means the word list is incomplete, not that the author misspelled anything — which is why the exercise asks about both directions. Any comparison against a reference is also a test of the reference.
Trap
A program reports every word absent from words.txt as a spelling error.
Trust the reference data
Why: It is a word list, so what is not in it is presumably not a word.
Most of the output is proper names and period spellings, so the error count is dominated by things that are not errors — and the genuine typos are invisible among them.
Treat the difference as candidates, not conclusions.
Expect three categories
Why: Names, archaisms and real errors, which the exercise names explicitly.
Read the result as a test of both files
Why: A common word in the difference means the list is incomplete.
The program's job is to narrow 7,214 words down to a few hundred worth looking at. Deciding which are errors is a judgement it cannot make, and claiming otherwise turns a useful filter into a wrong answer.
Comparison
Fill the blanks. Each needs a different amount of work.
Comparison matrix
| Result | What it needs | How informative |
|---|---|---|
| total words | sum the histogram's values | the book's length, and nothing more |
| different words | count the histogram's items | vocabulary size, comparable only at equal lengths |
| commonest words | build tuples and sort | identifies the language, not the book |
| words not in the word list | a second file and a subtraction | the most revealing, and needs judgement |
The pattern is that the cheap results are the least interesting, which is normal: the informative question usually needs something from outside the data.
Explain it
The program is right and the result says nothing.
Discussion prompt
A classmate reports that their analysis found the is the most common word in their book. Explain why that is not a result, and suggest what to compute instead.
Hint: What would a different book give?
Answer:
Every English text gives the same answer, so the finding is about English rather than about their book. A result that would be identical for any input is not telling you about the input.
What distinguishes a text is where it differs from the ordinary — words it uses much more than usual, or words that appear in it and almost nowhere else.
Subtracting a word list is the first step in that direction, and it is what the chapter does next. The general principle is worth having: compare against a baseline, because an absolute count usually measures the baseline rather than the subject.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C treats the word list as complete and authoritative, and it is neither. The difference mixes proper names, words like rencontre that are no longer in common use, genuine typos, and ordinary words the list happens to omit. The program narrows 7,214 words down to a few hundred candidates; deciding which are errors is judgement it cannot make.
Comparison
Fill the blanks. Each is chosen for a different question.
Comparison matrix
| Question | Dictionary | List of tuples |
|---|---|---|
| Used for | counting and membership | ranking |
| Ordered? | no | yes, once sorted |
| Cost of a membership test | about constant | proportional to the length |
| Which functions use it? | process_line, subtract | most_common, print_most_common |
The program converts between them deliberately: count in a dictionary, rank in a list of tuples, test membership back in a dictionary.
Pattern
Five steps, and the second is the one that does the work.
Steps 2 and 5 must agree. Building (count, word) and unpacking as word, count is legal, silent, and prints the two columns the wrong way round.
Python documentation — collections — Container datatypes collections — Container datatypes
Check
The list is sorted by whichever field comes first.
t = []
for key, value in hist.items():
t.append((key, value))
t.sort(reverse=True)| Part | What happens | Result |
|---|---|---|
| (key, value) | the word is at position 0 | not the count |
| comparison | by position 0 | the word |
| the result | reverse alphabetical | not by frequency |
Check your understanding
What does this list end up sorted by?
Answer: A
Why: Tuples compare from the left, so position 0 dominates — and here that is the word. To sort by frequency the tuples must be built as (value, key), which is exactly what most_common does. The sort call is identical in both cases; only the construction differs.
Check
One of these definitions is rejected.
Check your understanding
Which definition raises a SyntaxError?
Answer: A
Why: If a function has both required and optional parameters, all the required parameters have to come first. Putting the optional one first makes a single-argument call ambiguous, so Python rejects the definition rather than the call — at definition time, which is as early as the error could possibly be caught.
Check
The values are set to None.
Check your understanding
Why does subtract set every value in its result to None?
Answer: A
Why: Since we don't really care about the values, we set them all to None. The dictionary is being used as a collection of unique keys with fast membership testing — which is what a set is, a type the book has not yet introduced.
Real world
Rank by one field, break ties by another — the everyday shape of a sorted report.
Discussion prompt
Think of a ranked list you have seen — sales by product, scores by player, files by size. What was the tie-break, and would you have noticed if there had not been one?
Hint: What happens when two entries are equal?
Answer:
Most ranked lists have one: equal scores go alphabetically, equal sizes by name or date. Without it the order among ties is arbitrary and can change between runs.
That instability is genuinely annoying — a report that reorders its middle rows every time it is generated looks broken even when the ranking is correct.
Putting the fields in a tuple gives the tie-break for free: sort by the first, and ties resolve by the second. It is the same mechanism as most_common, and it is why choosing the order of fields is a decision worth making deliberately.
Commit first
Answer, then rate your confidence.
Predict first
In most_common, why does each tuple hold (frequency, word) rather than (word, frequency)?
Correct: Because tuples compare from the left, so position 0 decides what the sort orders by.
Why: Nothing in t.sort(reverse=True) mentions frequencies. The ordering comes entirely from the tuple comparison rule of lesson 12a: Python compares the first element from each sequence, and moves on only if they are equal. Putting the frequency at position 0 makes it the primary ordering and leaves the word as the tie-break; building the tuples the other way round would sort alphabetically, with the same sort call on the same data. This also means you can sort by any field at all without options — put it first. The one thing to watch is that the printing loop must unpack in the same order, since for word, freq would bind a number to word and print the columns backwards, silently.
Explain it
Three structures, chosen three times in one program.
Discussion prompt
A classmate asks why the program keeps converting between dictionaries and lists. Walk them through the three choices and what each one buys.
Hint: Each conversion happens just before an operation the other structure cannot do.
Answer:
Counting uses a dictionary, because reaching a word's counter has to be fast and there are 161,080 of those lookups.
Ranking uses a list of tuples, because a dictionary has no order and cannot be sorted — so the pairs are pulled out, swapped to put the frequency first, and sorted.
Membership testing goes back to a dictionary, because the word list is checked once for every distinct word in the book, and in on a list would scan it every time.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The accumulator depends on chapter 10's point that a function can modify what it is given, and it is worth being able to say why creating the dictionary in the wrong place breaks it. The tuple order in most_common is the chapter's neatest idea and the one most worth over-learning, since it lets you sort by any field with no options. Optional parameters are mechanically simple with one genuine trap — a mutable default. And dictionary subtraction is a set operation wearing a dictionary's clothes, which will make sets feel obvious when you meet them.
Connect it up
One page, from memory.
Draw it
Draw the program as three boxes — file, histogram, ranked list — with the function that produces each, and mark on each arrow what the conversion buys. Beside the ranked list, write one tuple as most_common builds it and circle the field that decides the sort. Underneath, write a function definition with one required and one optional parameter, and note what happens if you swap them. Finally write subtract in four lines and say in one sentence what the None values are for.
Recap
Three pages, and the book's answers to the exercises.
| If you remember one thing | It is this |
|---|---|
| From the accumulator | One dictionary, created once, modified by every call. |
| From most_common | The sort field goes at position 0. Nothing else tells sort anything. |
| From optional parameters | Required first, and the default should be immutable. |
| From subtract | A dictionary with None values is a set in disguise. |
| From the results | The is the commonest word in every English text, which is why it is not a finding. |
The next lesson turns the analysis around and generates text: random words weighted by frequency, then Markov analysis — which needs a dictionary mapping tuples of words to lists of the words that followed them, and is the chapter's real exercise in choosing a data structure.
Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-129 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.