13b Histograms, Most Common Words, and Optional Parameters

This lesson gives the book's solution to the word-frequency exercises: reading a file into a histogram, counting it two ways, finding the commonest words by sorting tuples, writing functions with optional parameters, and subtracting one dictionary from another.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 13b Histograms, Most Common Words, and Optional Parameters

Title

Python · Chapter 13 — Case study: data structure selection

§13.3-13.6, pp. 127-129

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-129 — the pages these objectives are drawn from

3. Before we start: sorting by the wrong field

Warm-up

You have a dictionary of word counts and want the ten commonest.

Discussion prompt

sorted(hist.items()) gives you the items in order — but in order of what? What would you have to change to order them by frequency instead?

Hint: Which element of each tuple comes first?

Answer:

Each item is a (word, count) tuple, and tuples compare element by element from the left — so it sorts alphabetically by word, which is not what you asked for.

To sort by frequency, the frequency has to be at position 0. Which means building the tuples the other way round: (count, word).

That is exactly what the book's most_common does, and it is why chapter 12's comparison rule was worth learning: the order of fields inside a tuple decides what a plain sort does.

4. The one idea behind this lesson: put the sort field first

Concept

To find the most common words, we can make a list of tuples, where each tuple contains a word and its frequency, and sort it. In each tuple, the frequency appears first, so the resulting list is sorted by frequency.

def most_common(hist):
    t = []
    for key, value in hist.items():
        t.append((value, key))     # frequency FIRST
    t.sort(reverse=True)
    return t
LineWhat it doesNote
hist.items()(word, frequency) pairsthe dictionary's order
(value, key)swapped: (frequency, word)the sort field first
t.sort(reverse=True)highest frequency firstone call, no options

Nothing tells sort what to order by. The ordering comes from the position of the frequency inside each tuple, which is a decision made two lines earlier.

Figure (svg): Two columns contrasting tuples built word-first with tuples built frequency-first

Same data, same sort call, and completely different results.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-128

5. Reading a file into a histogram

Section

Section 1

6. Two functions and an accumulator

Concept

Here is a program that reads a file and builds a histogram of the words in it. process_file loops through the lines of the file, passing them one at a time to process_line, and the histogram is being used as an accumulator.

import string

def process_file(filename):
    hist = dict()
    fp = open(filename)
    for line in fp:
        process_line(line, hist)
    return hist
LineWhat it doesNote
hist = dict()the accumulatorcreated once
for line in fpone line per passa file is iterable
process_line(line, hist)the SAME dictionary is passedmodified in place
return histeverything accumulatedone dictionary

The histogram is passed to process_line and modified there, which works because a dictionary is mutable and a function receives a reference to it — chapter 10's list-arguments point, in the place it is genuinely useful.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-127

7. Picture it: one dictionary, many lines

Picture it

Every line adds to the same accumulator.

Figure (svg): A call diagram showing process_file passing one histogram to many calls of process_line

One dictionary, created once and modified by every call.

Notice process_line returns nothing. It modifies the accumulator, which is a design decision the caller has to know about — and lesson 10c's rule says it should therefore return None, which it does.

8. Worked example: the cleaning, in three lines

Worked example

The steps lesson 13a described, as the book writes them.

def process_line(line, hist):
    line = line.replace('-', ' ')
    for word in line.split():
        word = word.strip(string.punctuation + string.whitespace)
        word = word.lower()
        hist[word] = hist.get(word, 0) + 1
StepWhat it doesNote
replace('-', ' ')hyphens become spacesbefore splitting
split()a list of wordson whitespace
strip(...)punctuation off both endsa new string
lower()case normalisedanother new string
hist.get(word, 0) + 1count itlesson 11a's idiom

Replace hyphens before splitting.

Why: process_line uses replace to replace hyphens with spaces before using split to break the line into a list of strings.

Clean each word.

Why: It traverses the list of words and uses strip and lower to remove punctuation and convert to lower case.

Count it.

Why: Finally, process_line updates the histogram by creating a new item or incrementing an existing one — the get idiom from lesson 11a.

Figure (svg): The state of the program after each line of Worked example the cleaning, in three lines, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Six lines that turn a raw line of a book into counted words. Every line of the file passes through this.

Verify: Check the reassignments.

Why: Each of the three cleaning calls assigns its result back to word, which it must: it is a shorthand to say that strings are converted, since strings are immutable and methods like strip and lower return new strings. Omitting any one assignment leaves that step undone, silently.

9. Predict: does the caller see the change?

Prediction

process_line modifies the dictionary it was passed.

def add_word(word, hist):
    hist[word] = hist.get(word, 0) + 1

h = {}
add_word('cat', h)
print(h)
StepWhat happensResult
the callpasses a referencenot a copy
hist[word] = ...modifies the objectin place
print(h)the same object{'cat': 1}

Predict first

What does this print?

  • {'cat': 1}
  • {}
  • None
  • A NameError

Correct: {'cat': 1} — the function modified the dictionary the caller passed, because both names refer to one object.

Why: A function receives a reference to a mutable argument, so modifying it is visible to the caller — which is exactly what makes the accumulator pattern work. Note that the function assigns to an item rather than to the name hist; hist = {} inside the function would repoint the parameter and leave the caller's dictionary empty.

10. Worked example: the two counts, and the real numbers

Worked example

Both are one line, and both come free from the histogram.

def total_words(hist):
    return sum(hist.values())

def different_words(hist):
    return len(hist)

# Total number of words: 161080
# Number of different words: 7214
MeasureWhat it countsEmma
sum of valuesevery occurrence161,080
number of itemsevery distinct word7,214
the ratioabout 22 uses per wordon average

Add the frequencies for the length.

Why: Every occurrence contributed 1 to some counter, so summing the counters recovers the total.

Count the items for the vocabulary.

Why: Each distinct word created exactly one item, so the number of items is the number of distinct words.

Look at the numbers.

Why: Emma is 161,080 words long and uses 7,214 different ones — so the average word is used about twenty-two times.

Figure (svg): A pipeline from a file through per-line processing to a histogram and its two totals

One pass over the file produces a structure that answers both questions instantly.

Two one-line functions and two numbers that mean quite different things. The histogram answered both without a second pass over the text.

Verify: Sanity-check the two against each other.

Why: The total must be at least the vocabulary, since every distinct word occurs at least once — and 161,080 against 7,214 is comfortably consistent. If they were close to equal, almost every word would be used once, which would suggest the cleaning had failed and near-duplicates were being counted separately.

11. Trap: creating the histogram inside process_line

Trap

The trap

A student moves hist = dict() into process_line, so each function owns its own data.

Keep a function's data inside it

Why: Which is normally good practice and avoids passing things around.

Every line then starts with an empty histogram and the counts are thrown away at the end of each call. The program runs, returns an empty dictionary, and reports a vocabulary of zero.

The fix

Create the accumulator once and pass it in.

hist = dict() in process_file, before the loop

Why: So that one dictionary survives across every line.

process_line modifies what it was given

Why: Which works because a function receives a reference to a mutable object.

This is the same lifetime question as chapter 11's memo: the accumulator has to outlive the call that adds to it, and here that is achieved by passing rather than by a global.

12. Rank: what process_line does to one line

Ranking

Five steps, in order.

Put in order

  1. replace hyphens with spaces
  2. split the line into a list of words
  3. strip punctuation and whitespace from a word
  4. convert the word to lowercase
  5. increment the word's counter in the histogram

Why: The hyphen replacement must precede the split, because it is what causes the split to happen at those points. Everything after operates on one word at a time: strip, lower, then count. Stripping and lowering could be swapped without changing the result, since neither affects what the other removes.

13. Complete it: count the word

Faded example

The idiom from lesson 11a, in its natural home.

Fill in the blanks

word = word.lower()
hist[word] = hist.get(word, 0) + 1

Why: The default of 0 is the count of a word not yet seen, so adding one gives 1 on a first sighting and the correct increment thereafter. This is what removes the need for a conditional distinguishing the first occurrence from the rest — and it is why the whole counting step fits on one line.

14. Explain it yourself: why pass the histogram rather than return it?

Explain it to yourself

process_line could have returned an updated dictionary.

Discussion prompt

Why does process_line take the histogram as an argument and modify it, rather than returning a new one for process_file to merge?

Hint: How many lines are there in a book?

Answer:

Emma has tens of thousands of lines. Returning a new dictionary per line would mean merging tens of thousands of dictionaries, which is far more work than adding to one.

Modifying in place costs nothing extra: the function already has a reference to the caller's dictionary, so each word goes straight into the structure that will be the answer.

The cost is that process_line has a side effect the caller must know about — which is exactly why it returns None and why the parameter name says what it is. Chapter 10's advice was to pick one contract and document it, and this function has picked the modifying one for a good reason.

15. The most common words

Section

Section 2

16. Build tuples, then sort them

Concept

To find the most common words, we can make a list of tuples, where each tuple contains a word and its frequency, and sort it. The following function takes a histogram and returns a list of word-frequency tuples.

def most_common(hist):
    t = []
    for key, value in hist.items():
        t.append((value, key))
    t.sort(reverse=True)
    return t
LineWhat it doesNote
hist.items()(word, frequency)the dictionary's own order
append((value, key))swapped to (frequency, word)sort field first
sort(reverse=True)descendingcommonest first
return ta list of tuplesordered

In each tuple, the frequency appears first, so the resulting list is sorted by frequency. Nothing tells sort what to compare — the tuple comparison rule from lesson 12a does it, using position 0.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-128

17. Picture it: the swap that decides the ordering

Picture it

items gives one order and the function wants the other.

Figure (svg): A pipeline showing dictionary items swapped into frequency-first tuples and then sorted

The swap is the whole trick. Everything after it is an ordinary sort and an ordinary slice.

And the tie-break comes free: two words with the same frequency are ordered by the word itself, because that is position 1.

18. Worked example: printing the top ten

Worked example

A slice, a loop, and one keyword argument.

t = most_common(hist)
print('The most common words are:')
for freq, word in t[:10]:
    print(word, freq, sep='\t')

# to    5242
# the   5205
# and   4897
PartWhat it doesNote
t[:10]the first ten tuplesalready sorted
for freq, word in ...tuple assignmentunpacked per pass
sep='\t'a tab between the columnsso they line up

Slice off the first ten.

Why: The list is already sorted in descending order, so the commonest words are at the front.

Unpack each tuple.

Why: for freq, word in ... binds the two fields in the order they appear in the tuple — frequency first, because that is how they were built.

Use the sep keyword argument.

Why: I use the keyword argument sep to tell print to use a tab character as a separator, rather than a space, so the second column is lined up.

Figure (svg): The state of the program after each line of Worked example printing the top ten, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The ten commonest words in Emma, in two aligned columns. The unpacking order must match the tuple order, or the columns swap.

Verify: Check the results against expectation.

Why: The top ten are all function words — to, the, and, of, i, a, it, her, was, she — which is what any English text produces. If a content word appeared in the top ten, that would suggest the header was not skipped or the cleaning had failed, so the plausibility of the list is itself a check.

19. Predict: what does the sort produce?

Prediction

The tuples were built frequency-first.

t = [(3, 'cat'), (10, 'the'), (1, 'zoo')]
t.sort(reverse=True)
print(t[0])
PartWhat happensResult
comparisonby position 0the frequency
reverse=Truedescendinglargest first
t[0]the largest frequency(10, 'the')

Predict first

What does this print?

  • (10, 'the')
  • (1, 'zoo')
  • (3, 'cat')
  • ('zoo', 1)

Correct: (10, 'the') — sorting descending by position 0 puts the highest frequency first.

Why: Tuples compare element by element from the left, so position 0 — the frequency — decides the order, and reverse=True makes it descending. Had the tuples been built word-first, position 0 would be the word and the answer would be ('zoo', 1), sorted alphabetically in reverse.

20. Worked example: why the frequency comes first

Worked example

Build the tuples the other way and the sort does something else.

# frequency first: sorts by frequency
t.append((value, key))
# [(5242, 'to'), (5205, 'the'), ...]

# word first: sorts alphabetically
t.append((key, value))
# [('a', 3130), ('and', 4897), ...]
Tuple orderWhat position 0 holdsSorted by
(frequency, word)position 0 is the countsorts by count
(word, frequency)position 0 is the wordsorts alphabetically
the sort callidentical in boththe tuples differ

Recall the comparison rule.

Why: Python compares the first element from each sequence and moves on only if they are equal — so position 0 dominates completely.

See what that means here.

Why: Whichever field is at position 0 becomes the primary ordering, and the other becomes the tie-break.

Note that sort takes no hint.

Why: The two versions call sort identically. The difference lives entirely in how the tuples were built.

Figure (svg): Two lists of the same data built with fields in opposite orders and their resulting sorts

The sort call is the same. The tuples are not.

Frequency-first sorts by frequency and word-first sorts alphabetically, from the same data and the same sort call. The ordering is decided at construction.

Verify: Check the tie-break behaviour.

Why: With (frequency, word), two words of equal frequency are ordered by the word — reversed, since sort was given reverse=True, so z comes before a among ties. That is a small oddity worth noticing: reverse applies to the whole comparison, not just the first field.

21. Trap: unpacking the tuples in the wrong order

Trap

The trap

The printing loop is written as for word, freq in t[:10], reading the names in the order they appear on the page.

Name them in the order you think of them

Why: Word then frequency is how you would say it aloud.

The tuples hold (frequency, word), so word binds to a number and freq to a string. The output prints the count first with no error, and the columns are silently swapped.

The fix

Match the order the tuples were built in.

for freq, word in t[:10]

Why: Because most_common appended (value, key).

Check one line of output

Why: A word where a number should be is immediately visible.

Tuple assignment binds strictly by position and cannot know what you meant — lesson 12a's silent failure, showing up here in the one place the tuples are deliberately built in an unusual order.

22. Complete it: build the tuple for sorting

Faded example

Which field must come first?

Fill in the blanks

for key, value in hist.items():
t.append((value, key))
t.sort(reverse=True)

Why: The value is the frequency, and putting it at position 0 is what makes the sort order by frequency — tuples compare from the left, and nothing else tells sort what to look at. Appending (key, value) instead would sort the words alphabetically, using the same sort call on the same data.

23. Discriminate: what does this sort by?

Discrimination

Read position 0 of each tuple.

Sort into buckets

For each list of tuples, what does sorting it order by?

by the number
[(count, word), ...]; [(score, name), ...]; [(year, title), ...]
by the text
[(word, count), ...]; [(name, score), ...]; [(title, year), ...]
num
In each, the number occupies position 0, so it dominates the comparison and the text becomes the tie-break only.
txt
In each, the text is at position 0, so the list orders alphabetically and the number never affects the result unless two texts are identical.

24. Explain it: how does sort know to use the frequency?

Explain it

Nothing in the sort call mentions it.

Discussion prompt

A classmate cannot see how t.sort() knows to order by frequency when the call takes no arguments about it. Explain.

Hint: What is being sorted?

Answer:

It does not know anything about frequencies. It is sorting tuples, and tuples compare element by element from the left — so position 0 decides the order whatever happens to be there.

The decision was made two lines earlier, when the tuples were built as (value, key) rather than (key, value). Putting the frequency at position 0 is what makes it the sort field.

Which means you can sort by any field you like without a single option: put it at position 0. The book mentions that sort also has a key parameter for doing this without rearranging the data — but the tuple trick works with what you already know.

25. Optional parameters

Section

Section 3

26. A default value in the definition

Concept

We have seen built-in functions and methods that take optional arguments. It is possible to write programmer-defined functions with optional arguments too.

def print_most_common(hist, num=10):
    t = most_common(hist)
    print('The most common words are:')
    for freq, word in t[:num]:
        print(word, freq, sep='\t')

print_most_common(hist)        # num gets 10
print_most_common(hist, 20)    # num gets 20
ParameterRequired?Note
histrequiredno default
num=10optionalthe default value
one argumentnum is 10the default applies
two argumentsnum is 20the argument overrides

The first parameter is required; the second is optional, with a default value of 10. If you only provide one argument, num gets the default value; if you provide two, the optional argument overrides the default.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 129-129

27. Picture it: where the value comes from

Picture it

The default is used only when no argument arrives.

Figure (svg): A flowchart showing a parameter taking either the supplied argument or its default

One function, two ways for a parameter to get its value.

The body cannot tell which happened, and does not need to — by the time it runs, num holds a number either way.

28. Worked example: required before optional

Worked example

The ordering rule, and what happens when you break it.

# legal: required first, then optional
def print_most_common(hist, num=10):
    ...

# illegal
def bad(num=10, hist):
    ...
# SyntaxError: non-default argument follows default argument
DefinitionLegal?Note
required then optionallegalthe required one is unambiguous
optional then requiredSyntaxErrorcaught at definition
the reasonwhich argument is which?position would be ambiguous

State the rule.

Why: If a function has both required and optional parameters, all the required parameters have to come first, followed by the optional ones.

See why it must be so.

Why: Arguments are matched by position, so if the optional one came first, a single argument would be ambiguous — is it the optional one or the required one?

Note when the error appears.

Why: At definition time, not at the call — which is unusually early and unusually helpful.

Figure (svg): The state of the program after each line of Worked example required before optional, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Required parameters first. The rule follows from positional matching, and Python enforces it when the function is defined rather than when it is called.

Verify: Check what a single argument would have to mean.

Why: In bad(num=10, hist), calling bad(5) could plausibly mean num=5 with hist missing, or hist=5 with num defaulted. There is no rule that would settle it, which is why the definition is rejected rather than the call.

29. Predict: what does num hold?

Prediction

Only one argument is supplied.

def show(hist, num=10):
    return num

print(show({}))
PartWhat happensResult
one argumentmatched to histpositionally
numno argument giventhe default applies
the result10the default value

Predict first

What does this print?

  • 10
  • None
  • {}
  • A TypeError about missing arguments

Correct: 10 — no second argument was given, so num gets its default value.

Why: The first parameter is required and receives the dictionary; the second is optional and falls back to its default. Providing a second argument would override it — the optional argument overrides the default, in the book's words. A TypeError would only occur if a required parameter were missing.

30. Worked example: why the default is a good idea here

Worked example

The function is more useful for having one.

print_most_common(hist)         # the usual case: ten
print_most_common(hist, 20)     # the exercise asked for twenty
print_most_common(hist, 1)      # just the commonest
print_most_common(hist, 100)    # a longer look
CallWhat it doesNote
no second argumentthe common casestays short
a second argumentthe unusual casestill available
without a defaultevery call must say 10noise at every call site

Identify the usual case.

Why: Ten is what you want almost every time, so making it the default removes an argument from almost every call.

Keep the general case available.

Why: Exercise 13.3 asks for twenty, and the same function serves both without modification.

Note what a default is not for.

Why: It should be the value that is right most of the time, not merely a value that avoids an error.

Figure (svg): Two columns comparing a function with a default parameter against one without

The default does not add capability; it removes repetition from the common case.

One function covering every case, with the common one requiring no argument. That is what an optional parameter buys.

Verify: Ask what happens with a num larger than the vocabulary.

Why: t[:100000] gives the whole list rather than raising, because slicing beyond the end is legal — lesson 10a's rule. So the function degrades gracefully on a bad argument, which is worth knowing before someone adds a length check that was never needed.

31. Trap: a mutable default value

Trap

The trap

A function is written as def collect(item, results=[]) so the caller need not supply a list.

Give the optional parameter an empty list

Why: It is the obvious default for something that accumulates.

The default is created once, when the function is defined, and shared by every call that omits the argument. Results accumulate across calls that were meant to be independent, which looks like data appearing from nowhere.

The fix

Default to None and create the list inside.

def collect(item, results=None)

Why: None is immutable and cannot accumulate anything.

Then: if results is None: results = []

Why: A fresh list per call, which is what was meant.

The rule is that a default value should be immutable. num=10 is safe for exactly this reason, and it is why the book's example never runs into the problem.

32. Error analysis: four function definitions

Error analysis

Mark each and say whether it is legal.

Annotate

  • Line 1 is the book's form: a required parameter followed by an optional one. Legal and idiomatic.
  • Line 2 raises SyntaxError: non-default argument follows default argument. All the required parameters have to come first, because arguments are matched by position.
  • Line 3 is legal and dangerous: the empty list is created once at definition time and shared by every call that omits the argument, so it accumulates across calls.
  • Line 4 is legal — both parameters are optional, which is allowed since there is no required one to be ambiguous about.
  • So one raises at definition time, one is a genuine trap that never raises, and two are fine.
  • The two rules are: required parameters first, and default values should be immutable.

Line 3 is the one to watch. It produces no error and results from earlier calls appear in later ones.

33. Complete it: an optional count

Faded example

Ten unless the caller says otherwise.

Fill in the blanks

def print_most_common(hist, num=10):
t = most_common(hist)
for freq, word in t[:num]:
print(word, freq, sep='\t')

Why: An equals sign and a value in the parameter list make the parameter optional with that default. Without it, num would be required and every call site would have to supply a number — including the many that just want the usual ten.

34. Think it through: why must required parameters come first?

Socratic

The rule looks arbitrary until you try to break it.

Discussion prompt

Suppose def f(num=10, hist) were allowed. What would the call f(5) mean?

Hint: There are two readings and no rule to choose between them.

Answer:

It could mean num=5 with hist not supplied — which would be an error, since hist is required. Or it could mean hist=5 with num defaulted, matching by skipping the optional one.

Nothing decides between those readings. Arguments are matched by position, and position cannot express skip this one.

So the language rules out the definition rather than trying to resolve the call, and it does so at definition time — before the ambiguous call has even been written. That is the earliest and cheapest place to catch it.

35. Dictionary subtraction

Section

Section 4

36. Which keys are in one and not the other

Concept

Finding the words from the book that are not in the word list is a problem you might recognise as set subtraction: we want to find all the words from one set — the words in the book — that are not in the other.

def subtract(d1, d2):
    res = dict()
    for key in d1:
        if key not in d2:
            res[key] = None
    return res
LineWhat it doesNote
for key in d1every key in the firstthe book's words
if key not in d2a fast membership testthe word list
res[key] = Nonekeep the keythe value is unused

subtract takes dictionaries d1 and d2 and returns a new dictionary containing all the keys from d1 that are not in d2. Since we don't really care about the values, we set them all to None.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 129-129

37. Picture it: the keys that survive

Picture it

Everything in the book, minus everything in the word list.

Figure (svg): Two columns showing which words survive the subtraction and which are removed

What survives is mostly proper names, archaic spellings and typos.

Which is the exercise's real question: how many are typos, how many are common words missing from the list, and how many are genuinely obscure?

38. Worked example: why the values are None

Worked example

The dictionary is being used for its keys alone.

res[key] = None

# the result is used only for its keys:
for word in diff:
    print(word, end=' ')
PartWhat it is forNote
the valueNonenever read
the keythe wordall that matters
why a dictionary at allfast membership, no duplicatesnot for the values

Note what the structure is for.

Why: Since we don't really care about the values, we set them all to None — the dictionary is holding a collection of keys.

Note why that is reasonable.

Why: A dictionary gives fast membership testing and automatic removal of duplicates, both of which are wanted here.

Note how the result is used.

Why: Looping over the result yields the keys, and nothing ever looks at a value.

Figure (svg): A dictionary whose values are all None, used purely as a collection of keys

A dictionary standing in for a set: keys that matter, values that do not.

A dictionary used as a set of words. The values exist because a dictionary requires them, not because the program wants them.

Verify: Ask what Python offers for this directly.

Why: A set, which is exactly a collection of unique hashable things with fast membership and no values at all. The book has not introduced it, so a dictionary with None values is the standing substitute — and recognising that this is what the code is doing makes the eventual introduction of sets read as a simplification rather than a new idea.

39. Predict: what survives the subtraction?

Prediction

Keys in the first and not the second.

d1 = {'a': 1, 'b': 2, 'c': 3}
d2 = {'b': 99}
print(sorted(subtract(d1, d2)))
KeyTestResult
'a'not in d2kept
'b'in d2dropped
'c'not in d2kept

Predict first

What does this print?

  • ['a', 'c']
  • ['b']
  • ['a', 'b', 'c']
  • [1, 3]

Correct: ['a', 'c'] — the keys of d1 that do not appear as keys of d2.

Why: The function keeps a key when it is absent from the second dictionary, so 'b' is dropped and the other two survive. Note that d2's value of 99 is never consulted: only membership matters, which is why the word list's counts are irrelevant and why its values in the result are set to None.

40. Worked example: why the second structure must be a dictionary

Worked example

The membership test runs once per word in the book.

words = process_file('words.txt')   # a dictionary
diff = subtract(hist, words)

# the test inside: key not in d2
# 7,214 tests, each about constant time

# if words were a LIST of 100,000 entries:
# 7,214 tests x up to 100,000 comparisons each
StructureCost per testNote
words as a dictionaryeach test about constantfast
words as a listeach test scansup to 100,000 comparisons
the totalhundreds of millionsagainst a few thousand

Count the tests.

Why: One per distinct word in the book — 7,214 for Emma.

Cost each test.

Why: For a dictionary, in takes about the same amount of time no matter how many items there are; for a list, it searches in order.

Multiply.

Why: The dictionary version does a few thousand fast lookups; the list version does up to hundreds of millions of comparisons.

Figure (svg): The state of the program after each line of Worked example why the second structure must be a dictionary, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The same algorithm, made practical by the structure of its second argument. This is precisely the data-structure selection the chapter is named for.

Verify: Check that the book builds the word list the same way.

Why: It calls process_file on words.txt, producing a dictionary rather than a list — which is a deliberate reuse and a deliberate structure choice. The counts in that histogram are meaningless, and the fast membership is the whole reason for it.

41. Trap: subtracting values instead of keys

Trap

The trap

A student writes if d1[key] not in d2, comparing counts rather than words.

Use the value, since that is what the dictionary holds

Why: The key is the index, so the value feels like the content.

That asks whether a frequency appears as a word in the word list, which is nearly always false — so almost every word survives and the result is meaningless.

The fix

Test the keys.

if key not in d2

Why: The words are the keys, in both dictionaries.

Remember what in checks

Why: It checks keys, so the test is already asking the right question about d2.

This is lesson 11a's asymmetry again: a dictionary is built to be asked about its keys, and here both dictionaries are keyed by word for exactly that reason.

42. Complete it: keep the keys that are missing

Faded example

The values do not matter.

Fill in the blanks

def subtract(d1, d2):
res = dict()
for key in d1:
if key not in d2:
res[key] = None
return res

Why: The function keeps the keys of d1 that are absent from d2, so the test is for absence. Using in instead would compute the intersection — the words the book and the list have in common — which is a perfectly good function and not the one the exercise asks for.

43. Sort: dictionary or list for this job?

Sorting

Ask whether the structure is searched repeatedly.

Sort into buckets

For each use, which structure is right?

a dictionary
a word list checked once per word in a book; counting how often each word appears; checking whether a word has been seen before
a list
the words of a book, in the order written; the ten commonest words, in order; the lines of a file, to be processed in turn
dict
Each needs fast lookup by a key: repeated membership tests, or a counter reached by word. A list would turn every one of these into a scan.
list
Each needs order — the sequence words were written in, a ranking, or the order of lines in a file. A dictionary has no order to offer.

44. Where set subtraction is the question

Real world

What is in one and not the other is a very common question.

Discussion prompt

Think of a situation where you compared two collections to find what was missing from one. What made the comparison slow or fast?

Hint: How did you look each item up?

Answer:

Checking a guest list against arrivals, reconciling two records, finding which files were not backed up — all the same shape: for everything in A, is it in B?

What decides the speed is how B is organised. Scanning an unsorted list for every item in A is the slow way, and it is what people do by hand.

Indexing B first — alphabetising, or building a dictionary — turns each check into a direct lookup. It costs one pass to build and saves a scan on every one of the checks, which is the same trade as the memo in chapter 11.

45. Reading the results

Section

Section 5

46. The numbers are the output, and they need interpreting

Concept

The program produces four kinds of result, and each one answers a question the exercises actually asked.

The last one is the most interesting, because the exercise asks you to sort them out: how many are typos, how many are common words that should be in the word list, and how many are really obscure?

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 128-129

47. Picture it: what each number is measuring

Picture it

Four outputs, four different questions about one text.

Figure (svg): Four outputs of the program paired with the question each one answers

The first two come free from the histogram; the third needs a sort and the fourth a second file.

Notice that only the fourth needed data from outside the book. The others are all read off a structure built in a single pass.

48. Worked example: why the top ten are all function words

Worked example

The result is the same for almost any English text.

# The most common words are:
# to    5242
# the   5205
# and   4897
# of    4295
# i     3191
# a     3130
ObservationWhat is trueNote
the top tenarticles, prepositions, pronounsno content words
their shareabout a fifth of the bookfrom ten words
what it says about Emmaalmost nothingany novel looks like this

Look at what is there.

Why: Every one is a function word — an article, preposition, conjunction or pronoun — and not one is about the story.

Add up their counts.

Why: The top ten alone account for roughly thirty-five thousand of the book's 161,080 words, about a fifth of the text.

Draw the conclusion.

Why: The commonest words tell you the text is English, and nothing else. Distinguishing texts needs a different measure.

Figure (svg): A ladder showing the cumulative share of the text taken by the commonest words

A handful of words carry a fifth of the text, and thousands of words appear once each.

A ranking dominated by structural words, which is true of nearly every English text. The result is correct and not very informative, which is itself worth noticing.

Verify: Ask what would be informative instead.

Why: The words unusually common in this book compared with English generally — which is exactly what subtracting a word list starts to approximate, since it removes everything ordinary. That progression, from a correct-but-dull result to a more useful one, is the shape of most data analysis.

49. Predict: what dominates the top ten?

Prediction

The commonest words in a novel.

# The most common words are:
# to    5242
# the   5205
# and   4897
ObservationWhat is trueNote
the wordsarticles, prepositionsstructural
content wordsabsent from the topfar rarer
any English textthe same patternnot specific to Emma

Predict first

What do the ten commonest words in Emma tell you about the book?

  • Almost nothing — they are function words that dominate any English text
  • That it is about relationships, since 'her' and 'she' appear
  • That the cleaning failed, since no content words appear
  • That the vocabulary is unusually small

Correct: Almost nothing — they are function words that dominate any English text.

Why: Articles, prepositions and pronouns are the commonest words in essentially all English prose, so the ranking identifies the language rather than the book. That is why the exercise moves on to subtracting a word list: removing the ordinary words is what leaves something characteristic of this text.

50. Worked example: interpreting the leftovers

Worked example

The words not in the word list fall into three kinds.

words = process_file('words.txt')
diff = subtract(hist, words)
for word in diff:
    print(word, end=' ')

# proper names, archaisms, and typos - mixed
CategoryExampleNote
proper namesWoodhouse, Highburyabsent from any word list
archaic formsold spellings and contractionscorrect in 1815
typosgenuine errors in the filewhat the exercise hunts

Expect proper names to dominate.

Why: A general word list contains no names, so every character and place in the novel appears in the difference.

Expect period spellings.

Why: A book from 1815 uses forms a modern list omits, and they are not errors.

Look for the genuine typos among them.

Why: The exercise asks how many are typos, how many are common words that should be in the word list, and how many are really obscure.

Figure (svg): The state of the program after each line of Worked example interpreting the leftovers, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A mixed list requiring human judgement to sort out. The program narrows the whole vocabulary down to a few hundred candidates, and cannot do the last step.

Verify: Ask what the result says about the word list itself.

Why: As much as it says about the book. A common word appearing in the difference means the word list is incomplete, not that the author misspelled anything — which is why the exercise asks about both directions. Any comparison against a reference is also a test of the reference.

51. Trap: treating the word list as authoritative

Trap

The trap

A program reports every word absent from words.txt as a spelling error.

Trust the reference data

Why: It is a word list, so what is not in it is presumably not a word.

Most of the output is proper names and period spellings, so the error count is dominated by things that are not errors — and the genuine typos are invisible among them.

The fix

Treat the difference as candidates, not conclusions.

Expect three categories

Why: Names, archaisms and real errors, which the exercise names explicitly.

Read the result as a test of both files

Why: A common word in the difference means the list is incomplete.

The program's job is to narrow 7,214 words down to a few hundred worth looking at. Deciding which are errors is a judgement it cannot make, and claiming otherwise turns a useful filter into a wrong answer.

52. Compare: the four results

Comparison

Fill the blanks. Each needs a different amount of work.

Comparison matrix

ResultWhat it needsHow informative
total wordssum the histogram's valuesthe book's length, and nothing more
different wordscount the histogram's itemsvocabulary size, comparable only at equal lengths
commonest wordsbuild tuples and sortidentifies the language, not the book
words not in the word lista second file and a subtractionthe most revealing, and needs judgement

The pattern is that the cheap results are the least interesting, which is normal: the informative question usually needs something from outside the data.

53. Explain it: why is *the* being commonest not a finding?

Explain it

The program is right and the result says nothing.

Discussion prompt

A classmate reports that their analysis found the is the most common word in their book. Explain why that is not a result, and suggest what to compute instead.

Hint: What would a different book give?

Answer:

Every English text gives the same answer, so the finding is about English rather than about their book. A result that would be identical for any input is not telling you about the input.

What distinguishes a text is where it differs from the ordinary — words it uses much more than usual, or words that appear in it and almost nowhere else.

Subtracting a word list is the first step in that direction, and it is what the chapter does next. The general principle is worth having: compare against a baseline, because an absolute count usually measures the baseline rather than the subject.

54. Two truths and a lie: reading the output

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. Some of the words missing from the word list are names and possessives
  • B. A few of them are common words that should really be in the list
  • C. Every word missing from the word list is a spelling error in the book

Survives elimination: C

Why: C treats the word list as complete and authoritative, and it is neither. The difference mixes proper names, words like rencontre that are no longer in common use, genuine typos, and ordinary words the list happens to omit. The program narrows 7,214 words down to a few hundred candidates; deciding which are errors is judgement it cannot make.

55. Compare: the structures this chapter uses

Comparison

Fill the blanks. Each is chosen for a different question.

Comparison matrix

QuestionDictionaryList of tuples
Used forcounting and membershipranking
Ordered?noyes, once sorted
Cost of a membership testabout constantproportional to the length
Which functions use it?process_line, subtractmost_common, print_most_common

The program converts between them deliberately: count in a dictionary, rank in a list of tuples, test membership back in a dictionary.

56. The procedure: ranking the contents of a histogram

Pattern

Five steps, and the second is the one that does the work.

  1. Start from a histogram: the thing counted as the key, the count as the value.
  2. Build a list of tuples with the count FIRST, since position 0 decides the ordering.
  3. Sort with reverse=True for descending, which needs no key function or option.
  4. Slice off as many as you want; slicing past the end is legal and gives everything.
  5. Unpack in the same order the tuples were built — count first — or the columns silently swap.

Steps 2 and 5 must agree. Building (count, word) and unpacking as word, count is legal, silent, and prints the two columns the wrong way round.

Python documentation — collections — Container datatypes collections — Container datatypes

57. Check yourself 1 of 3: the tuple order

Check

The list is sorted by whichever field comes first.

t = []
for key, value in hist.items():
    t.append((key, value))
t.sort(reverse=True)
PartWhat happensResult
(key, value)the word is at position 0not the count
comparisonby position 0the word
the resultreverse alphabeticalnot by frequency

Check your understanding

What does this list end up sorted by?

  • A. The words, in reverse alphabetical order (correct)
  • B. The frequencies, highest first
  • C. The frequencies, lowest first
  • D. Nothing — the order is unpredictable

Answer: A

Why: Tuples compare from the left, so position 0 dominates — and here that is the word. To sort by frequency the tuples must be built as (value, key), which is exactly what most_common does. The sort call is identical in both cases; only the construction differs.

Why B tempts people
This is what (value, key) would give. The fields are the other way round here.
Why C tempts people
That would need reverse=False as well as the fields swapped.
Why D tempts people
Unpredictable order belongs to dictionaries. A sorted list has a completely defined order.

58. Check yourself 2 of 3: optional parameters

Check

One of these definitions is rejected.

Check your understanding

Which definition raises a SyntaxError?

  • A. def f(num=10, hist): (correct)
  • B. def f(hist, num=10):
  • C. def f(hist=None, num=10):
  • D. def f(hist, num):

Answer: A

Why: If a function has both required and optional parameters, all the required parameters have to come first. Putting the optional one first makes a single-argument call ambiguous, so Python rejects the definition rather than the call — at definition time, which is as early as the error could possibly be caught.

Why B tempts people
This is the book's own form: required first, then optional.
Why C tempts people
Both parameters are optional, which is fine — there is no required one to be ambiguous about.
Why D tempts people
Both are required, which is the ordinary case and needs no rule at all.

59. Check yourself 3 of 3: dictionary subtraction

Check

The values are set to None.

Check your understanding

Why does subtract set every value in its result to None?

  • A. Because the values are not needed — the result is used only for its keys (correct)
  • B. Because None is faster to store than an integer
  • C. Because a dictionary cannot hold two different value types
  • D. Because the counts from d1 would be wrong after subtraction

Answer: A

Why: Since we don't really care about the values, we set them all to None. The dictionary is being used as a collection of unique keys with fast membership testing — which is what a set is, a type the book has not yet introduced.

Why B tempts people
Storage cost is not the reason, and the difference would be irrelevant at this scale.
Why C tempts people
A dictionary can hold values of any mixture of types. Nothing constrains them.
Why D tempts people
The counts would be perfectly meaningful if kept; they are simply not wanted for this question.

60. Where this shows up outside this course

Real world

Rank by one field, break ties by another — the everyday shape of a sorted report.

Discussion prompt

Think of a ranked list you have seen — sales by product, scores by player, files by size. What was the tie-break, and would you have noticed if there had not been one?

Hint: What happens when two entries are equal?

Answer:

Most ranked lists have one: equal scores go alphabetically, equal sizes by name or date. Without it the order among ties is arbitrary and can change between runs.

That instability is genuinely annoying — a report that reorders its middle rows every time it is generated looks broken even when the ranking is correct.

Putting the fields in a tuple gives the tie-break for free: sort by the first, and ties resolve by the second. It is the same mechanism as most_common, and it is why choosing the order of fields is a decision worth making deliberately.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

In most_common, why does each tuple hold (frequency, word) rather than (word, frequency)?

  • Because tuples compare from the left, so position 0 decides what the sort orders by
  • Because sort requires numbers before strings
  • Because dictionaries store values before keys
  • Because reverse=True only works on numbers

Correct: Because tuples compare from the left, so position 0 decides what the sort orders by.

Why: Nothing in t.sort(reverse=True) mentions frequencies. The ordering comes entirely from the tuple comparison rule of lesson 12a: Python compares the first element from each sequence, and moves on only if they are equal. Putting the frequency at position 0 makes it the primary ordering and leaves the word as the tie-break; building the tuples the other way round would sort alphabetically, with the same sort call on the same data. This also means you can sort by any field at all without options — put it first. The one thing to watch is that the printing loop must unpack in the same order, since for word, freq would bind a number to word and print the columns backwards, silently.

62. Explain it to someone else

Explain it

Three structures, chosen three times in one program.

Discussion prompt

A classmate asks why the program keeps converting between dictionaries and lists. Walk them through the three choices and what each one buys.

Hint: Each conversion happens just before an operation the other structure cannot do.

Answer:

Counting uses a dictionary, because reaching a word's counter has to be fast and there are 161,080 of those lookups.

Ranking uses a list of tuples, because a dictionary has no order and cannot be sorted — so the pairs are pulled out, swapped to put the frequency first, and sorted.

Membership testing goes back to a dictionary, because the word list is checked once for every distinct word in the book, and in on a list would scan it every time.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • The accumulator: one histogram passed between functions
  • most_common, and why the frequency comes first in each tuple
  • Optional parameters, and the required-before-optional rule
  • Dictionary subtraction, and why the values are None

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The accumulator depends on chapter 10's point that a function can modify what it is given, and it is worth being able to say why creating the dictionary in the wrong place breaks it. The tuple order in most_common is the chapter's neatest idea and the one most worth over-learning, since it lets you sort by any field with no options. Optional parameters are mechanically simple with one genuine trap — a mutable default. And dictionary subtraction is a set operation wearing a dictionary's clothes, which will make sets feel obvious when you meet them.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the program as three boxes — file, histogram, ranked list — with the function that produces each, and mark on each arrow what the conversion buys. Beside the ranked list, write one tuple as most_common builds it and circle the field that decides the sort. Underneath, write a function definition with one required and one optional parameter, and note what happens if you swap them. Finally write subtract in four lines and say in one sentence what the None values are for.

65. What you can do now

Recap

Three pages, and the book's answers to the exercises.

If you remember one thingIt is this
From the accumulatorOne dictionary, created once, modified by every call.
From most_commonThe sort field goes at position 0. Nothing else tells sort anything.
From optional parametersRequired first, and the default should be immutable.
From subtractA dictionary with None values is a set in disguise.
From the resultsThe is the commonest word in every English text, which is why it is not a finding.

The next lesson turns the analysis around and generates text: random words weighted by frequency, then Markov analysis — which needs a dictionary mapping tuples of words to lists of the words that followed them, and is the chapter's real exercise in choosing a data structure.

Think Python, 2nd edition — Allen B. Downey §13.3-13.6, pp. 127-129 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §13.3-13.6, pp. 127-129
  2. Python documentation — collections — Container datatypes
  3. Python documentation — Data Structures

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108