This lesson opens the case study by setting out the word-frequency problem and the string tools for cleaning text, then introduces determinism, pseudorandom numbers, and the random module's three core functions.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 13 — Case study: data structure selection
§13.1-13.2, pp. 125-126
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126 — the pages these objectives are drawn from
Warm-up
A question that sounds simple and is not.
Discussion prompt
You are asked how many words are in a book. Before writing any code, list three decisions you would have to make before the number is even well defined.
Hint: Is The the same word as the?
Answer:
Case: are The and the the same word? Punctuation: is dog, the word dog? Hyphens: is well-known one word or two?
And a fourth: do you mean how many words were written, or how many different words were used? Those are completely different numbers and both are called how many words.
This chapter is a case study in exactly these decisions. The programming is straightforward; choosing what to count and which structure to count it in is the actual work.
Concept
At this point you have learned about Python's core data structures, and you have seen some of the algorithms that use them. This chapter presents a case study with exercises that let you think about choosing data structures and practice using them.
Every question in this chapter can be answered with a list, a dictionary, or a list of tuples — and the choice decides whether the program takes a second or an hour. That is what the chapter is about, and it is why it comes after all three types rather than alongside any one of them.
Figure (svg): Two columns pairing a question about a text with the structure that answers it well
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-125
Section
Section 1
Concept
The first exercise sets the whole problem: write a program that reads a file, breaks each line into words, strips whitespace and punctuation from the words, and converts them to lowercase.
Each step exists because of a way two occurrences of the same word can fail to look alike. Skip any one of them and the histogram counts dog and dog, and Dog as three different words.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-125
Picture it
Each stage removes one kind of difference.
Figure (svg): A pipeline showing a raw line reduced to a clean lowercase word
Note that each stage returns a new string rather than changing one — it is a shorthand to say strings are converted, since strings are immutable.
Worked example
You do not have to write out the punctuation yourself.
>>> import string
>>> string.punctuation
'!"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~'
>>> string.whitespace
' \t\n\r\x0b\x0c'| Constant | What it holds | Note |
|---|---|---|
| string.punctuation | every punctuation character | 32 of them |
| string.whitespace | space, tab, newline and more | the invisible ones |
| together | everything to strip | concatenate them |
Import the module.
Why: The string module provides a string named whitespace, which contains space, tab, newline and so on, and punctuation which contains the punctuation characters.
Look at what they hold.
Why: Both are ordinary strings, so they can be concatenated, indexed, and passed to any method that takes one.
Use them together.
Why: strip takes a string of characters to remove, so punctuation + whitespace removes both in one call.
Figure (svg): The state of the program after each line of Worked example the string module's two useful constants, drawn as a ladder with one rung per traced line
Two ready-made strings of characters. The book's aside is that this is how you can make Python swear — printing string.punctuation produces a line of symbols.
Verify: Check that strip removes from both ends only.
Why: 'a,b'.strip(string.punctuation) gives 'a,b' unchanged, because the comma is in the middle. strip removes characters from the ends until it meets one that is not in the set — which is exactly what you want for words, and would be wrong for removing punctuation everywhere.
Prediction
One cleaning step has been left out.
for word in 'The dog. A dog'.split():
word = word.strip('.')
hist[word] = hist.get(word, 0) + 1
# lower() was omitted| Word | After cleaning | Counted as |
|---|---|---|
| 'The' | stripped, not lowered | counted as 'The' |
| 'dog.' | the full stop removed | 'dog' |
| 'dog' | already clean | 'dog' again: 2 |
Predict first
How many different words does the histogram end up with?
Correct: 3 — 'The', 'A' and 'dog', because the two occurrences of dog merge but nothing merges the capitalised words.
Why: Without lower(), any word that appears capitalised at the start of a sentence and lowercase elsewhere is counted as two different words — which inflates the vocabulary count substantially in real text. Here it happens not to bite, because 'The' and 'A' each appear once, but in a whole book the effect is large.
Worked example
strip cannot help with punctuation in the middle.
>>> line = 'well-known dog,'
>>> line.split()
['well-known', 'dog,']
>>> line.replace('-', ' ').split()
['well', 'known', 'dog,']| Approach | What happens | Note |
|---|---|---|
| split alone | 'well-known' stays one word | the hyphen is not whitespace |
| replace first | the hyphen becomes a space | then split separates them |
| the decision | one word or two? | a judgement, not a bug |
Notice split only separates on whitespace.
Why: A hyphen is punctuation in the middle of a word, so split leaves it attached.
Notice strip cannot reach it either.
Why: strip works from the ends inward and stops at the first character it is not removing.
Replace before splitting.
Why: Turning hyphens into spaces makes split treat the two halves as separate words — which is the decision this program makes.
Figure (svg): Two columns comparing the effect of splitting hyphenated words or keeping them whole
Two words rather than one. The book's own solution replaces hyphens with spaces before splitting, and that is a choice about what counts as a word rather than a technical necessity.
Verify: Ask whether the choice is right.
Why: It depends on the question. Counting vocabulary, well-known is arguably one word; counting word frequencies for text generation, splitting is more useful because the halves recur separately. Being able to name the decision, rather than absorbing it as a step, is the point of a case study.
Trap
A program calls word.strip(string.punctuation) on its own line and then uses word.
Treat strip like a list method
Why: append and sort modify in place, so strip looks like it should too.
Strings are immutable, so strip returns a new string and the original is untouched. The word keeps its punctuation and the histogram counts dog, separately from dog, with no error at all.
Assign the result.
word = word.strip(...)
Why: The book's own solution writes it this way, on each of the three cleaning lines.
Remember why it must be so
Why: It is a shorthand to say that strings are converted; since strings are immutable, methods like strip and lower return new strings.
This is lesson 10c's returning-versus-modifying distinction in the place it causes the quietest damage: the program runs, produces a histogram, and every count is subtly wrong.
Ranking
Four steps, and the order matters for one of them.
Put in order
Why: The hyphen replacement must come before the split, because it is what causes the split to happen in the right places. Stripping and lowering both operate on individual words, so they come after — and their order relative to each other does not matter, since neither affects what the other removes.
Faded example
One call, two constants.
Fill in the blanks
import string
word = word.strip(string.punctuation + string.whitespace)
Why: string.whitespace contains space, tab, newline and the other invisible characters, and concatenating it with punctuation gives strip one set covering both. Note the assignment: strip returns a new string, so calling it without assigning leaves the word exactly as it was.
Real world
The same data, cleaned differently, gives different results.
Discussion prompt
Think of a situation where counting things from text gave a surprising answer. What kind of cleaning decision could have caused it?
Hint: What counts as the same thing?
Answer:
Search results that miss an obvious match because of a plural or an accent; a survey where N/A, n/a and NA are three answers; a tally where trailing spaces split one category into two.
Every one is a normalisation failure — two things that mean the same and do not look the same, so a program treats them as different.
Which is why this exercise puts the cleaning first. The counting is trivial and the counting is not where the errors are; deciding what counts as the same word is the whole difficulty, and it never has a purely technical answer.
Section
Section 2
Concept
The second exercise asks you to count the total number of words in the book, and the number of times each word is used — and then to print the number of different words used.
# total words: add up every count
sum(hist.values())
# different words: how many items
len(hist)| Measure | What it counts | Emma |
|---|---|---|
| sum of the values | every occurrence | 161,080 for Emma |
| number of items | every distinct word | 7,214 for Emma |
| the ratio | each word used ~22 times | on average |
Both are called how many words in English and they are not the same question. One measures the length of the book and the other measures the size of its vocabulary.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126
Picture it
Two numbers from one histogram, computed in one line each.
Figure (svg): A histogram with the two different totals derived from it marked
A dictionary answers both questions immediately, which is one reason it is the right structure here. A plain list of every word would answer the first easily and the second only by searching.
Worked example
Each is one line, and each reads a different part of the dictionary.
def total_words(hist):
return sum(hist.values())
def different_words(hist):
return len(hist)| Expression | What it reads | Note |
|---|---|---|
| hist.values() | the counts | one per distinct word |
| sum | adds them | every occurrence, once each |
| len(hist) | the number of items | the vocabulary |
Add up the frequencies for the total.
Why: To count the total number of words in the file, we can add up the frequencies in the histogram — every occurrence contributed 1 to some counter.
Count the items for the vocabulary.
Why: The number of different words is just the number of items in the dictionary, since each distinct word created exactly one item.
Notice both are free.
Why: Neither needs a loop or a second pass over the text; the histogram already contains both answers.
Figure (svg): The state of the program after each line of Worked example the two functions, drawn as a ladder with one rung per traced line
161,080 total words and 7,214 different words for Emma. Two one-line functions, reading the two halves of the same dictionary.
Verify: Check the two against each other.
Why: The total must be at least the number of different words, since each distinct word occurs at least once — and here it is more than twenty times larger. If total_words ever returned less than different_words, some counter would have to hold zero, which the histogram construction makes impossible.
Prediction
Two ways of counting one histogram.
hist = {'the': 5, 'dog': 2, 'ran': 1}
print(sum(hist.values()), len(hist))| Measure | Computation | Result |
|---|---|---|
| sum of values | 5 + 2 + 1 | 8 |
| len | three items | 3 |
| the relationship | total >= different | always |
Predict first
What does this print?
Correct: 8 3 — eight words were written, using three different words.
Why: sum(hist.values()) adds the counts and gives the length of the text; len(hist) counts the items and gives the vocabulary. The total can never be smaller than the number of different words, since every distinct word contributes at least 1 — a relationship worth using as a sanity check on any histogram.
Worked example
The exercise's real question, and why the raw number misleads.
# book A: 161,080 words, 7,214 different
# book B: 40,000 words, 5,100 different
# which author has the larger vocabulary?
# 7214 / 161080 = 0.045
# 5100 / 40000 = 0.128| Measure | Which wins | Note |
|---|---|---|
| by raw count | book A has more | 7,214 against 5,100 |
| by ratio | book B is richer | 0.128 against 0.045 |
| why they differ | a longer book repeats more | length inflates the count |
Take the raw vocabulary count.
Why: The longer book has more different words, which is almost inevitable — more text means more chances for a rare word to appear.
Divide by the length.
Why: The proportion of distinct words to total words is a fairer comparison, and it reverses the answer here.
Note the remaining problem.
Why: Even the ratio falls as a text gets longer, because common words keep recurring — so comparing books of very different lengths is genuinely hard.
Figure (svg): Two columns comparing raw vocabulary count against vocabulary as a proportion of length
The two measures disagree, and neither is simply right. The exercise asks which author uses the most extensive vocabulary, and answering it honestly means saying how you measured.
Verify: Test the ratio's weakness deliberately.
Why: Take the first 40,000 words of the longer book and count again: the ratio rises sharply, because the same vocabulary is now spread over less text. That confirms the measure depends on length, which is why serious comparisons use samples of equal size — a data decision, not a programming one.
Trap
A program reads a Project Gutenberg file straight through and reports the vocabulary.
Read the whole file
Why: It is a text file and the words are in it.
The file begins with several hundred lines of licensing header, so words like copyright, gutenberg and ebook enter the histogram and the vocabulary count includes text the author never wrote.
Skip over the header information at the beginning of the file.
Find the marker line that ends it
Why: Gutenberg files carry a recognisable start-of-text line to look for.
Ignore everything until you have seen it
Why: A boolean flag switched on at the marker, which is the flag use from lesson 11c.
The exercise names this explicitly, and it is the kind of thing that produces a plausible wrong answer: the counts are all slightly off and nothing looks broken.
Discrimination
Two numbers, and each answers different questions.
Sort into buckets
For each question, which measure answers it?
Faded example
Add up every counter in the histogram.
Fill in the blanks
def total_words(hist):
return sum(hist.values())
Why: The values are the counts, so summing them gives every occurrence. sum(hist) would add the keys instead and raise a TypeError, since they are strings; len(hist) would give the vocabulary rather than the length. Each of the three is one word different and answers a different question.
Socratic
A list of every word would also work.
Discussion prompt
You could store every word in a list, in order, and compute both counts from it. What would that cost, and what would it gain?
Hint: How would you count the different words?
Answer:
The length would be free — len of the list. But counting different words would mean checking each word against everything seen so far, which is a linear search inside a loop over 161,080 words.
The dictionary gives both answers immediately, because the merging happened as the text was read: each occurrence either created an item or incremented one.
What the list gains is order, which the dictionary throws away. If you wanted to generate text — which 13c does — you would need that order back, and the chapter's answer is to keep a different structure for that question. Choosing per question rather than once is the case study's real lesson.
Section
Section 3
Concept
Given the same inputs, most computer programs generate the same outputs every time, so they are said to be deterministic. Determinism is usually a good thing, since we expect the same calculation to yield the same result. For some applications, though, we want the computer to be unpredictable.
deterministic — Pertaining to a program that does the same thing each time it runs, given the same inputs.
Making a program truly nondeterministic turns out to be difficult, but there are ways to make it at least seem nondeterministic. One of them is to use algorithms that generate pseudorandom numbers — not truly random, because they are generated by a deterministic computation, but just by looking at the numbers it is all but impossible to distinguish them from random.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-126
Picture it
A calculation you could repeat exactly, producing output you could not predict.
Figure (svg): A flowchart showing a deterministic computation producing a sequence indistinguishable from random
The reproducibility is a feature rather than a flaw: starting from the same value gives the same sequence, which is what makes a program using randomness testable at all.
Worked example
Two programs, and only one of them should surprise you.
# deterministic: the same answer every time
def total_words(hist):
return sum(hist.values())
# nondeterministic by design
import random
def roll():
return random.randint(1, 6)| Function | Behaviour | Note |
|---|---|---|
| total_words | same input, same output | and it must be |
| roll | same call, different output | and it must be |
| the difference | what the function is for | not a property of good code |
Note the default expectation.
Why: Determinism is usually a good thing, since we expect the same calculation to yield the same result — a word count that varied between runs would be broken.
Note where it is wrong.
Why: Games are an obvious example: a die that rolled the same number every time would not be a die.
Note that unpredictability is deliberate.
Why: It has to be asked for, by importing a module and calling a function. Nothing becomes random by accident.
Figure (svg): The state of the program after each line of Worked example why determinism is usually what you want, drawn as a ladder with one rung per traced line
Two correct functions with opposite requirements. Which behaviour is right depends entirely on what the function is for.
Verify: Ask what makes the random version testable.
Why: Because it is pseudorandom, fixing the starting value makes the sequence repeat exactly — so a test can check a specific outcome. A truly random source could not be tested that way at all, which is one practical advantage of pseudo.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C is false: each time you call random you get the next number in a long series, and each is computed from the one before by a fixed rule. That relationship is exactly what makes the sequence reproducible — which is a nuisance for cryptography and an advantage for testing and debugging.
Worked example
Not random, and indistinguishable from random.
import random
for i in range(10):
x = random.random()
print(x)
# 0.1836...
# 0.9573...
# 0.2214... each is the next in a long series| Aspect | What is true | Note |
|---|---|---|
| each call | the next number in a long series | not a fresh coin flip |
| the series | computed deterministically | from the one before |
| to an observer | indistinguishable from random | which is the point |
Note the sequence is fixed.
Why: Each time you call random, you get the next number in a long series — the series exists whether or not you look at it.
Note the computation is ordinary.
Why: Pseudorandom numbers are generated by a deterministic computation, the same kind of arithmetic as anything else in the program.
Note why it is good enough.
Why: Just by looking at the numbers it is all but impossible to distinguish them from random, which is all most applications need.
Figure (svg): Two columns contrasting truly random with pseudorandom generation
A deterministic sequence that passes for random. The distinction matters for cryptography and rarely for anything else you will write.
Verify: Ask what would break if it were genuinely random.
Why: Reproducibility. A bug that appears one run in fifty would be almost impossible to investigate if the run could not be repeated — whereas with a pseudorandom source, recording the starting value lets you reproduce the exact sequence that caused it. The determinism underneath is what makes debugging possible.
Trap
A student decides a function using random cannot be tested, because the output changes every run.
Assume unpredictable means uncheckable
Why: A test needs an expected answer, and there is not one.
Plenty is still checkable: that every result is within range, that all possible outcomes eventually appear, that the proportions are about right over many calls. Abandoning testing entirely gives up all of it.
Test the properties rather than the value.
Check the range
Why: Every randint(1, 6) must be between 1 and 6 — an invariant that holds on every run.
Check the distribution over many calls
Why: Ten thousand rolls should hit all six faces, roughly evenly.
And because the numbers are pseudorandom rather than truly random, fixing the starting value makes the whole sequence repeat — so an exact test is possible too. The determinism underneath is what rescues the testing.
Prediction
The same program, run twice.
def total_words(hist):
return sum(hist.values())
# run twice on the same histogram| Aspect | What is true | Note |
|---|---|---|
| the inputs | identical | the same dictionary |
| the computation | deterministic | no randomness |
| the outputs | identical | necessarily |
Predict first
Will the two runs give the same answer?
Correct: Yes — given the same inputs, most computer programs generate the same outputs every time, and nothing here introduces randomness.
Why: Dictionary order is unpredictable, which is why option B is tempting — but addition does not care about order, so the sum is the same however the items are traversed. Determinism is the default and unpredictability has to be asked for explicitly, by importing random and calling one of its functions.
Explain it to yourself
The book says it turns out to be difficult, without saying why.
Discussion prompt
Why can a program not simply produce a truly random number, when it can produce anything else you ask for?
Hint: What does a program have to work from?
Answer:
A program is a fixed sequence of operations on values it has. Every output is computed from something, and anything computed from known inputs is predictable in principle.
So randomness cannot come from the computation itself — it has to come from outside, from something physically unpredictable like electrical noise or the timing of external events.
That is why true randomness needs special hardware or operating-system support, and why the ordinary route is a pseudorandom algorithm that is deterministic but looks random. The book's word difficult is really about where the unpredictability could possibly come from.
Discrimination
Ask what the caller expects of a repeated run.
Sort into buckets
For each program, should the same inputs give the same output?
Section
Section 4
Concept
The random module provides functions that generate pseudorandom numbers. Three of them cover almost everything you will need.
import random
random.random() # a float, 0.0 <= x < 1.0
random.randint(5, 10) # an integer, 5 <= n <= 10
random.choice([1, 2, 3]) # one element of the sequence| Function | What it returns | Range |
|---|---|---|
| random() | a float between 0.0 and 1.0 | including 0.0, excluding 1.0 |
| randint(low, high) | an integer | including BOTH ends |
| choice(t) | an element | chosen uniformly |
The module also provides functions to generate random values from continuous distributions including Gaussian, exponential, gamma and a few more — but these three are the ones this chapter uses.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-126
Picture it
The two numeric functions differ in whether they include the top.
Figure (svg): Three random functions with their return types and inclusive or exclusive bounds
That difference catches people, because range(5, 10) excludes 10 and randint(5, 10) includes it — two functions with similar-looking arguments and opposite conventions.
Worked example
Unlike range, which everyone has already learned.
>>> random.randint(5, 10)
5
>>> random.randint(5, 10)
9
>>> list(range(5, 10))
[5, 6, 7, 8, 9] # 10 is NOT included| Call | Possible values | Note |
|---|---|---|
| randint(5, 10) | 5, 6, 7, 8, 9 or 10 | six possible values |
| range(5, 10) | 5 through 9 | five values |
| the difference | randint includes high | range excludes stop |
Read randint's contract.
Why: It takes parameters low and high and returns an integer between low and high, including both.
Compare with range.
Why: range(5, 10) stops before 10, which is the convention everywhere else in Python.
Remember the exception.
Why: randint is the one that includes its upper bound, and it is the source of a great many off-by-one errors.
Figure (svg): The state of the program after each line of Worked example randint includes both ends, drawn as a ladder with one rung per traced line
Six possible values from randint and five from range, for the same-looking arguments. The two conventions genuinely differ.
Verify: Check by simulating a die.
Why: randint(1, 6) gives the six faces of a die correctly; range(1, 6) would give five. Using a die as the test case makes the convention memorable, because everyone knows how many faces there should be.
Prediction
randint includes both ends.
random.randint(5, 10)| Bound | Included? | Note |
|---|---|---|
| low | 5, and it is possible | included |
| high | 10, and it is possible | also included |
| the count | high - low + 1 | 6 |
Predict first
How many different values can this call return?
Correct: 6 — the values 5, 6, 7, 8, 9 and 10, since randint includes both ends.
Why: randint takes parameters low and high and returns an integer between them, including both — which is high minus low plus one possible values. This differs from range(5, 10), which gives five values and stops before 10. The inconsistency between the two is a genuine source of off-by-one errors.
Worked example
One element, chosen uniformly.
>>> t = [1, 2, 3]
>>> random.choice(t)
2
>>> random.choice(t)
3
>>> random.choice(['a', 'a', 'b']) # 'a' twice as likely| Call | What happens | Note |
|---|---|---|
| choice(t) | one element | each equally likely |
| a list with repeats | repeats raise the odds | 'a' has two chances in three |
| the mechanism | uniform over POSITIONS | not over distinct values |
Note it chooses an element.
Why: To choose an element from a sequence at random, you can use choice — each position is equally likely.
Note what repeats do.
Why: A value appearing twice occupies two positions, so it is twice as likely to be chosen. The weighting comes from the data, not from an option.
See where this is heading.
Why: That observation is the whole of exercise 13.5: a list containing each word once per occurrence would give exactly the frequencies of the text.
Figure (svg): Two columns comparing weighting a random choice by repeating list entries or by using a histogram
A uniformly chosen element — which becomes a weighted choice if the sequence contains repeats. That is the bridge to the next idea.
Verify: Check the two ways to weight a choice.
Why: Building a list with repeats and calling choice gives the right proportions and uses memory in proportion to the total word count; computing from a histogram gives the same proportions from a structure the size of the vocabulary. Both are correct, and for Emma that is 161,080 entries against 7,214 — which is exactly the data structure decision the chapter is about.
Trap
A program simulates a die with random.randint(1, 7), by analogy with range(1, 7).
Apply Python's usual half-open convention
Why: range, slices and every other range-like thing exclude the upper end.
randint includes both ends, so this die has seven faces. The program runs, produces plausible numbers, and is wrong about once in seven rolls — which is exactly the frequency that makes a bug hard to notice and hard to dismiss.
randint(low, high) includes high.
randint(1, 6) for a die
Why: Six values, both ends included.
Check the count of possible values
Why: high - low + 1, which is the giveaway that both ends are in.
If the half-open convention is what you want, randrange(1, 7) behaves like range. Having both available means the choice is yours — but randint is the one the book uses and the one that breaks the pattern.
Sorting
A float, an integer, or an element.
Sort into buckets
For each task, which random function fits?
Faded example
Six faces, both ends included.
Fill in the blanks
import random
def roll():
return random.randint(1, 6)
Why: randint includes both ends, so randint(1, 6) gives exactly the six faces. Writing 7 by analogy with range(1, 7) would produce a seven-sided die — a bug that shows up about one roll in seven and looks like bad luck rather than a mistake.
Explain it
Two similar-looking calls with opposite conventions.
Discussion prompt
A classmate's die simulation occasionally produces a 7. Explain what happened and give them a way to remember the difference.
Hint: Count the possible values in each.
Answer:
They wrote randint(1, 7), copying range's convention. randint includes both ends, so 7 is a possible result — the die has seven faces.
The rule is that randint is the exception: everything else in Python that takes a range excludes the upper end, and randint does not.
The way to remember it is to count: randint(low, high) can return high - low + 1 different values, and that plus one is the tell. Checking against a die, where you know there should be six, catches it every time.
Section
Section 5
Concept
Exercise 13.5 asks for a function named choose_from_hist that takes a histogram and returns a random value from it, chosen with probability in proportion to frequency.
>>> t = ['a', 'a', 'b']
>>> hist = histogram(t)
>>> hist
{'a': 2, 'b': 1}
# choose_from_hist(hist) should return
# 'a' with probability 2/3 and 'b' with probability 1/3| Part | What it means | Note |
|---|---|---|
| the histogram | 'a': 2, 'b': 1 | three occurrences in total |
| 'a' | two of the three | probability 2/3 |
| 'b' | one of the three | probability 1/3 |
This is where the chapter's two halves meet: a histogram from chapter 11, and a random choice from this section. It is also the piece that makes the text generator of the last lesson possible.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-127
Picture it
Each word occupies a share of the line in proportion to its count.
Figure (svg): A ladder showing cumulative ranges for two words weighted by their counts
That is the key move: the randomness stays uniform, and the weighting comes from how much of the range each word occupies.
Worked example
Rebuild the list and let choice do the work.
import random
def choose_from_hist(hist):
t = []
for word, freq in hist.items():
t.extend([word] * freq)
return random.choice(t)| Part | What it does | Note |
|---|---|---|
| the loop | one entry per occurrence | 'a', 'a', 'b' |
| choice(t) | uniform over positions | 'a' has two of three |
| the cost | a list the size of the text | 161,080 for Emma |
Undo the histogram.
Why: Each word is repeated as many times as its count, rebuilding the list of every occurrence.
Let choice weight it.
Why: choice is uniform over positions, and a word occupying more positions is proportionally more likely.
Note the cost.
Why: The list has one entry per word occurrence, so the memory is proportional to the length of the text rather than the size of the vocabulary.
Figure (svg): The state of the program after each line of Worked example the straightforward solution, drawn as a ladder with one rung per traced line
Correct probabilities, obtained by throwing away the compression the histogram provided. It works, and it undoes the reason the histogram existed.
Verify: Check the proportions empirically.
Why: Calling it thirty thousand times on {'a': 2, 'b': 1} should give roughly twenty thousand 'a's — a consistency check of the kind lesson 11c described, and the only practical way to test a function whose output is random. A result near ten thousand would mean the weighting was ignored.
Prediction
Two words, three occurrences.
hist = {'a': 2, 'b': 1}
# choose_from_hist(hist)| Word | Share | Probability |
|---|---|---|
| total occurrences | 2 + 1 | 3 |
| 'a' | 2 of the 3 | 2/3 |
| 'b' | 1 of the 3 | 1/3 |
Predict first
With what probability should the function return 'a'?
Correct: 2/3 — 'a' accounts for two of the three occurrences, and the choice is proportional to frequency.
Why: The exercise states this exactly: the function should return 'a' with probability 2/3 and 'b' with probability 1/3. Option B is what choosing uniformly from the keys would give — two distinct words, one chance each — which ignores the frequencies entirely and is the commonest wrong answer to this exercise.
Worked example
Same probabilities, without rebuilding the text.
def choose_from_hist(hist):
total = sum(hist.values())
n = random.randint(1, total)
for word, freq in hist.items():
n = n - freq
if n <= 0:
return word| Step | What it does | Note |
|---|---|---|
| total | every occurrence | 3 for our example |
| n = randint(1, 3) | a uniform position | 1, 2 or 3 |
| subtracting freq | walks the ranges | 'a' covers 1 and 2 |
| n <= 0 | this word's range contains n | return it |
Find the total.
Why: The sum of the frequencies is the number of occurrences, which is the size of the range to pick from.
Pick a uniform position in that range.
Why: randint(1, total) with both ends included, so every occurrence has an equal chance.
Walk the words, subtracting.
Why: Each word's frequency is the width of its range; when n drops to zero or below, the chosen position fell inside this word's share.
Figure (svg): A flowchart showing a random position walked down through each word's frequency
The same probabilities, using a structure the size of the vocabulary rather than of the text. For Emma that is 7,214 items instead of 161,080.
Verify: Check the two ends of the range.
Why: n = 1 must return the first word and n = total the last, and both do — the first because subtracting its frequency takes n to zero or below immediately, the last because everything before it was subtracted without reaching zero. Testing the extremes is what confirms an off-by-one has not crept into the comparison.
Trap
A function returns random.choice(list(hist)) to pick a word from a histogram.
Pick a random word from the words
Why: The keys are the words, so choosing among them looks right.
That gives every distinct word an equal chance, so the is exactly as likely as a word used once. The frequencies — the entire content of the histogram — are ignored, and the output looks nothing like the text.
Weight by frequency, which is what the histogram is for.
The probability must be proportional to the count
Why: 'a' with 2 and 'b' with 1 means two chances in three.
Either rebuild the occurrences or walk the cumulative counts
Why: Both give the same distribution; they differ in memory.
The tell is that the output uses rare words as often as common ones. Sampled text that reads as a list of unusual words has almost always been sampled uniformly from the keys.
Invariant
Step through the cumulative subtraction.
Step through it
Which values of n return 'a', and which return 'b'?
n of 1 or 2 returns 'a' and n of 3 returns 'b' — two of three positions against one, which is exactly the 2/3 and 1/3 the exercise asks for. The randomness never stopped being uniform; the widths of the ranges did the weighting.
Faded example
How many occurrences are there altogether?
Fill in the blanks
def choose_from_hist(hist):
total = sum(hist.values())
n = random.randint(1, total)
for word, freq in hist.items():
n = n - freq
if n <= 0:
return word
Why: The values are the frequencies, so summing them gives the number of occurrences to draw a position from — the same total_words computation as idea 2. Using len(hist) instead would draw from the vocabulary size and give every word an equal chance, which is exactly the bug that ignores the weighting.
Real world
Picking fairly and picking proportionally are different things.
Discussion prompt
Think of a situation where a random choice should not be uniform. What decides the weights, and what would go wrong with an equal chance?
Hint: Anything where some options should come up more often.
Answer:
A raffle where more tickets mean better odds; a playlist that favours what you listen to; a simulation drawing from observed frequencies rather than assuming everything is equally likely.
The weights come from the data — tickets held, plays recorded, occurrences counted — which is exactly what a histogram is.
With an equal chance the rare cases would appear as often as the common ones, which for a simulation means the model no longer resembles what it is modelling. That is precisely the failure that makes uniformly-sampled text unreadable, and it is why the next lessons build on this function rather than on choice alone.
Comparison
Fill the blanks. Two return numbers and one returns an element.
Comparison matrix
| Question | random() | randint(a, b) | choice(t) |
|---|---|---|---|
| What does it return? | a float | an integer | an element of t |
| Its range | 0.0 included, 1.0 excluded | a and b both included | the sequence's elements |
| Uniform over what? | the interval | the integers in range | the positions in t |
| How do you weight it? | compare against a threshold | walk cumulative counts | repeat entries in t |
The last row is the chapter's point: all three are uniform, and weighting always comes from arranging the data, never from the function.
Pattern
Six steps, and the first three are all about deciding what counts as one word.
Step 3's three operations all return new strings rather than modifying, so each result must be assigned. Forgetting one of the three is silent and skews every count.
Python documentation — Input and Output Input and Output
Check
Both ends are included.
Check your understanding
How many values can random.randint(1, 6) return?
Answer: A
Why: randint takes parameters low and high and returns an integer between them, including both — so 1, 2, 3, 4, 5 and 6, which is exactly the six faces of a die. This differs from range(1, 6), which gives five values because it stops before the upper bound.
Check
One histogram, two questions.
hist = {'the': 100, 'cat': 3, 'sat': 1}
print(sum(hist.values()), len(hist))| Measure | Computation | Result |
|---|---|---|
| sum of values | 100 + 3 + 1 | 104 |
| len | three items | 3 |
| what each means | length, then vocabulary | different questions |
Check your understanding
What does this print?
Answer: A
Why: Summing the values gives every occurrence — the length of the text — and len gives the number of items, which is the vocabulary. Both are called how many words in English and they answer different questions, which is why the exercise asks for both.
Check
One of these ignores the frequencies.
Check your understanding
Which expression picks a word from a histogram in proportion to how often it appears?
Answer: B
Why: Drawing a position uniformly from the total number of occurrences and then walking the frequencies gives each word a share of the range equal to its count — which is exactly proportional weighting, and it keeps the histogram rather than rebuilding the text.
Real world
Counting things from messy text is one of the most common real tasks there is.
Discussion prompt
Think of a time something was counted from text or forms and the total looked wrong. What normalisation step was probably missing?
Hint: Two entries that meant the same and did not match.
Answer:
Categories split by capitalisation or a trailing space; the same company appearing three times with different punctuation; a survey where free text was tallied without any cleaning at all.
Every one is the same failure as counting Dog, dog and dog, separately — two things that mean the same and do not look the same.
Which is why the cleaning comes first and takes most of the code. The counting is four characters of arithmetic; deciding what counts as the same thing is where the judgement and nearly all the errors are.
Commit first
Answer, then rate your confidence.
Predict first
How many different values can random.randint(1, 6) return?
Correct: 6 — randint takes low and high and returns an integer between them, including both.
Why: This is the one place where Python breaks its own convention. range(1, 6), slices, and everything else range-like exclude the upper bound; randint includes it. The count of possible values is high minus low plus one, and that plus one is the giveaway. Writing randint(1, 7) for a die — the natural translation from range — produces a seven-sided die that rolls wrong about one time in seven, which is frequent enough to matter and rare enough to look like chance. If the half-open convention is what you want, randrange(1, 7) behaves exactly like range.
Explain it
A random choice that is not uniform.
Discussion prompt
A classmate's text generator produces sentences full of unusual words. Diagnose it and tell them what to change.
Hint: How is it choosing?
Answer:
They are almost certainly choosing uniformly from the histogram's keys, which gives every distinct word an equal chance — so the, used five thousand times, is exactly as likely as a word used once.
Since most of a vocabulary is rare words, uniform sampling produces mostly rare words. The output is a list of oddities rather than anything resembling the text.
The fix is to weight by frequency: draw a position from the total number of occurrences and walk the counts until the position is used up. The randomness stays uniform, and the widths of the ranges do the weighting.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The cleaning is where most of the code and nearly all the judgement live, and every step exists because of a specific way two occurrences of one word can fail to match. The two counts are trivial to compute and easy to confuse, which is why the exercise asks for both. The pseudorandom idea matters mostly for understanding why a random program is still testable. And the weighted choice is the piece the rest of the chapter is built on — get that one solid and the text generator two lessons from now is mostly bookkeeping.
Connect it up
One page, from memory.
Draw it
Draw the pipeline from a raw line of text to a clean word, labelling each stage with the method that performs it and the kind of difference it removes. Beside it, write the two counts and what each measures. Underneath, list the three random functions with their return types and mark which one includes both ends of its range. Finally, sketch the range-walking picture that turns a uniform draw into a choice weighted by frequency.
Recap
Two pages, and the case study's problem is set.
| If you remember one thing | It is this |
|---|---|
| From cleaning | Every step exists because two occurrences of a word can fail to look alike. |
| From the counts | How many words is two questions, and they differ by a factor of twenty. |
| From pseudorandomness | Deterministic underneath, which is what makes a random program testable. |
| From randint | It includes both ends. range does not. |
| From weighted choice | The randomness stays uniform; the data supply the weights. |
The next lesson gives the book's solution to these exercises: process_file and process_line, the two counting functions, most_common with its sorted list of tuples, optional parameters, and dictionary subtraction for finding words that are not in the word list.
Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.