13a Word Frequency Analysis and Random Numbers

This lesson opens the case study by setting out the word-frequency problem and the string tools for cleaning text, then introduces determinism, pseudorandom numbers, and the random module's three core functions.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 13a Word Frequency Analysis and Random Numbers

Title

Python · Chapter 13 — Case study: data structure selection

§13.1-13.2, pp. 125-126

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126 — the pages these objectives are drawn from

3. Before we start: how many words is that?

Warm-up

A question that sounds simple and is not.

Discussion prompt

You are asked how many words are in a book. Before writing any code, list three decisions you would have to make before the number is even well defined.

Hint: Is The the same word as the?

Answer:

Case: are The and the the same word? Punctuation: is dog, the word dog? Hyphens: is well-known one word or two?

And a fourth: do you mean how many words were written, or how many different words were used? Those are completely different numbers and both are called how many words.

This chapter is a case study in exactly these decisions. The programming is straightforward; choosing what to count and which structure to count it in is the actual work.

4. The one idea behind this chapter: choosing the structure is the problem

Concept

At this point you have learned about Python's core data structures, and you have seen some of the algorithms that use them. This chapter presents a case study with exercises that let you think about choosing data structures and practice using them.

Every question in this chapter can be answered with a list, a dictionary, or a list of tuples — and the choice decides whether the program takes a second or an hour. That is what the chapter is about, and it is why it comes after all three types rather than alongside any one of them.

Figure (svg): Two columns pairing a question about a text with the structure that answers it well

Four questions about one text, and four different shapes.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-125

5. Cleaning words out of text

Section

Section 1

6. Four steps between a line and its words

Concept

The first exercise sets the whole problem: write a program that reads a file, breaks each line into words, strips whitespace and punctuation from the words, and converts them to lowercase.

Each step exists because of a way two occurrences of the same word can fail to look alike. Skip any one of them and the histogram counts dog and dog, and Dog as three different words.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-125

7. Picture it: one word through the pipeline

Picture it

Each stage removes one kind of difference.

Figure (svg): A pipeline showing a raw line reduced to a clean lowercase word

Split, strip, lower. Miss a stage and the same word is counted twice.

Note that each stage returns a new string rather than changing one — it is a shorthand to say strings are converted, since strings are immutable.

8. Worked example: the string module's two useful constants

Worked example

You do not have to write out the punctuation yourself.

>>> import string
>>> string.punctuation
'!"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~'
>>> string.whitespace
' \t\n\r\x0b\x0c'
ConstantWhat it holdsNote
string.punctuationevery punctuation character32 of them
string.whitespacespace, tab, newline and morethe invisible ones
togethereverything to stripconcatenate them

Import the module.

Why: The string module provides a string named whitespace, which contains space, tab, newline and so on, and punctuation which contains the punctuation characters.

Look at what they hold.

Why: Both are ordinary strings, so they can be concatenated, indexed, and passed to any method that takes one.

Use them together.

Why: strip takes a string of characters to remove, so punctuation + whitespace removes both in one call.

Figure (svg): The state of the program after each line of Worked example the string module's two useful constants, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Two ready-made strings of characters. The book's aside is that this is how you can make Python swear — printing string.punctuation produces a line of symbols.

Verify: Check that strip removes from both ends only.

Why: 'a,b'.strip(string.punctuation) gives 'a,b' unchanged, because the comma is in the middle. strip removes characters from the ends until it meets one that is not in the set — which is exactly what you want for words, and would be wrong for removing punctuation everywhere.

9. Predict: what does the histogram count?

Prediction

One cleaning step has been left out.

for word in 'The dog. A dog'.split():
    word = word.strip('.')
    hist[word] = hist.get(word, 0) + 1
# lower() was omitted
WordAfter cleaningCounted as
'The'stripped, not loweredcounted as 'The'
'dog.'the full stop removed'dog'
'dog'already clean'dog' again: 2

Predict first

How many different words does the histogram end up with?

  • 3
  • 2
  • 4
  • 1

Correct: 3 — 'The', 'A' and 'dog', because the two occurrences of dog merge but nothing merges the capitalised words.

Why: Without lower(), any word that appears capitalised at the start of a sentence and lowercase elsewhere is counted as two different words — which inflates the vocabulary count substantially in real text. Here it happens not to bite, because 'The' and 'A' each appear once, but in a whole book the effect is large.

10. Worked example: why the hyphen needs separate treatment

Worked example

strip cannot help with punctuation in the middle.

>>> line = 'well-known dog,'
>>> line.split()
['well-known', 'dog,']
>>> line.replace('-', ' ').split()
['well', 'known', 'dog,']
ApproachWhat happensNote
split alone'well-known' stays one wordthe hyphen is not whitespace
replace firstthe hyphen becomes a spacethen split separates them
the decisionone word or two?a judgement, not a bug

Notice split only separates on whitespace.

Why: A hyphen is punctuation in the middle of a word, so split leaves it attached.

Notice strip cannot reach it either.

Why: strip works from the ends inward and stops at the first character it is not removing.

Replace before splitting.

Why: Turning hyphens into spaces makes split treat the two halves as separate words — which is the decision this program makes.

Figure (svg): Two columns comparing the effect of splitting hyphenated words or keeping them whole

Neither is wrong. The right answer depends on what the count is for.

Two words rather than one. The book's own solution replaces hyphens with spaces before splitting, and that is a choice about what counts as a word rather than a technical necessity.

Verify: Ask whether the choice is right.

Why: It depends on the question. Counting vocabulary, well-known is arguably one word; counting word frequencies for text generation, splitting is more useful because the halves recur separately. Being able to name the decision, rather than absorbing it as a step, is the point of a case study.

11. Trap: forgetting that string methods return rather than modify

Trap

The trap

A program calls word.strip(string.punctuation) on its own line and then uses word.

Treat strip like a list method

Why: append and sort modify in place, so strip looks like it should too.

Strings are immutable, so strip returns a new string and the original is untouched. The word keeps its punctuation and the histogram counts dog, separately from dog, with no error at all.

The fix

Assign the result.

word = word.strip(...)

Why: The book's own solution writes it this way, on each of the three cleaning lines.

Remember why it must be so

Why: It is a shorthand to say that strings are converted; since strings are immutable, methods like strip and lower return new strings.

This is lesson 10c's returning-versus-modifying distinction in the place it causes the quietest damage: the program runs, produces a histogram, and every count is subtly wrong.

12. Rank: the cleaning steps, in order

Ranking

Four steps, and the order matters for one of them.

Put in order

  1. replace hyphens with spaces
  2. split the line into words
  3. strip punctuation and whitespace from each word
  4. convert to lowercase

Why: The hyphen replacement must come before the split, because it is what causes the split to happen in the right places. Stripping and lowering both operate on individual words, so they come after — and their order relative to each other does not matter, since neither affects what the other removes.

13. Complete it: strip both kinds of character

Faded example

One call, two constants.

Fill in the blanks

import string
word = word.strip(string.punctuation + string.whitespace)

Why: string.whitespace contains space, tab, newline and the other invisible characters, and concatenating it with punctuation gives strip one set covering both. Note the assignment: strip returns a new string, so calling it without assigning leaves the word exactly as it was.

14. Where text cleaning decides the answer

Real world

The same data, cleaned differently, gives different results.

Discussion prompt

Think of a situation where counting things from text gave a surprising answer. What kind of cleaning decision could have caused it?

Hint: What counts as the same thing?

Answer:

Search results that miss an obvious match because of a plural or an accent; a survey where N/A, n/a and NA are three answers; a tally where trailing spaces split one category into two.

Every one is a normalisation failure — two things that mean the same and do not look the same, so a program treats them as different.

Which is why this exercise puts the cleaning first. The counting is trivial and the counting is not where the errors are; deciding what counts as the same word is the whole difficulty, and it never has a purely technical answer.

15. Two different counts

Section

Section 2

16. Total words and different words

Concept

The second exercise asks you to count the total number of words in the book, and the number of times each word is used — and then to print the number of different words used.

# total words: add up every count
sum(hist.values())

# different words: how many items
len(hist)
MeasureWhat it countsEmma
sum of the valuesevery occurrence161,080 for Emma
number of itemsevery distinct word7,214 for Emma
the ratioeach word used ~22 timeson average

Both are called how many words in English and they are not the same question. One measures the length of the book and the other measures the size of its vocabulary.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126

17. Picture it: length against vocabulary

Picture it

Two numbers from one histogram, computed in one line each.

Figure (svg): A histogram with the two different totals derived from it marked

The values added give the length; the items counted give the vocabulary.

A dictionary answers both questions immediately, which is one reason it is the right structure here. A plain list of every word would answer the first easily and the second only by searching.

18. Worked example: the two functions

Worked example

Each is one line, and each reads a different part of the dictionary.

def total_words(hist):
    return sum(hist.values())

def different_words(hist):
    return len(hist)
ExpressionWhat it readsNote
hist.values()the countsone per distinct word
sumadds themevery occurrence, once each
len(hist)the number of itemsthe vocabulary

Add up the frequencies for the total.

Why: To count the total number of words in the file, we can add up the frequencies in the histogram — every occurrence contributed 1 to some counter.

Count the items for the vocabulary.

Why: The number of different words is just the number of items in the dictionary, since each distinct word created exactly one item.

Notice both are free.

Why: Neither needs a loop or a second pass over the text; the histogram already contains both answers.

Figure (svg): The state of the program after each line of Worked example the two functions, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

161,080 total words and 7,214 different words for Emma. Two one-line functions, reading the two halves of the same dictionary.

Verify: Check the two against each other.

Why: The total must be at least the number of different words, since each distinct word occurs at least once — and here it is more than twenty times larger. If total_words ever returned less than different_words, some counter would have to hold zero, which the histogram construction makes impossible.

19. Predict: which is larger?

Prediction

Two ways of counting one histogram.

hist = {'the': 5, 'dog': 2, 'ran': 1}
print(sum(hist.values()), len(hist))
MeasureComputationResult
sum of values5 + 2 + 18
lenthree items3
the relationshiptotal >= differentalways

Predict first

What does this print?

  • 8 3
  • 3 8
  • 8 8
  • 3 3

Correct: 8 3 — eight words were written, using three different words.

Why: sum(hist.values()) adds the counts and gives the length of the text; len(hist) counts the items and gives the vocabulary. The total can never be smaller than the number of different words, since every distinct word contributes at least 1 — a relationship worth using as a sanity check on any histogram.

20. Worked example: comparing authors

Worked example

The exercise's real question, and why the raw number misleads.

# book A: 161,080 words, 7,214 different
# book B:  40,000 words, 5,100 different

# which author has the larger vocabulary?
# 7214 / 161080 = 0.045
# 5100 /  40000 = 0.128
MeasureWhich winsNote
by raw countbook A has more7,214 against 5,100
by ratiobook B is richer0.128 against 0.045
why they differa longer book repeats morelength inflates the count

Take the raw vocabulary count.

Why: The longer book has more different words, which is almost inevitable — more text means more chances for a rare word to appear.

Divide by the length.

Why: The proportion of distinct words to total words is a fairer comparison, and it reverses the answer here.

Note the remaining problem.

Why: Even the ratio falls as a text gets longer, because common words keep recurring — so comparing books of very different lengths is genuinely hard.

Figure (svg): Two columns comparing raw vocabulary count against vocabulary as a proportion of length

Both are one line of code. Choosing between them is the actual work.

The two measures disagree, and neither is simply right. The exercise asks which author uses the most extensive vocabulary, and answering it honestly means saying how you measured.

Verify: Test the ratio's weakness deliberately.

Why: Take the first 40,000 words of the longer book and count again: the ratio rises sharply, because the same vocabulary is now spread over less text. That confirms the measure depends on length, which is why serious comparisons use samples of equal size — a data decision, not a programming one.

21. Trap: forgetting to skip the header

Trap

The trap

A program reads a Project Gutenberg file straight through and reports the vocabulary.

Read the whole file

Why: It is a text file and the words are in it.

The file begins with several hundred lines of licensing header, so words like copyright, gutenberg and ebook enter the histogram and the vocabulary count includes text the author never wrote.

The fix

Skip over the header information at the beginning of the file.

Find the marker line that ends it

Why: Gutenberg files carry a recognisable start-of-text line to look for.

Ignore everything until you have seen it

Why: A boolean flag switched on at the marker, which is the flag use from lesson 11c.

The exercise names this explicitly, and it is the kind of thing that produces a plausible wrong answer: the counts are all slightly off and nothing looks broken.

22. Discriminate: which count answers this?

Discrimination

Two numbers, and each answers different questions.

Sort into buckets

For each question, which measure answers it?

sum(hist.values())
how long is this book?; how many words would I have to read?; how many times was any word written?
len(hist)
how large is this author's vocabulary?; how many entries would the index have?; how many distinct terms need defining?
total
Each asks about occurrences — how much text there is. Every appearance of every word counts separately, which is what adding the frequencies gives.
diff
Each asks about distinct words. Repetition is irrelevant: a word used five thousand times contributes exactly one item to the dictionary, and one entry to an index.

23. Complete it: the total word count

Faded example

Add up every counter in the histogram.

Fill in the blanks

def total_words(hist):
return sum(hist.values())

Why: The values are the counts, so summing them gives every occurrence. sum(hist) would add the keys instead and raise a TypeError, since they are strings; len(hist) would give the vocabulary rather than the length. Each of the three is one word different and answers a different question.

24. Think it through: why is a dictionary the right structure here?

Socratic

A list of every word would also work.

Discussion prompt

You could store every word in a list, in order, and compute both counts from it. What would that cost, and what would it gain?

Hint: How would you count the different words?

Answer:

The length would be free — len of the list. But counting different words would mean checking each word against everything seen so far, which is a linear search inside a loop over 161,080 words.

The dictionary gives both answers immediately, because the merging happened as the text was read: each occurrence either created an item or incremented one.

What the list gains is order, which the dictionary throws away. If you wanted to generate text — which 13c does — you would need that order back, and the chapter's answer is to keep a different structure for that question. Choosing per question rather than once is the case study's real lesson.

25. Determinism and pseudorandom numbers

Section

Section 3

26. Programs repeat; sometimes you do not want them to

Concept

Given the same inputs, most computer programs generate the same outputs every time, so they are said to be deterministic. Determinism is usually a good thing, since we expect the same calculation to yield the same result. For some applications, though, we want the computer to be unpredictable.

deterministic — Pertaining to a program that does the same thing each time it runs, given the same inputs.

Making a program truly nondeterministic turns out to be difficult, but there are ways to make it at least seem nondeterministic. One of them is to use algorithms that generate pseudorandom numbers — not truly random, because they are generated by a deterministic computation, but just by looking at the numbers it is all but impossible to distinguish them from random.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-126

27. Picture it: deterministic underneath, unpredictable on the surface

Picture it

A calculation you could repeat exactly, producing output you could not predict.

Figure (svg): A flowchart showing a deterministic computation producing a sequence indistinguishable from random

Every number follows from the last by a fixed rule, and the sequence is still unpredictable to look at.

The reproducibility is a feature rather than a flaw: starting from the same value gives the same sequence, which is what makes a program using randomness testable at all.

28. Worked example: why determinism is usually what you want

Worked example

Two programs, and only one of them should surprise you.

# deterministic: the same answer every time
def total_words(hist):
    return sum(hist.values())

# nondeterministic by design
import random
def roll():
    return random.randint(1, 6)
FunctionBehaviourNote
total_wordssame input, same outputand it must be
rollsame call, different outputand it must be
the differencewhat the function is fornot a property of good code

Note the default expectation.

Why: Determinism is usually a good thing, since we expect the same calculation to yield the same result — a word count that varied between runs would be broken.

Note where it is wrong.

Why: Games are an obvious example: a die that rolled the same number every time would not be a die.

Note that unpredictability is deliberate.

Why: It has to be asked for, by importing a module and calling a function. Nothing becomes random by accident.

Figure (svg): The state of the program after each line of Worked example why determinism is usually what you want, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Two correct functions with opposite requirements. Which behaviour is right depends entirely on what the function is for.

Verify: Ask what makes the random version testable.

Why: Because it is pseudorandom, fixing the starting value makes the sequence repeat exactly — so a test can check a specific outcome. A truly random source could not be tested that way at all, which is one practical advantage of pseudo.

29. Two truths and a lie: pseudorandom numbers

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. They are generated by a deterministic computation
  • B. Just by looking at them it is all but impossible to distinguish them from random
  • C. Each call produces a number unrelated to the ones before it

Survives elimination: C

Why: C is false: each time you call random you get the next number in a long series, and each is computed from the one before by a fixed rule. That relationship is exactly what makes the sequence reproducible — which is a nuisance for cryptography and an advantage for testing and debugging.

30. Worked example: what pseudorandom means precisely

Worked example

Not random, and indistinguishable from random.

import random
for i in range(10):
    x = random.random()
    print(x)

# 0.1836...
# 0.9573...
# 0.2214...   each is the next in a long series
AspectWhat is trueNote
each callthe next number in a long seriesnot a fresh coin flip
the seriescomputed deterministicallyfrom the one before
to an observerindistinguishable from randomwhich is the point

Note the sequence is fixed.

Why: Each time you call random, you get the next number in a long series — the series exists whether or not you look at it.

Note the computation is ordinary.

Why: Pseudorandom numbers are generated by a deterministic computation, the same kind of arithmetic as anything else in the program.

Note why it is good enough.

Why: Just by looking at the numbers it is all but impossible to distinguish them from random, which is all most applications need.

Figure (svg): Two columns contrasting truly random with pseudorandom generation

The reproducibility in the right-hand column is a feature everywhere except security.

A deterministic sequence that passes for random. The distinction matters for cryptography and rarely for anything else you will write.

Verify: Ask what would break if it were genuinely random.

Why: Reproducibility. A bug that appears one run in fifty would be almost impossible to investigate if the run could not be repeated — whereas with a pseudorandom source, recording the starting value lets you reproduce the exact sequence that caused it. The determinism underneath is what makes debugging possible.

31. Trap: expecting a random program to be untestable

Trap

The trap

A student decides a function using random cannot be tested, because the output changes every run.

Assume unpredictable means uncheckable

Why: A test needs an expected answer, and there is not one.

Plenty is still checkable: that every result is within range, that all possible outcomes eventually appear, that the proportions are about right over many calls. Abandoning testing entirely gives up all of it.

The fix

Test the properties rather than the value.

Check the range

Why: Every randint(1, 6) must be between 1 and 6 — an invariant that holds on every run.

Check the distribution over many calls

Why: Ten thousand rolls should hit all six faces, roughly evenly.

And because the numbers are pseudorandom rather than truly random, fixing the starting value makes the whole sequence repeat — so an exact test is possible too. The determinism underneath is what rescues the testing.

32. Predict: will these two runs agree?

Prediction

The same program, run twice.

def total_words(hist):
    return sum(hist.values())

# run twice on the same histogram
AspectWhat is trueNote
the inputsidenticalthe same dictionary
the computationdeterministicno randomness
the outputsidenticalnecessarily

Predict first

Will the two runs give the same answer?

  • Yes — given the same inputs, the program is deterministic
  • No — dictionary order is unpredictable, so the sum varies
  • Only if the dictionary was built the same way
  • Only on the same machine

Correct: Yes — given the same inputs, most computer programs generate the same outputs every time, and nothing here introduces randomness.

Why: Dictionary order is unpredictable, which is why option B is tempting — but addition does not care about order, so the sum is the same however the items are traversed. Determinism is the default and unpredictability has to be asked for explicitly, by importing random and calling one of its functions.

33. Explain it yourself: why is true randomness hard?

Explain it to yourself

The book says it turns out to be difficult, without saying why.

Discussion prompt

Why can a program not simply produce a truly random number, when it can produce anything else you ask for?

Hint: What does a program have to work from?

Answer:

A program is a fixed sequence of operations on values it has. Every output is computed from something, and anything computed from known inputs is predictable in principle.

So randomness cannot come from the computation itself — it has to come from outside, from something physically unpredictable like electrical noise or the timing of external events.

That is why true randomness needs special hardware or operating-system support, and why the ordinary route is a pseudorandom algorithm that is deterministic but looks random. The book's word difficult is really about where the unpredictability could possibly come from.

34. Discriminate: should this be deterministic?

Discrimination

Ask what the caller expects of a repeated run.

Sort into buckets

For each program, should the same inputs give the same output?

deterministic
counting the words in a book; computing an average; sorting a list of names
unpredictable by design
dealing a hand of cards; shuffling a playlist; picking a daily puzzle from a bank
det
Each is a calculation whose answer is a fact about the input. We expect the same calculation to yield the same result, and a word count that varied between runs would simply be broken.
non
Each would be useless if it repeated. Games are the obvious example, and a shuffle or a daily selection that gave the same answer every time would defeat its own purpose.

35. The random module's three functions

Section

Section 4

36. A float, an integer, and an element

Concept

The random module provides functions that generate pseudorandom numbers. Three of them cover almost everything you will need.

import random

random.random()          # a float, 0.0 <= x < 1.0
random.randint(5, 10)    # an integer, 5 <= n <= 10
random.choice([1, 2, 3]) # one element of the sequence
FunctionWhat it returnsRange
random()a float between 0.0 and 1.0including 0.0, excluding 1.0
randint(low, high)an integerincluding BOTH ends
choice(t)an elementchosen uniformly

The module also provides functions to generate random values from continuous distributions including Gaussian, exponential, gamma and a few more — but these three are the ones this chapter uses.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-126

37. Picture it: the three ranges, and their ends

Picture it

The two numeric functions differ in whether they include the top.

Figure (svg): Three random functions with their return types and inclusive or exclusive bounds

randint includes both ends. random excludes its top. The inconsistency is worth memorising.

That difference catches people, because range(5, 10) excludes 10 and randint(5, 10) includes it — two functions with similar-looking arguments and opposite conventions.

38. Worked example: randint includes both ends

Worked example

Unlike range, which everyone has already learned.

>>> random.randint(5, 10)
5
>>> random.randint(5, 10)
9

>>> list(range(5, 10))
[5, 6, 7, 8, 9]        # 10 is NOT included
CallPossible valuesNote
randint(5, 10)5, 6, 7, 8, 9 or 10six possible values
range(5, 10)5 through 9five values
the differencerandint includes highrange excludes stop

Read randint's contract.

Why: It takes parameters low and high and returns an integer between low and high, including both.

Compare with range.

Why: range(5, 10) stops before 10, which is the convention everywhere else in Python.

Remember the exception.

Why: randint is the one that includes its upper bound, and it is the source of a great many off-by-one errors.

Figure (svg): The state of the program after each line of Worked example randint includes both ends, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Six possible values from randint and five from range, for the same-looking arguments. The two conventions genuinely differ.

Verify: Check by simulating a die.

Why: randint(1, 6) gives the six faces of a die correctly; range(1, 6) would give five. Using a die as the test case makes the convention memorable, because everyone knows how many faces there should be.

39. Predict: how many possible values?

Prediction

randint includes both ends.

random.randint(5, 10)
BoundIncluded?Note
low5, and it is possibleincluded
high10, and it is possiblealso included
the counthigh - low + 16

Predict first

How many different values can this call return?

  • 6
  • 5
  • 10
  • 11

Correct: 6 — the values 5, 6, 7, 8, 9 and 10, since randint includes both ends.

Why: randint takes parameters low and high and returns an integer between them, including both — which is high minus low plus one possible values. This differs from range(5, 10), which gives five values and stops before 10. The inconsistency between the two is a genuine source of off-by-one errors.

40. Worked example: choice, and what it assumes

Worked example

One element, chosen uniformly.

>>> t = [1, 2, 3]
>>> random.choice(t)
2
>>> random.choice(t)
3

>>> random.choice(['a', 'a', 'b'])   # 'a' twice as likely
CallWhat happensNote
choice(t)one elementeach equally likely
a list with repeatsrepeats raise the odds'a' has two chances in three
the mechanismuniform over POSITIONSnot over distinct values

Note it chooses an element.

Why: To choose an element from a sequence at random, you can use choice — each position is equally likely.

Note what repeats do.

Why: A value appearing twice occupies two positions, so it is twice as likely to be chosen. The weighting comes from the data, not from an option.

See where this is heading.

Why: That observation is the whole of exercise 13.5: a list containing each word once per occurrence would give exactly the frequencies of the text.

Figure (svg): Two columns comparing weighting a random choice by repeating list entries or by using a histogram

The same probabilities from structures twenty times different in size.

A uniformly chosen element — which becomes a weighted choice if the sequence contains repeats. That is the bridge to the next idea.

Verify: Check the two ways to weight a choice.

Why: Building a list with repeats and calling choice gives the right proportions and uses memory in proportion to the total word count; computing from a histogram gives the same proportions from a structure the size of the vocabulary. Both are correct, and for Emma that is 161,080 entries against 7,214 — which is exactly the data structure decision the chapter is about.

41. Trap: assuming randint excludes its upper bound

Trap

The trap

A program simulates a die with random.randint(1, 7), by analogy with range(1, 7).

Apply Python's usual half-open convention

Why: range, slices and every other range-like thing exclude the upper end.

randint includes both ends, so this die has seven faces. The program runs, produces plausible numbers, and is wrong about once in seven rolls — which is exactly the frequency that makes a bug hard to notice and hard to dismiss.

The fix

randint(low, high) includes high.

randint(1, 6) for a die

Why: Six values, both ends included.

Check the count of possible values

Why: high - low + 1, which is the giveaway that both ends are in.

If the half-open convention is what you want, randrange(1, 7) behaves like range. Having both available means the choice is yours — but randint is the one the book uses and the one that breaks the pattern.

42. Sort: which function do you need?

Sorting

A float, an integer, or an element.

Sort into buckets

For each task, which random function fits?

randint
simulate one roll of a six-sided die; pick a page number between 1 and 200
choice
pick a word from a list of words; choose a card from a deck list
random
decide something with probability 0.3; generate a proportion between 0 and 1
ri
Both want an integer in a range, with both ends possible — a die face and a page number are exactly that.
ch
Both pick an existing element from a sequence rather than generating a number, which is what choice does.
rf
Both want a float between 0.0 and 1.0 — one as a probability threshold, one as a proportion in its own right.

43. Complete it: roll a die

Faded example

Six faces, both ends included.

Fill in the blanks

import random

def roll():
return random.randint(1, 6)

Why: randint includes both ends, so randint(1, 6) gives exactly the six faces. Writing 7 by analogy with range(1, 7) would produce a seven-sided die — a bug that shows up about one roll in seven and looks like bad luck rather than a mistake.

44. Explain it: randint or range?

Explain it

Two similar-looking calls with opposite conventions.

Discussion prompt

A classmate's die simulation occasionally produces a 7. Explain what happened and give them a way to remember the difference.

Hint: Count the possible values in each.

Answer:

They wrote randint(1, 7), copying range's convention. randint includes both ends, so 7 is a possible result — the die has seven faces.

The rule is that randint is the exception: everything else in Python that takes a range excludes the upper end, and randint does not.

The way to remember it is to count: randint(low, high) can return high - low + 1 different values, and that plus one is the tell. Checking against a die, where you know there should be six, catches it every time.

45. Choosing from a histogram

Section

Section 5

46. A random word, weighted by how often it appears

Concept

Exercise 13.5 asks for a function named choose_from_hist that takes a histogram and returns a random value from it, chosen with probability in proportion to frequency.

>>> t = ['a', 'a', 'b']
>>> hist = histogram(t)
>>> hist
{'a': 2, 'b': 1}

# choose_from_hist(hist) should return
# 'a' with probability 2/3 and 'b' with probability 1/3
PartWhat it meansNote
the histogram'a': 2, 'b': 1three occurrences in total
'a'two of the threeprobability 2/3
'b'one of the threeprobability 1/3

This is where the chapter's two halves meet: a histogram from chapter 11, and a random choice from this section. It is also the piece that makes the text generator of the last lesson possible.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 126-127

47. Picture it: the weighting made explicit

Picture it

Each word occupies a share of the line in proportion to its count.

Figure (svg): A ladder showing cumulative ranges for two words weighted by their counts

Turning counts into ranges is what makes a uniform pick produce a weighted result.

That is the key move: the randomness stays uniform, and the weighting comes from how much of the range each word occupies.

48. Worked example: the straightforward solution

Worked example

Rebuild the list and let choice do the work.

import random

def choose_from_hist(hist):
    t = []
    for word, freq in hist.items():
        t.extend([word] * freq)
    return random.choice(t)
PartWhat it doesNote
the loopone entry per occurrence'a', 'a', 'b'
choice(t)uniform over positions'a' has two of three
the costa list the size of the text161,080 for Emma

Undo the histogram.

Why: Each word is repeated as many times as its count, rebuilding the list of every occurrence.

Let choice weight it.

Why: choice is uniform over positions, and a word occupying more positions is proportionally more likely.

Note the cost.

Why: The list has one entry per word occurrence, so the memory is proportional to the length of the text rather than the size of the vocabulary.

Figure (svg): The state of the program after each line of Worked example the straightforward solution, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Correct probabilities, obtained by throwing away the compression the histogram provided. It works, and it undoes the reason the histogram existed.

Verify: Check the proportions empirically.

Why: Calling it thirty thousand times on {'a': 2, 'b': 1} should give roughly twenty thousand 'a's — a consistency check of the kind lesson 11c described, and the only practical way to test a function whose output is random. A result near ten thousand would mean the weighting was ignored.

49. Predict: what are the odds?

Prediction

Two words, three occurrences.

hist = {'a': 2, 'b': 1}
# choose_from_hist(hist)
WordShareProbability
total occurrences2 + 13
'a'2 of the 32/3
'b'1 of the 31/3

Predict first

With what probability should the function return 'a'?

  • 2/3
  • 1/2
  • 1/3
  • 2

Correct: 2/3 — 'a' accounts for two of the three occurrences, and the choice is proportional to frequency.

Why: The exercise states this exactly: the function should return 'a' with probability 2/3 and 'b' with probability 1/3. Option B is what choosing uniformly from the keys would give — two distinct words, one chance each — which ignores the frequencies entirely and is the commonest wrong answer to this exercise.

50. Worked example: the version that keeps the histogram

Worked example

Same probabilities, without rebuilding the text.

def choose_from_hist(hist):
    total = sum(hist.values())
    n = random.randint(1, total)
    for word, freq in hist.items():
        n = n - freq
        if n <= 0:
            return word
StepWhat it doesNote
totalevery occurrence3 for our example
n = randint(1, 3)a uniform position1, 2 or 3
subtracting freqwalks the ranges'a' covers 1 and 2
n <= 0this word's range contains nreturn it

Find the total.

Why: The sum of the frequencies is the number of occurrences, which is the size of the range to pick from.

Pick a uniform position in that range.

Why: randint(1, total) with both ends included, so every occurrence has an equal chance.

Walk the words, subtracting.

Why: Each word's frequency is the width of its range; when n drops to zero or below, the chosen position fell inside this word's share.

Figure (svg): A flowchart showing a random position walked down through each word's frequency

A uniform pick, converted into a weighted one by walking the cumulative counts.

The same probabilities, using a structure the size of the vocabulary rather than of the text. For Emma that is 7,214 items instead of 161,080.

Verify: Check the two ends of the range.

Why: n = 1 must return the first word and n = total the last, and both do — the first because subtracting its frequency takes n to zero or below immediately, the last because everything before it was subtracted without reaching zero. Testing the extremes is what confirms an off-by-one has not crept into the comparison.

51. Trap: choosing uniformly from the keys

Trap

The trap

A function returns random.choice(list(hist)) to pick a word from a histogram.

Pick a random word from the words

Why: The keys are the words, so choosing among them looks right.

That gives every distinct word an equal chance, so the is exactly as likely as a word used once. The frequencies — the entire content of the histogram — are ignored, and the output looks nothing like the text.

The fix

Weight by frequency, which is what the histogram is for.

The probability must be proportional to the count

Why: 'a' with 2 and 'b' with 1 means two chances in three.

Either rebuild the occurrences or walk the cumulative counts

Why: Both give the same distribution; they differ in memory.

The tell is that the output uses rare words as often as common ones. Sampled text that reads as a list of unusual words has almost always been sampled uniformly from the keys.

52. Watch the walk: converting a uniform pick into a weighted one

Invariant

Step through the cumulative subtraction.

Step through it

Which values of n return 'a', and which return 'b'?

  1. Three occurrences in total, so the position is drawn from a range of three.
  2. Every position from 1 to 3 was equally likely, and 2 came up. The randomness is uniform.
  3. The first word's frequency is subtracted, taking n from 2 to 0.
  4. n has reached zero, which means the drawn position fell within this word's share of the range, so 'a' is returned.

n of 1 or 2 returns 'a' and n of 3 returns 'b' — two of three positions against one, which is exactly the 2/3 and 1/3 the exercise asks for. The randomness never stopped being uniform; the widths of the ranges did the weighting.

53. Complete it: the total to draw from

Faded example

How many occurrences are there altogether?

Fill in the blanks

def choose_from_hist(hist):
total = sum(hist.values())
n = random.randint(1, total)
for word, freq in hist.items():
n = n - freq
if n <= 0:
return word

Why: The values are the frequencies, so summing them gives the number of occurrences to draw a position from — the same total_words computation as idea 2. Using len(hist) instead would draw from the vocabulary size and give every word an equal chance, which is exactly the bug that ignores the weighting.

54. Where weighted random choice is used

Real world

Picking fairly and picking proportionally are different things.

Discussion prompt

Think of a situation where a random choice should not be uniform. What decides the weights, and what would go wrong with an equal chance?

Hint: Anything where some options should come up more often.

Answer:

A raffle where more tickets mean better odds; a playlist that favours what you listen to; a simulation drawing from observed frequencies rather than assuming everything is equally likely.

The weights come from the data — tickets held, plays recorded, occurrences counted — which is exactly what a histogram is.

With an equal chance the rare cases would appear as often as the common ones, which for a simulation means the model no longer resembles what it is modelling. That is precisely the failure that makes uniformly-sampled text unreadable, and it is why the next lessons build on this function rather than on choice alone.

55. Compare: the three random functions

Comparison

Fill the blanks. Two return numbers and one returns an element.

Comparison matrix

Questionrandom()randint(a, b)choice(t)
What does it return?a floatan integeran element of t
Its range0.0 included, 1.0 excludeda and b both includedthe sequence's elements
Uniform over what?the intervalthe integers in rangethe positions in t
How do you weight it?compare against a thresholdwalk cumulative countsrepeat entries in t

The last row is the chapter's point: all three are uniform, and weighting always comes from arranging the data, never from the function.

56. The procedure: counting words in a text

Pattern

Six steps, and the first three are all about deciding what counts as one word.

  1. Skip the header, if the file has one, so that only the author's words are counted.
  2. Read a line at a time and replace hyphens with spaces, before splitting.
  3. Split into words, then strip punctuation and whitespace from each and convert it to lowercase.
  4. Accumulate into a dictionary: hist[word] = hist.get(word, 0) + 1.
  5. Read the length off as sum(hist.values()) and the vocabulary as len(hist).
  6. For a weighted random word, draw a position from the total and walk the frequencies.

Step 3's three operations all return new strings rather than modifying, so each result must be assigned. Forgetting one of the three is silent and skews every count.

Python documentation — Input and Output Input and Output

57. Check yourself 1 of 3: randint's range

Check

Both ends are included.

Check your understanding

How many values can random.randint(1, 6) return?

  • A. 6 (correct)
  • B. 5
  • C. 7
  • D. It depends on the seed

Answer: A

Why: randint takes parameters low and high and returns an integer between them, including both — so 1, 2, 3, 4, 5 and 6, which is exactly the six faces of a die. This differs from range(1, 6), which gives five values because it stops before the upper bound.

Why B tempts people
This is what range(1, 6) would give. randint is the exception to Python's usual half-open convention.
Why C tempts people
Seven would need randint(1, 7), which is the classic off-by-one that produces a seven-sided die.
Why D tempts people
The seed determines which value comes up, not how many are possible. The range is fixed by the arguments.

58. Check yourself 2 of 3: two counts

Check

One histogram, two questions.

hist = {'the': 100, 'cat': 3, 'sat': 1}
print(sum(hist.values()), len(hist))
MeasureComputationResult
sum of values100 + 3 + 1104
lenthree items3
what each meanslength, then vocabularydifferent questions

Check your understanding

What does this print?

  • A. 104 3 (correct)
  • B. 3 104
  • C. 104 104
  • D. A TypeError

Answer: A

Why: Summing the values gives every occurrence — the length of the text — and len gives the number of items, which is the vocabulary. Both are called how many words in English and they answer different questions, which is why the exercise asks for both.

Why B tempts people
This reverses them. sum reads the counts and len reads the number of entries.
Why C tempts people
These can only be equal if every word appears exactly once, which no real text manages.
Why D tempts people
sum over the values adds integers, which is fine. It would be a TypeError only if applied to the keys.

59. Check yourself 3 of 3: weighted choice

Check

One of these ignores the frequencies.

Check your understanding

Which expression picks a word from a histogram in proportion to how often it appears?

  • A. random.choice(list(hist))
  • B. walking the frequencies after drawing randint(1, sum(hist.values())) (correct)
  • C. random.choice(list(hist.values()))
  • D. random.random() * len(hist)

Answer: B

Why: Drawing a position uniformly from the total number of occurrences and then walking the frequencies gives each word a share of the range equal to its count — which is exactly proportional weighting, and it keeps the histogram rather than rebuilding the text.

Why A tempts people
This picks uniformly among the distinct words, so a word used once is as likely as one used five thousand times.
Why C tempts people
This picks a random count rather than a word, returning a number with no way back to which word it belonged to.
Why D tempts people
This produces a float, not a word, and nothing about it consults the frequencies.

60. Where this shows up outside this course

Real world

Counting things from messy text is one of the most common real tasks there is.

Discussion prompt

Think of a time something was counted from text or forms and the total looked wrong. What normalisation step was probably missing?

Hint: Two entries that meant the same and did not match.

Answer:

Categories split by capitalisation or a trailing space; the same company appearing three times with different punctuation; a survey where free text was tallied without any cleaning at all.

Every one is the same failure as counting Dog, dog and dog, separately — two things that mean the same and do not look the same.

Which is why the cleaning comes first and takes most of the code. The counting is four characters of arithmetic; deciding what counts as the same thing is where the judgement and nearly all the errors are.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

How many different values can random.randint(1, 6) return?

  • 6 — both ends are included
  • 5 — the upper bound is excluded, as in range
  • 7 — both ends plus zero
  • It depends on the random seed

Correct: 6 — randint takes low and high and returns an integer between them, including both.

Why: This is the one place where Python breaks its own convention. range(1, 6), slices, and everything else range-like exclude the upper bound; randint includes it. The count of possible values is high minus low plus one, and that plus one is the giveaway. Writing randint(1, 7) for a die — the natural translation from range — produces a seven-sided die that rolls wrong about one time in seven, which is frequent enough to matter and rare enough to look like chance. If the half-open convention is what you want, randrange(1, 7) behaves exactly like range.

62. Explain it to someone else

Explain it

A random choice that is not uniform.

Discussion prompt

A classmate's text generator produces sentences full of unusual words. Diagnose it and tell them what to change.

Hint: How is it choosing?

Answer:

They are almost certainly choosing uniformly from the histogram's keys, which gives every distinct word an equal chance — so the, used five thousand times, is exactly as likely as a word used once.

Since most of a vocabulary is rare words, uniform sampling produces mostly rare words. The output is a list of oddities rather than anything resembling the text.

The fix is to weight by frequency: draw a position from the total number of occurrences and walk the counts until the position is used up. The randomness stays uniform, and the widths of the ranges do the weighting.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • The cleaning steps, and why each one is needed
  • Total words against different words
  • Determinism, and what pseudorandom means
  • The three random functions, and weighting a choice by frequency

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The cleaning is where most of the code and nearly all the judgement live, and every step exists because of a specific way two occurrences of one word can fail to match. The two counts are trivial to compute and easy to confuse, which is why the exercise asks for both. The pseudorandom idea matters mostly for understanding why a random program is still testable. And the weighted choice is the piece the rest of the chapter is built on — get that one solid and the text generator two lessons from now is mostly bookkeeping.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the pipeline from a raw line of text to a clean word, labelling each stage with the method that performs it and the kind of difference it removes. Beside it, write the two counts and what each measures. Underneath, list the three random functions with their return types and mark which one includes both ends of its range. Finally, sketch the range-walking picture that turns a uniform draw into a choice weighted by frequency.

65. What you can do now

Recap

Two pages, and the case study's problem is set.

If you remember one thingIt is this
From cleaningEvery step exists because two occurrences of a word can fail to look alike.
From the countsHow many words is two questions, and they differ by a factor of twenty.
From pseudorandomnessDeterministic underneath, which is what makes a random program testable.
From randintIt includes both ends. range does not.
From weighted choiceThe randomness stays uniform; the data supply the weights.

The next lesson gives the book's solution to these exercises: process_file and process_line, the two counting functions, most_common with its sorted list of tuples, optional parameters, and dictionary subtraction for finding words that are not in the word list.

Think Python, 2nd edition — Allen B. Downey §13.1-13.2, pp. 125-126 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §13.1-13.2, pp. 125-126
  2. Python documentation — Input and Output
  3. Python documentation — random — Generate pseudo-random numbers

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108