This lesson opens a real word list with 113,809 entries, reads it line by line, and then writes four search functions that are the same pattern with one thing changed each time — ending with the development plan of reducing a problem to one already solved.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 9 — Case study: word play
§9.1-9.3, pp. 83-85
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 83-85 — the pages these objectives are drawn from
Warm-up
Every string you have processed so far was typed into the program. That stops now.
Discussion prompt
You are asked to find every English word longer than twenty letters. What is the first thing you need, and why can it not be a string literal in your program?
Hint: How many words are there?
Answer:
A list of English words — and the one this chapter uses has 113,809 entries, which nobody is going to type into a program.
So the data has to come from a file, which means the program has to be able to open one and read it.
Chapter 14 explains files properly. This chapter borrows just enough — open, readline, and a for loop over a file — because the puzzles are the point and the data has to come from somewhere.
Concept
This chapter presents a second case study, solving word puzzles by searching for words with certain properties — and with it a third program development plan: reduction to a previously solved problem.
reduction to a previously solved problem — Solving a problem by expressing it as an instance of a problem you have already solved.
Chapter 4 gave encapsulation and generalization, which build a solution up from working code. This one is different in kind: it is about noticing that no new solution is needed at all.
Figure (svg): A diagram showing three development plans in the order the book introduces them
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 83-83
Section
Section 1
Concept
The built-in function open takes the name of the file as a parameter and returns a file object you can use to read the file. The file object provides several methods for reading, including readline.
file object — A value that represents an open file.
>>> fin = open('words.txt')
>>> fin.readline()
'aa\n'
>>> fin.readline()
'aah\n'| Line | What happens | Result |
|---|---|---|
| open('words.txt') | returns a file object | assigned to fin |
| fin.readline() | reads up to a newline | 'aa' plus the newline |
| fin.readline() again | the file remembers where it is | the next word |
fin is a common name for a file object used for input. readline reads characters from the file until it gets to a newline and returns the result as a string — including the newline, which is usually not wanted.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 83-83
Picture it
Two identical calls give different answers, because the object keeps track.
Figure (svg): A diagram showing a file object with a position marker that advances after each readline call
That is why the same expression gives a different answer each time — which is unlike anything you have met so far, where an expression's value depended only on its parts.
Worked example
readline gives you one character you did not ask for. Remove it.
>>> line = fin.readline()
>>> line
'aahed\n'
>>> word = line.strip()
>>> word
'aahed'| Value | What it is | Note |
|---|---|---|
| line | what readline returned | includes the newline |
| line.strip() | a string method | returns a new string without it |
| word | the clean word | 'aahed' |
Notice the extra character.
Why: The sequence backslash-n represents the newline character that separates this word from the next. readline includes it.
Remove it with strip.
Why: A string method, exactly as in lesson 8c — and like every string method it returns a new string rather than modifying the original.
Assign the result.
Why: line.strip() on its own would compute a clean word and discard it, which is the wasted-return-value mistake from lesson 8c.
Figure (svg): The state of the program after each line of Worked example stripping the newline, drawn as a ladder with one rung per traced line
'aahed' without the newline. strip returns a new string, so the result has to be assigned or it is lost.
Verify: Compare the lengths before and after.
Why: len(line) is one more than len(word), which is the newline. Checking the length is a quicker confirmation than looking at the printed forms, since a trailing newline is nearly invisible on screen — and an unstripped word is a classic source of comparisons that mysteriously fail.
Prediction
One thing more than the word.
Predict first
The first line of words.txt contains the word aa. What does the first call to readline return?
Correct: 'aa\n' — the word plus the newline character that separated it from the next line.
Why: readline reads characters from the file until it gets to a newline and returns the result as a string, including that newline. It is a single character even though it is typed as two, which is the escape sequence from lesson 5c. Removing it with strip is the first thing nearly every file-reading loop does.
Worked example
A file object works in a for loop, exactly like a string.
fin = open('words.txt')
for line in fin:
word = line.strip()
print(word)| Part | What it does | Note |
|---|---|---|
| for line in fin | one pass per line of the file | 113,809 passes |
| line | includes its newline | 'aa\n' |
| line.strip() | the clean word | 'aa' |
Use the file object where a sequence would go.
Why: You can also use a file object as part of a for loop, and the loop variable holds one line per pass.
Strip each line.
Why: Every line arrives with its newline attached, so the strip goes inside the loop.
Notice the shape.
Why: This is the traversal pattern from lesson 8a with a file instead of a string — same loop, different sequence.
Figure (svg): A diagram showing a file being opened, traversed line by line, and each line stripped to a clean word
Every word in the file, one per line, with the newlines removed. The for loop handles a file exactly as it handles a string.
Verify: Compare with the readline version.
Why: Calling readline in a while loop would need an explicit test for the end of the file. The for loop knows when the file is exhausted, which is the same advantage it had over an index traversal in lesson 8a: the bookkeeping is built in and cannot be got wrong.
Trap
A program reads words from a file and compares them with string literals, and none of the comparisons ever succeeds.
Assume readline gives you the word
Why: The printed output looks exactly like the word, because a newline is invisible.
'aa\n' is not equal to 'aa'. The comparison is False for every word in the file, and nothing on screen shows why.
Strip every line as soon as it arrives.
Do it in the loop, before anything else
Why: word = line.strip() as the first statement of the body, so the rest of the code deals only in clean words.
Check with len when a comparison fails unexpectedly
Why: A word one character longer than it looks is a trailing newline.
This is the same convert at the boundary habit as lesson 5c's advice about input: clean the data where it enters the program, so that nothing downstream has to remember.
Discrimination
Ask whether the trailing newline would change the answer.
Sort into buckets
For each use of a line read from a file, is stripping necessary?
Faded example
Two blanks: how to traverse, and how to clean.
Fill in the blanks
fin = open('words.txt')
for line in fin:
word = line.strip()
print(word)
Why: A file object can be used directly in a for loop, giving one line per pass, and strip removes the trailing newline. Note that strip is a method call and needs its parentheses, and that its result must be assigned — line.strip() on its own would compute the clean word and discard it, which is lesson 8c's most common string mistake.
Socratic
Every expression you have met so far depended only on its parts.
Discussion prompt
fin.readline() is the same expression each time and returns something different on each call. Explain what makes that possible, and name one thing it makes harder.
Hint: The file object is remembering something.
Answer:
The file object keeps track of where it is in the file, and reading advances that position. So the expression's value depends on state held inside the object, not only on the expression itself.
What it makes harder is reasoning: you cannot work out what a line does by looking at it, because the answer depends on what has happened before.
This is the first stateful thing in the book — every earlier expression was a pure computation on its arguments. Chapter 15 makes objects with state a deliberate subject, and it is worth noticing here that files got there first.
Section
Section 2
Concept
All of the chapter's exercises have something in common: they can be solved with the search pattern from chapter 8. The simplest example returns False as soon as it finds a disqualifying letter, and True only after the loop has finished.
def has_no_e(word):
for letter in word:
if letter == 'e':
return False
return True| Part | When it happens | Note |
|---|---|---|
| the traversal | one pass per letter | the for loop |
| return False | as soon as an 'e' is found | inside the loop |
| return True | the loop finished without finding one | after the loop |
If we find the letter e we can immediately return False; otherwise we go to the next letter. If we exit the loop normally, that means we did not find an e, so we return True. You could write this more concisely with the in operator, but this version demonstrates the logic of the search pattern.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 85-85
Picture it
One counterexample disqualifies; only an exhausted loop qualifies.
Figure (svg): A flow chart showing a search that returns False on the first disqualifying letter and True after the loop completes
This is exactly lesson 8b's search pattern with the two return values swapped: there, finding something returned a position; here, finding something disqualifies the word.
Worked example
One change turns has_no_e into a function that handles any letters at all.
def avoids(word, forbidden):
for letter in word:
if letter in forbidden:
return False
return True| Aspect | What it is | Note |
|---|---|---|
| the structure | identical to has_no_e | same pattern |
| the condition | letter in forbidden, not letter == 'e' | the only change |
| the second parameter | a string of letters to avoid | generalization |
Notice what stayed the same.
Why: The traversal, the early return of False, and the final return of True. avoids is a more general version of has_no_e but it has the same structure.
Notice what changed.
Why: The condition uses the in operator from lesson 8c to test membership in a string of forbidden letters, instead of comparing against one fixed letter.
Recognise the move.
Why: This is generalization, from lesson 4b: a fixed value became a parameter.
Figure (svg): Two columns comparing has_no_e with avoids, showing that only the condition differs
The same search pattern with the condition generalized. has_no_e(word) is exactly avoids(word, 'e'), which is worth noticing before writing both.
Verify: Check that the general version reproduces the specific one.
Why: avoids('banana', 'e') gives True and avoids('bee', 'e') gives False, matching has_no_e on both. A generalization must reproduce the special case it came from — lesson 4b's verification step, applied here.
Prediction
The return True has escaped into the loop.
def has_no_e(word):
for letter in word:
if letter == 'e':
return False
else:
return True| Input | What happens | Result |
|---|---|---|
| 'bee' | first letter is 'b', not 'e' | the else runs |
| return True | on the first pass | before any 'e' is seen |
| the result | True | wrong: 'bee' has two e's |
Predict first
What does this version return for 'bee'?
Correct: True, incorrectly — the else returns on the first letter, which is a 'b', before any 'e' is examined.
Why: The function reports that 'bee' has no e. One acceptable letter is not evidence that all the letters are acceptable, so the True return cannot go inside the loop. Note that the bug is invisible for any word starting with 'e', which would correctly return False — so a badly chosen test would pass.
Worked example
uses_only asks the opposite question with the same machinery.
def uses_only(word, available):
for letter in word:
if letter not in available:
return False
return True| Function | The condition | What disqualifies |
|---|---|---|
| avoids | if letter in forbidden | disqualify on a match |
| uses_only | if letter not in available | disqualify on a NON-match |
| the difference | the word not | the sense is reversed |
Compare the two conditions.
Why: uses_only is similar except that the sense of the condition is reversed.
State what each list means.
Why: Instead of a list of forbidden letters, we have a list of available letters. If we find a letter in word that is not in available, we can return False.
Notice the structure is untouched.
Why: Same traversal, same two returns, same places. Only the test changed.
Figure (svg): The state of the program after each line of Worked example reversing the sense of the condition, drawn as a ladder with one rung per traced line
The same pattern with a negated condition. avoids disqualifies a word for containing a listed letter; uses_only disqualifies it for containing an unlisted one.
Verify: Check the two on the same inputs and confirm they differ.
Why: avoids('hoe', 'acefhlo') is False, because h, o and e are all forbidden; uses_only('hoe', 'acefhlo') is True, because every letter is available. Same arguments, opposite answers — which confirms that the negation genuinely reverses the question rather than merely rewording it.
Trap
A student writes an else that returns True when a letter is acceptable.
Treat each letter as decisive in both directions
Why: It feels symmetric: a bad letter disqualifies, so a good letter should qualify.
It does not. One acceptable letter proves nothing about the rest of the word, so the function returns True after examining only the first letter — reporting success for words that contain forbidden letters later on.
The two conclusions need different amounts of evidence.
Return False from inside the loop
Why: One disqualifying letter is enough to settle it.
Return True after the loop
Why: Because the word only qualifies once every letter has been checked.
This is the same asymmetry as lesson 8b's search, and it is worth saying in general terms: a claim about EVERY element can be disproved by one counterexample and proved only by exhaustion.
Matching
Four functions, four conditions, one shared structure.
Match the pairs
Why: Three of the four traverse the word; the fourth traverses the required letters instead, which is why the book says it reverses the role of the word and the string of letters. Noticing that the loop changes what it traverses — not merely what it tests — is the step that leads to the reduction in the next section.
Faded example
Traverse the required letters, not the word.
Fill in the blanks
def uses_all(word, required):
for letter in required:
if letter not in word:
return False
return True
Why: The loop traverses the required letters rather than the word, because the question is whether each required letter appears — not whether each of the word's letters is required. Traversing the word instead would answer uses_only, which is a different question. This swap of which sequence is traversed is exactly what makes the reduction in the next section possible.
Explain it to yourself
State the general principle, not just the rule.
Discussion prompt
In all four of these functions, False is returned inside the loop and True after it. Explain the general principle behind that, in terms of what each conclusion requires as evidence.
Hint: One conclusion is about every letter; the other is about one letter.
Answer:
True is a claim about EVERY letter — that none is forbidden, or that all are available. A claim about every element can only be established by checking every element, so it has to wait until the loop is done.
False is a claim about ONE letter — that this particular one disqualifies the word. One counterexample settles it, so it can be returned immediately.
The general principle is that universal claims need exhaustion and existential claims need one example. That asymmetry decides where every return in a search goes, and it is the same fact that made the not-found return in lesson 8b sit after the loop.
Section
Section 3
Concept
If you were really thinking like a computer scientist, you would have recognised that uses_all was an instance of a previously solved problem, and you would have written it in one line.
def uses_all(word, required):
return uses_only(required, word)| Function | What it asks | Note |
|---|---|---|
| uses_only(a, b) | every letter of a is in b | the general shape |
| uses_all(word, required) | every letter of required is in word | the same shape |
| the reduction | swap the arguments | no new logic at all |
This is an example of a program development plan called reduction to a previously solved problem, which means that you recognise the problem you are working on as an instance of a solved problem and apply an existing solution.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 85-85
Picture it
Both ask whether every letter of one string appears in the other.
Figure (svg): Two columns showing uses_only and uses_all asking the same question with the arguments in opposite orders
Once the shape is visible, uses_all(word, required) is just uses_only(required, word) — and the second function needs no body of its own.
Worked example
Write both questions in the same form and the swap becomes obvious.
# uses_only(word, available):
# is every letter of WORD in AVAILABLE?
# uses_all(word, required):
# is every letter of REQUIRED in WORD?
# same shape: is every letter of X in Y?| Function | What X and Y are | Note |
|---|---|---|
| uses_only | X = word, Y = available | the original |
| uses_all | X = required, Y = word | the same with X and Y swapped |
| therefore | uses_all(w, r) = uses_only(r, w) | the reduction |
Write both questions in a common form.
Why: Is every letter of X in Y fits both, with different things playing X and Y.
Compare the assignments.
Why: In one, the word plays X; in the other it plays Y. That is the whole difference.
Write the reduction.
Why: uses_all(word, required) is uses_only(required, word). No new loop, no new condition, no new returns.
Figure (svg): A call diagram showing uses_all delegating directly to uses_only with its arguments swapped
One line. Rewriting both problems in a common form is what makes the reduction visible — and the rewriting is the hard part, not the swap.
Verify: Test the reduced version against the hand-written one.
Why: Both give True for uses_all('banana', 'ban') and False for uses_all('banana', 'bane'). Agreement on both a positive and a negative case is the minimum check, and it is worth doing because a swapped-argument bug would pass one of them by luck.
Prediction
Two functions with the same shape and different questions.
def avoids(word, forbidden):
return not uses_only(word, forbidden)| Expression | What it actually asks | Note |
|---|---|---|
| uses_only(word, forbidden) | every letter of word IS forbidden | a different question |
| not that | at least one letter is not forbidden | not what avoids means |
| avoids | NO letter is forbidden | a third question |
Predict first
Is this a correct reduction?
Correct: No — negating every letter is forbidden gives at least one letter is not forbidden, which is a much weaker claim than no letter is forbidden.
Why: The negation of a universal claim is an existential one, not the universal claim about the opposite. avoids('bee', 'e') should be False, and uses_only('bee', 'e') is also False since 'b' is not in 'e', so the negation gives True — exactly wrong. This is why the sentence-form test matters: two functions can share a structure and ask questions that do not relate by negation at all.
Worked example
Three benefits, and one of them is about bugs.
# hand-written uses_all: 5 lines, its own loop and returns
# reduced uses_all: 1 line, delegating
# benefits:
# less code to write
# less code to get wrong
# a fix to uses_only fixes both| Benefit | What it means | Note |
|---|---|---|
| writing | one line instead of five | less work |
| correctness | no new loop to get wrong | fewer bugs |
| maintenance | one implementation, two callers | a fix helps both |
Count the code.
Why: One line instead of five, which is the visible benefit and the least important.
Count the opportunities for error.
Why: The hand-written version has a loop, a condition and two returns, each of which can be wrong. The reduced version has none of them.
Consider the future.
Why: If uses_only turns out to have a bug, fixing it fixes uses_all too — which is lesson 4b's de-duplication argument arriving in a new form.
Figure (svg): The state of the program after each line of Worked example what reduction is worth, drawn as a ladder with one rung per traced line
Less code, fewer places to be wrong, and one implementation to maintain. The middle benefit is the one that matters most, because a loop you did not write cannot have an off-by-one error.
Verify: Ask what the reduction costs.
Why: One extra function call, which is negligible, and a small loss of directness — a reader has to look at uses_only to see how it works. That is the same indirection cost as any function, and lesson 3c's four reasons for writing functions already accepted it.
Trap
A student notices two functions look similar and writes one in terms of the other without checking they ask the same question.
Match on shape rather than on meaning
Why: Two search functions look alike, so one must be an instance of the other.
uses_only and avoids have the same structure and ask opposite questions. Reducing one to the other would give exactly wrong answers, with no error to indicate it.
Write both problems out in a common form before reducing.
State each as a sentence with the same shape
Why: Is every letter of X in Y fits uses_only and uses_all. It does not fit avoids, whose answer is about NONE rather than every.
Then test the reduction on a positive and a negative case
Why: Agreement on both is the minimum evidence that the reduction is genuine.
Recognising an instance of a solved problem is a real skill, and the failure mode is recognising one that is not there. The sentence-form test is cheap and catches it.
Sorting
Write each pair as a sentence and compare.
Sort into buckets
For each proposed reduction, is it correct?
Analogy
Match each plan to the move it describes.
Match the pairs
Why: The first three produce code and the fourth avoids producing it, which is what makes reduction different in kind. It is also the hardest to apply deliberately, because it depends on recognising a similarity rather than on following a procedure — which is why the book presents it through a worked example rather than as a set of steps.
Real world
It is one of the most general problem-solving moves there is.
Discussion prompt
Describe a time you solved a problem by recognising it as one you had already solved — in any field. What made the recognition possible, and what would have happened without it?
Hint: The recognition usually needs the problems to be stated in a common form.
Answer:
The common answer is that somebody restated the problem in different words and it became familiar — which is exactly what writing both questions as is every letter of X in Y did here.
What makes recognition possible is having the earlier solution stated generally enough. A solution written for one specific case cannot be recognised in another, which is why generalization from lesson 4b makes reduction possible later.
Without the recognition you solve it again, usually differently, and now there are two implementations that can drift apart. That is the maintenance cost, and it is why reduction is worth a name of its own.
Section
Section 4
Concept
The chapter sets six exercises, and their common structure is the point. Each asks whether a word has some property, and each is answered by the same search pattern with a different condition.
Five of these six are the search pattern with a different condition. The sixth is not, and that is why it gets a section to itself in the next lesson — comparing adjacent letters needs something the others do not.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 84-84
Picture it
The odd one out is the one that needs to look at two letters at once.
Figure (svg): Two columns separating the five exercises solved by the search pattern from the one that is not
Recognising which exercises are instances of a pattern and which are not is itself the skill this chapter is teaching — and it is what the reduction move depends on.
Worked example
Two parts: a search function, and a count over the whole file.
def has_no_e(word):
for letter in word:
if letter == 'e':
return False
return True
total = 0
without_e = 0
for line in open('words.txt'):
word = line.strip()
total = total + 1
if has_no_e(word):
without_e = without_e + 1| Part | Which pattern | Note |
|---|---|---|
| has_no_e | the search pattern | one word at a time |
| the loop | the counter pattern | over the whole file |
| two counters | total, and those without an e | for the percentage |
Write the search function first, and test it alone.
Why: Lesson 6a's incremental development: get one piece working before building on it.
Wrap it in a counting loop.
Why: Lesson 8b's counter pattern, over the file rather than over a string.
Keep two counters.
Why: A percentage needs both a numerator and a denominator, and counting the total is easier than assuming you know it.
Figure (svg): The state of the program after each line of Worked example exercise 9.2, and the percentage, drawn as a ladder with one rung per traced line
A search function inside a counter loop. The two patterns compose directly: the counter's condition is a call to the search.
Verify: Sanity-check the percentage against expectation.
Why: The book notes that e is the most common letter in English, so the fraction without an e should be a minority — around a third. A result of 90 percent or 2 percent would signal a bug, and having a rough expectation before running is what makes the check possible at all.
Discrimination
Ask whether a single letter can settle the question.
Sort into buckets
For each exercise, is the plain search pattern enough?
Worked example
The last sentence of the exercise is a different kind of question.
# Write a program that prompts the user for forbidden letters
# and prints how many words avoid all of them.
#
# Can you find a combination of 5 forbidden letters that
# excludes the smallest number of words?| Part | What kind of question | Note |
|---|---|---|
| the first part | a program, with a definite answer | input, count, print |
| the second part | a search over combinations | no single right answer |
| what it teaches | using a program as a tool | rather than as the answer |
Write the program.
Why: Prompt for the letters, traverse the file, count the words that avoid them all. This is the composition from the previous example with input added.
Notice what the second question asks.
Why: It has no closed-form answer. You are meant to run the program repeatedly with different letters and compare.
See what that teaches.
Why: The program is a tool for exploring rather than a computation that produces the answer — which is a genuinely different way to use one.
Figure (svg): The state of the program after each line of Worked example exercise 9.3, and the open-ended part, drawn as a ladder with one rung per traced line
A counting program, then an experimental search using it. The rare letters — j, q, x, z and one more — exclude the fewest words, and finding them means running the program rather than reasoning.
Verify: Predict which letters exclude fewest words before testing.
Why: The rarest letters in English are the obvious candidates, and testing confirms it. Having a prediction before running turns the experiment into a test of your model rather than a lookup — which is lesson 6a's advice about knowing the right answer, applied to an open question.
Trap
The solutions are in the very next section, and reading straight on is effortless.
Treat the exercises as examples of what follows
Why: The layout does not stop you, and the solutions are short.
The book is explicit: there are solutions to these exercises in the next section, and you should at least attempt each one before you read them.
Attempt each one, even badly, before reading the solution.
Write something that runs, however clumsy
Why: The solutions are worth far more once you have a version to compare against.
Then read for the PATTERN, not the code
Why: The book's has_no_e is not much better than yours. What it demonstrates is that all four functions are the same shape.
This matters more here than in chapter 4, because the recognition being taught — that these are all one pattern — only lands if you have written several of them and felt the repetition.
Prediction
Check every letter of the word against the forbidden set.
>>> avoids('banana', 'xyz')| Letters checked | Result | What happens |
|---|---|---|
| b, a, n, a, n, a | none is x, y or z | no disqualification |
| the loop finishes | no early return | reach the end |
| return True | after the loop | the word avoids them all |
Predict first
What does avoids('banana', 'xyz') return?
Correct: True — none of the word's letters is in the forbidden string, so the loop completes and the final return runs.
Why: Every letter is checked against the forbidden set and none matches, so no early return happens and control reaches the return True after the loop. Note that a word can only earn True by being fully examined, which is the universal-claim asymmetry — whereas avoids('banana', 'a') would return False after just two letters.
Faded example
A search inside a counter.
Fill in the blanks
count = 0
for line in open('words.txt'):
word = line.strip()
if has_no_e(word):
count = count + 1
print(count)
Why: The counter's condition is a call to the search function, which is how the two patterns compose. Note the division of labour: has_no_e decides about one word and knows nothing about files, and the loop handles the file and knows nothing about e's. That separation is what makes each piece testable on its own, which is lesson 4b's interface argument in miniature.
Explain it
The recognition is the whole point of the section.
Discussion prompt
A classmate has written four of the exercises as four completely different functions and does not see the connection. Point out what they share, and say what noticing it buys.
Hint: Look at where the returns are, not at the conditions.
Answer:
Say: look at the structure rather than the condition. All four traverse something, return False from inside the loop, and return True after it. Only the condition differs.
What noticing it buys: the next one is quicker to write, because you start from the pattern and only have to decide the condition.
And it buys the reduction. Once you see that uses_only and uses_all are the same question with the arguments swapped, one of them stops needing a body at all — which you cannot see while they look like four unrelated functions.
Section
Section 5
Concept
A complete answer to one of these exercises composes three things you already have: a file traversal, a search function, and a counter — each of which knows nothing about the others.
def has_no_e(word):
for letter in word:
if letter == 'e':
return False
return True
count = 0
for line in open('words.txt'):
if has_no_e(line.strip()):
count = count + 1
print(count)| Component | Which pattern | What it knows about |
|---|---|---|
| has_no_e | the search pattern | decides about one word |
| the for loop | the file traversal | supplies the words |
| count | the counter pattern | accumulates the answer |
Each piece has one job. The search function knows nothing about files, and the loop knows nothing about the letter e — which is why either can be changed without touching the other.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 83-85
Picture it
The separation is what makes a 113,809-line input manageable.
Figure (svg): A three-stage diagram showing a file supplying words, a search function judging each, and a counter accumulating the result
Testing has_no_e on five words by hand is quick; testing it on 113,809 is not. The separation lets you do the first and trust the second.
Worked example
Lesson 6a's method, applied to a program with real data.
# stage 1: does the file open and read?
# print the first five words and stop
# stage 2: does has_no_e work?
# test it on 'banana' and 'bee' by hand
# stage 3: does the count work?
# run it on a short file first
# stage 4: run it on the real file| Stage | What is being tested | Note |
|---|---|---|
| stage 1 | the file layer alone | no search, no count |
| stage 2 | the search alone | no file |
| stage 3 | both, on small data | checkable by hand |
| stage 4 | the real run | nothing new is being tested |
Test each layer separately first.
Why: The file layer and the search function have nothing to do with each other, so a bug in one cannot be confused with a bug in the other.
Then combine on data you can check.
Why: A five-word file gives an answer you can verify by counting on your fingers.
Only then run on the real file.
Why: By that point nothing new is being tested — you are collecting an answer rather than debugging.
Figure (svg): The state of the program after each line of Worked example building it incrementally, drawn as a ladder with one rung per traced line
Four stages, three of which use no real data. The full run is the last step, not the first, which is what makes an unexpected number meaningful rather than mysterious.
Verify: Consider what happens if you skip to stage 4.
Why: A wrong count over 113,809 words gives no information about which layer is at fault, and no way to check the answer by hand. Working up from small verified pieces is the only way to have an expectation — which lesson 7b's bisection also depended on.
Prediction
An extreme result implicates a specific layer.
Predict first
A program counting words with no 'e' reports 0 out of 113,809. Where is the bug most likely to be?
Correct: Any of these — and printing the first few words distinguishes them immediately.
Why: All three layers could produce a zero, which is why the aggregate number alone is not a diagnosis. Printing the first few words tests the file layer; calling the search on a known word tests the second; and watching the counter on a short file tests the third. This is lesson 6c's three possibilities applied to a three-layer program, and the cheapest check comes first.
Worked example
Two plausible wrong results, and what each one implicates.
# expected: roughly a third of words have no 'e'
# result 0 -> the condition never succeeds
# result 113809 -> the condition always succeeds
# result ~37000 -> plausible| Result | What it means | Where to look |
|---|---|---|
| 0 | no word passed | the search always returns False |
| 113809 | every word passed | the search always returns True |
| a plausible fraction | some passed | worth checking further |
Form an expectation before running.
Why: e is the most common letter in English, so most words contain one — the answer should be a minority but not a tiny one.
Read an extreme result as a structural bug.
Why: Zero or the total means the condition is constant, which points at the search function rather than at the loop.
Read a plausible result as worth checking further.
Why: It is not proof. A count that is 5 percent off would look entirely reasonable.
Figure (svg): A panel listing three possible counts and what each implies about where the bug is
Extremes implicate the condition; a plausible number means you still have to check by other means. Having an expectation is what makes any of this possible.
Verify: Check a plausible answer by testing individual words.
Why: Running has_no_e on ten words chosen by hand — some with an e at the start, some at the end, some without — is a check the aggregate count cannot provide. That is the subject of the next lesson's debugging section, and Dijkstra's remark about testing applies directly.
Trap
A student writes the whole program and runs it on all 113,809 words immediately.
Test on the real data because that is the point
Why: The exercise is about the word list, so using it seems like the natural first step.
Any bug now produces a single wrong number with no way to attribute it, and no way to check the right answer by hand.
Work up to the real data through pieces you can verify.
Test the search function on a handful of words you chose
Why: One with the letter at the start, one at the end, one without, and the empty string.
Then run the loop on a short file
Why: Five words, whose answer you can count yourself.
This is incremental development from lesson 6a with one addition: the test case has to be one whose answer you know, and for a hundred thousand words you never will.
Ranking
Each stage should be verifiable before the next is added.
Put in order
Why: The file layer and the search function are independent, so either could come first — but both must precede the combination, and the combination must be checked on data small enough to verify by hand before the real run. The last step tests nothing new; it collects the answer. Reversing any two would mean debugging two things at once.
Comparison
Fill the blanks. Each layer knows nothing about the others.
Comparison matrix
| Layer | What it does | What it does NOT know |
|---|---|---|
| the file loop | supplies one clean word per pass | anything about the property being tested |
| has_no_e | judges one word | anything about files or counting |
| the counter | accumulates how many passed | why any particular word passed |
The right-hand column is what makes each layer testable alone. A layer that knew about the others could not be checked without them.
Real world
Supply, judge, accumulate is a shape you will meet constantly.
Discussion prompt
Describe a real process with the same three layers: something that supplies items, something that judges each one, and something that accumulates a result. What happens when the judging layer is entangled with the supplying one?
Hint: Quality control, marking, sorting post.
Answer:
Quality control on a production line is the standard example: the line supplies items, an inspection judges each, and a tally accumulates the reject rate.
If the inspection is entangled with the supply — an inspector who can only judge items from one particular machine — then you cannot test the inspection separately, and you cannot reuse it.
That is exactly the argument for keeping has_no_e ignorant of files. A judging step that knows only about the thing it judges can be tested on five hand-made examples, which is the only way anybody ever gains confidence in it.
Comparison
Fill the blanks. One structure, four conditions, and one of them traverses something different.
Comparison matrix
| Function | What it traverses | What disqualifies |
|---|---|---|
| has_no_e(word) | the word | a letter equal to 'e' |
| avoids(word, forbidden) | the word | a letter that is in forbidden |
| uses_only(word, available) | the word | a letter that is not in available |
| uses_all(word, required) | the REQUIRED letters | a required letter not in the word |
The bottom row is the one that traverses something else, and noticing that is what makes the reduction to uses_only visible.
Pattern
Six steps, and the first is the one that saves the most work.
Step 2 is the one this chapter adds. It is also the one that is easiest to skip, because writing the function feels like progress and looking for an existing solution does not.
Python documentation — Input and Output Input and Output
Check
readline gives you one extra character.
>>> line = fin.readline()
>>> line
'aahed\n'| Step | What happens | Result |
|---|---|---|
| readline | reads up to and including the newline | 'aahed\n' |
| strip | returns a new string without it | 'aahed' |
| assign | or the clean word is lost | word = line.strip() |
Check your understanding
What removes the trailing newline from a line read with readline?
Answer: A
Why: strip is a string method that returns a new string with whitespace removed from both ends, including the trailing newline. Like every string method it returns rather than modifies, so the result must be assigned — line.strip() on its own computes the clean word and discards it.
Check
One conclusion needs one example; the other needs all of them.
def has_no_e(word):
for letter in word:
if letter == 'e':
return False
return True| Statement | Where | Why |
|---|---|---|
| return False | inside the loop | one 'e' settles it |
| return True | after the loop | needs every letter checked |
| swapping them | would report success after one letter | wrong |
Check your understanding
Why can return True not go inside the loop?
Answer: B
Why: True is a claim about every letter, and a claim about every element can only be established by checking every element. False is a claim about one letter, so a single counterexample settles it and can be returned immediately. Putting the True inside would report success after examining the first letter alone.
Check
Two functions asking the same question with the arguments swapped.
Check your understanding
uses_only(word, available) asks whether every letter of the word is in available. What is uses_all(word, required) in terms of it?
Answer: B
Why: uses_all asks whether every letter of required appears in the word, which is the same shape with required playing the first role and the word playing the second. Swapping the arguments is the whole reduction, and it leaves uses_all with no logic of its own.
Real world
Reduction is the most valuable of the development plans and the least teachable.
Discussion prompt
Think of a field where recognising a problem as a familiar one is most of the expertise — medicine, law, chess, repair work. What does the expert have that a beginner does not, and how did they get it?
Hint: It is not more solutions; it is better recognition.
Answer:
What the expert has is a stock of problems stated generally enough to be recognised in new clothes. A beginner has solutions attached to the specific situations where they learned them.
They got it by seeing the same problem in several forms — which is exactly what this chapter does with four search functions that look different and are not.
So the practical advice is the same in both cases: after solving something, restate the problem in the most general terms you can. A solution filed under its general form can be recognised later, and one filed under its specific circumstances cannot.
Commit first
Answer, then rate your confidence. The reasoning matters more than the answer.
Predict first
A student writes avoids(word, forbidden) as return not uses_only(word, forbidden). Is this correct?
Correct: No — negating every letter is forbidden gives at least one letter is not forbidden, which is a far weaker claim than no letter is forbidden.
Why: The two functions have the same structure and ask questions that do not relate by negation. Test it: avoids('bee', 'e') should be False, since the word contains a forbidden letter. uses_only('bee', 'e') is also False, since 'b' is not in 'e', so the negation gives True — exactly wrong. This is the failure mode of reduction: matching on shape rather than on meaning. The defence is to write both questions as sentences in a common form before reducing, and then to test the reduction on one positive and one negative case.
Explain it
Reduction is worth explaining because it is the plan people never think to apply.
Discussion prompt
A classmate is about to write a fifth search function from scratch. Ask them the one question that might save them the work, and explain what makes the question worth asking every time.
Hint: The question is about a problem they have already solved.
Answer:
Ask: is this the same question as one you have already answered, possibly with the arguments in a different order or the sense reversed?
It is worth asking every time because the cost is ten seconds and the payoff is an entire function you do not write — and, more importantly, do not have to get right.
Add the caveat, because it is the failure mode: write both questions out as sentences before deciding they are the same. Two functions can share a structure and ask genuinely different things, and reducing one to the other then gives confident wrong answers with nothing to indicate it.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The file handling is deliberately minimal here and gets a full chapter later, so a rough grasp is enough for now — except for the strip, which causes real bugs. The search pattern's asymmetry is a general logical point worth over-learning. Swapping which sequence you traverse is the specific insight that makes uses_all work, and it is the step most people miss. And reduction is the hardest of the four and the most valuable — it is the one that separates programmers who write a lot of code from those who write a little.
Connect it up
One page, from memory.
Draw it
Write the four search functions as skeletons, side by side, with only the differing line filled in for each. Mark which one traverses something other than the word. Then draw an arrow between the two that reduce to each other and label it with the swap. Finally, at the bottom, write the three-layer structure of a program over the word file — supply, judge, accumulate — and note what each layer must not know about the others.
Recap
Three pages, and programs that work on real data rather than on literals you typed.
| If you remember one thing | It is this |
|---|---|
| From files | readline includes the newline. Strip at the boundary, before anything else. |
| From the search | A universal claim needs the whole loop; one counterexample ends it. |
| From the four functions | One pattern, four conditions. The structure is the reusable part. |
| From reduction | Before writing a function, ask whether you have already written it. |
| From the layers | The judging step should know nothing about where its input came from. |
The next lesson finishes the chapter with the puzzles a plain for loop cannot solve — comparing adjacent letters, and comparing a word with itself from both ends — and with what it means to test a program you cannot check by hand.
Think Python, 2nd edition — Allen B. Downey §9.1-9.3, pp. 83-85 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.