This lesson introduces string methods and the dot notation that invokes them, meets optional arguments through the built-in find, adds the in operator, explains why uppercase letters sort before lowercase ones, and works a live diagnosis of an index traversal with two bugs in it.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 8 — Strings
§8.8-8.11, pp. 75-78
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-78 — the pages these objectives are drawn from
Warm-up
The dot arrives again, on a third kind of thing.
Discussion prompt
You have used a dot in two situations already: with a module in chapter 3, and with a turtle in chapter 4. Write down both, and say what the general rule was.
Hint: Lesson 3a gave a single reading that covers both.
Answer:
math.sqrt(2) reaches into a module, and bob.fd(100) asks a turtle to do something.
The rule was: look inside the thing on the left for the name on the right. It covered both without needing to know what an object is.
This lesson applies it to strings, which is where it becomes an everyday tool — because unlike turtles, you use strings in every program you write.
Concept
Strings provide methods that perform a variety of useful operations. A method is similar to a function — it takes arguments and returns a value — but the syntax is different: instead of upper(word), it is word.upper().
invocation — A method call. We say that we invoke a method on an object.
This form of dot notation specifies the name of the method and the name of the string to apply it to. The empty parentheses indicate that the method takes no arguments — and they are required, exactly as they were for a function.
Figure (svg): Two columns contrasting function syntax with method syntax for the same operation
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-76
Section
Section 1
Concept
A method is similar to a function: it takes arguments and returns a value. What differs is the syntax — the string the method applies to goes before the dot rather than inside the parentheses.
>>> word = 'banana'
>>> new_word = word.upper()
>>> new_word
'BANANA'
>>> word
'banana'| Expression | What happens | Result |
|---|---|---|
| word.upper() | invoke upper on word | returns a NEW string |
| new_word | the returned value | 'BANANA' |
| word | unchanged | strings are immutable |
A method call is called an invocation; in this case, we would say we are invoking upper on word. The method takes a string and returns a new string with all uppercase letters — and, because strings are immutable, the original is untouched.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-76
Picture it
Three kinds of thing on the left, one rule.
Figure (svg): A diagram showing the dot used on a module, on a turtle and on a string, with the same meaning each time
Chapter 15 explains what a string and a turtle have in common that a module does not. Until then, the reading covers every case you will meet.
Worked example
Immutability from the previous lesson, meeting method syntax.
>>> word = 'banana'
>>> word.upper()
'BANANA'
>>> word
'banana'
>>> word = word.upper()
>>> word
'BANANA'| Line | What happens | State |
|---|---|---|
| word.upper() | returns a new string | the original is unchanged |
| word alone | still lowercase | nothing was modified |
| word = word.upper() | the result is assigned | now word refers to the new string |
Invoke the method and look at the result.
Why: It returns a new string with all uppercase letters. At the prompt, the returned value is displayed.
Check the original.
Why: Unchanged. Strings are immutable, so no method can modify one — every string method returns a new string.
Assign the result if you want to keep it.
Why: word = word.upper() repoints the name at the new string, which is reassignment rather than modification.
Figure (svg): The state of the program after each line of Worked example a method returns rather than modifies, drawn as a ladder with one rung per traced line
The method returns a new uppercase string and leaves the original alone. Keeping the result requires an assignment, which is lesson 6a's fruitful-function rule applied to methods.
Verify: Call the method without assigning and check nothing changed.
Why: word.upper() on its own computes a new string and discards it — the same wasted-return-value mistake as calling math.sqrt on its own in lesson 6a. It produces no error, which is exactly why it is worth causing once deliberately.
Prediction
The method returns; it does not modify.
word = 'banana'
word.upper()
print(word)| Line | What happens | Output |
|---|---|---|
| word.upper() | computes 'BANANA' | and discards it |
| no assignment | the result goes nowhere | word is untouched |
| print(word) | the original | banana |
Predict first
What does this print?
Correct: banana — the method returned a new uppercase string, which was discarded, and the original is unchanged.
Why: Strings are immutable, so upper cannot modify anything; it produces a new string. Since nothing assigned that string, it is lost — the same wasted return value as calling math.sqrt on its own. The fix is word = word.upper(), and the absence of any error message is what makes this worth meeting deliberately.
Worked example
Even with no arguments, they are required. Work out what they do.
>>> word = 'banana'
>>> word.upper
<built-in method upper of str object at 0x...>
>>> word.upper()
'BANANA'| Expression | What it means | Result |
|---|---|---|
| word.upper | names the method | the method object itself |
| word.upper() | calls it | the uppercase string |
| the parentheses | the call operator | required even when empty |
Write the name without parentheses.
Why: This names the method without invoking it, exactly as math.sqrt without parentheses named the function in lesson 3a.
Add the parentheses.
Why: The empty parentheses indicate that this method takes no arguments — and they are what makes it a call.
Recall the general rule.
Why: The parentheses are the operator that means run this. Their emptiness says there is nothing to supply, not that they are optional.
Figure (svg): Two columns comparing naming a method with invoking it
Without parentheses you have named the method; with them you have invoked it. The rule is identical to the one for functions, which is one fewer thing to learn.
Verify: Compare with the same mistake on a function.
Why: math.sqrt without parentheses displays a function object, and word.upper without them displays a method object. Both are legal and neither computes anything — the same symptom, and the same fix.
Trap
A student writes word.upper() on a line by itself and then prints word, expecting it to be uppercase.
Read a method call as an instruction to the string
Why: The syntax reads like a command — word, uppercase yourself — which makes modification the natural expectation.
Strings are immutable, so nothing changed. The new string was created and immediately discarded, with no error to indicate it.
Every string method returns a new string and modifies nothing.
Assign the result
Why: word = word.upper(), which repoints the name at the new string.
Treat a string method call as an expression
Why: Its value is the point, exactly as with any fruitful function. A call whose result is discarded did nothing.
This is one of the two commonest string mistakes, and the other is the same mistake on the other side: assigning the result of a method that returns None. Both come from being unsure whether an operation modifies or returns.
Matching
Four expressions built from one method name.
Match the pairs
Why: The pattern is entirely mechanical, and it is the same as for functions in lesson 3b: parentheses mean invoke, and what happens to the result depends on what surrounds the call. Note that only the third one changes what word refers to, and it does so by reassignment rather than by modifying any string.
Comparison
Fill the blanks. The ideas are identical; only the arrangement differs.
Comparison matrix
| Question | Function | Method |
|---|---|---|
| How is it written? | len(word) | word.upper() |
| Where does the string go? | inside the parentheses, as an argument | before the dot |
| Does it return a value? | yes | yes — the same as a function |
| Are the parentheses required? | yes, even with no arguments | yes, even with no arguments |
Three of the four rows are identical. The one real difference is where the value being operated on is written, which is why the book says a method is similar to a function but the syntax is different.
Socratic
len(word) and word.upper() do comparable things in different shapes.
Discussion prompt
Python has both len(word) and word.upper(). Suggest one reason a language might attach some operations to the value and leave others as free functions.
Hint: Which operations make sense for many different types?
Answer:
len works on strings, lists, dictionaries and more, so it belongs to no single type — making it a free function lets one name cover all of them.
upper only makes sense for strings, so attaching it to strings keeps it out of the way and makes it discoverable: everything a string can do is reachable through the dot.
There is a practical benefit too: two different types can each have a method called find with behaviour appropriate to that type, without colliding. Chapter 17 shows how that works, and it is the reason methods exist at all.
Section
Section 2
Concept
There is a string method named find that is remarkably similar to the function written in the previous lesson — and it is more general in three ways.
>>> word = 'banana'
>>> word.find('a')
1
>>> word.find('na')
2
>>> word.find('na', 3)
4
>>> 'bob'.find('b', 1, 2)
-1| Call | What it does | Result |
|---|---|---|
| find('a') | a single character | 1 |
| find('na') | a SUBSTRING, not just a character | 2 |
| find('na', 3) | start looking at index 3 | 4 |
| find('b', 1, 2) | search only indices 1 up to but not including 2 | -1 |
The method is more general than the function we wrote: it can find substrings, not just characters. By default it starts at the beginning, but it can take a second argument, the index where it should start — and a third, the index where it should stop. This is an example of optional arguments.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 76-76
Picture it
Each addition is a capability the hand-written version lacks.
Figure (svg): Two columns comparing the hand-written find function with the built-in string method
Writing the simpler version first was not wasted: it is why you know what the built-in is doing, and why -1 means what it means.
Worked example
The same method, called three ways. Each adds one argument.
>>> word = 'banana'
>>> word.find('na')
2
>>> word.find('na', 3)
4
>>> word.find('na', 3, 4)
-1| Form | What it searches | Result |
|---|---|---|
| one argument | search the whole string from 0 | finds the first 'na' at 2 |
| two arguments | start at index 3 | finds the second 'na' at 4 |
| three arguments | search indices 3 up to but not including 4 | too narrow: -1 |
Call it with one argument.
Why: By default find starts at the beginning of the string, so it reports the first occurrence.
Add the start index.
Why: It can take a second argument, the index where it should start — which is how you find the second occurrence after finding the first.
Add the stop index.
Why: A third argument gives the index where it should stop, and searching up to but not including that index makes find consistent with the slice operator.
Figure (svg): The string banana with the two occurrences of na highlighted and the search region from index three marked
2, then 4, then -1. Each optional argument narrows the region searched, and the third one narrows it so far that nothing is found.
Verify: Check the consistency claim against a slice.
Why: word.find('na', 3, 5) searches the same region as word[3:5], which is 'an' — and 'na' is not in it, so both agree there is nothing to find. Searching up to but not including the second index is exactly the half-open convention from lesson 8b, applied here so the two operations agree.
Prediction
The search starts partway along.
>>> 'banana'.find('a', 2)| Step | What happens | Result |
|---|---|---|
| start at index 2 | skip the 'a' at index 1 | search from 'nana' |
| first 'a' from there | at index 3 | 3 |
| the result | an absolute index | not relative to the start |
Predict first
What does 'banana'.find('a', 2) return?
Correct: 3 — the first 'a' at or after index 2, reported as an index into the whole string.
Why: The start argument says where to begin looking, not where to count from: the returned index is still an offset into the original string. That distinction matters — if the result were relative to the start, chaining searches would need arithmetic on every call, and the enumerate-all-occurrences loop would be much harder to write.
Worked example
The start argument exists for exactly this. Use it in a loop.
word = 'banana'
index = word.find('a')
while index != -1:
print(index)
index = word.find('a', index + 1)| Call | What it searches | Result |
|---|---|---|
| find('a') | the first occurrence | 1 |
| find('a', 2) | search from just after it | 3 |
| find('a', 4) | again | 5 |
| find('a', 6) | nothing left | -1, so the loop ends |
Find the first occurrence.
Why: An ordinary call, giving index 1.
Search again from just past it.
Why: index + 1 rather than index, otherwise the same occurrence would be found forever — which is the classic infinite loop in this pattern.
Stop when the sentinel comes back.
Why: -1 means nothing further was found, so the loop condition tests for it.
Figure (svg): The state of the program after each line of Worked example finding every occurrence, drawn as a ladder with one rung per traced line
1, 3 and 5 — every position of an 'a'. The start argument turns a single search into a way of enumerating all occurrences.
Verify: Try it with index rather than index + 1 and predict what happens.
Why: It finds the same occurrence forever, because searching from index finds the thing at index. That is an infinite loop caused by an update that does not move — lesson 7a's requirement that the body must change what the condition depends on, in a form specific to searching.
Trap
A student writes if word.find('a'): to test whether the letter is present.
Treat a search result as a yes-or-no answer
Why: The question being asked really is yes-or-no, so a truthy test feels natural.
find returns a position. When the letter is at index 0 the result is 0, which counts as false — so the test reports not found for exactly the case where it was found first.
Compare the result with -1, or use the in operator.
Test explicitly against the sentinel
Why: if word.find('a') != -1: says what it means and cannot be confused by a zero.
Or use in, when the position does not matter
Why: if 'a' in word: is shorter and returns a genuine boolean.
This is lesson 5a's truthiness warning in its most damaging form: a valid answer of 0 counting as false. The book's advice to avoid relying on truthiness unless you know what you are doing is exactly about cases like this.
Discrimination
The optional arguments narrow the search region.
Sort into buckets
For word = 'banana', which of these calls return a position rather than -1?
Faded example
Use the position of the first as the starting point for the second.
Fill in the blanks
word = 'banana'
first = word.find('a')
second = word.find('a', first + 1)
Why: Searching from first + 1 skips past the occurrence already found. Searching from first itself would find the same one again, which is the classic infinite loop when this pattern goes in a while loop. The plus one is the update that moves the search forward, and it is the same requirement as lesson 7a's rule about the body changing what the condition depends on.
Explain it to yourself
The built-in is strictly better. Say what the exercise bought.
Discussion prompt
Python already had a find method when you wrote your own in the previous lesson. Name two things writing it gave you that using the built-in would not have.
Hint: One is about the pattern; one is about the -1.
Answer:
It taught the search pattern: traverse, return early on a match, return the sentinel after the loop. That pattern applies to lists, files and anything else you traverse, and no amount of calling the built-in would have shown it.
It also explains the -1. Having written the not-found return yourself, you know why the sentinel is negative and why it must be checked rather than used directly.
The general point is worth keeping: reimplementing something that already exists is a way of understanding it, not a waste. The book does it deliberately here, and again in chapter 11 with dictionaries.
Section
Section 3
Concept
The word in is a boolean operator that takes two strings and returns True if the first appears as a substring in the second.
>>> 'a' in 'banana'
True
>>> 'seed' in 'banana'
False
>>> 'nan' in 'banana'
True| Expression | What it asks | Result |
|---|---|---|
| 'a' in 'banana' | a single character | True |
| 'seed' in 'banana' | not present | False |
| 'nan' in 'banana' | a substring of any length | True |
With well-chosen variable names, Python sometimes reads like English. The book's example is a loop you can read as: for each letter in the first word, if the letter appears in the second word, print the letter.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 76-77
Picture it
Same question, two answers, and one of them is a genuine boolean.
Figure (svg): Two columns comparing the in operator with the find method for testing presence
Choosing between them is the same question as lesson 8b's search-versus-counter: what does the answer need to be? If a position is never used, in says so more clearly.
Worked example
The book's example, and it reads almost as English.
def in_both(word1, word2):
for letter in word1:
if letter in word2:
print(letter)| Letter from word1 | Is it in word2? | Output |
|---|---|---|
| 'a' from apples | 'a' in 'oranges'? yes | print a |
| 'p' | not in oranges | skip |
| 'e' | yes | print e |
| 's' | yes | print s |
Read the loop aloud.
Why: For each letter in the first word, if the letter appears in the second word, print the letter. The code and the sentence are nearly identical.
Notice the two different uses of the word in.
Why: The first is the for loop's keyword, meaning iterate over; the second is the boolean operator, meaning appears within. Same word, different jobs.
Trace the book's example.
Why: in_both('apples', 'oranges') prints a, e and s — the letters the two words have in common, in the order they appear in the first.
Figure (svg): A diagram showing each letter of the first word being tested for membership in the second and printed if present
a, e and s. The function traverses one word and tests each letter against the other, which is the counter pattern with printing instead of counting.
Verify: Check what happens with a repeated letter.
Why: in_both('aaa', 'ab') prints a three times, because the traversal visits each character and does not track which have already been reported. That is worth noticing: the function reports occurrences rather than distinct letters, which may or may not be what was wanted.
Prediction
Contiguous, and in order.
Predict first
What is 'ana' in 'banana'?
Correct: True — 'ana' appears starting at index 1, and again at index 3.
Why: The characters a, n, a appear consecutively and in that order, so the operator returns True. Note that it returns a boolean rather than a position, unlike find — so it can be used directly in a condition without any comparison. If you answered 1, you were thinking of find, which is the operator that reports where.
Worked example
It is not limited to single characters, which is easy to miss.
>>> 'nan' in 'banana'
True
>>> 'ban' in 'banana'
True
>>> 'nab' in 'banana'
False| Substring | Why | Result |
|---|---|---|
| 'nan' | appears at index 2 | True |
| 'ban' | appears at index 0 | True |
| 'nab' | those letters are all present, in that ORDER they are not | False |
Test a multi-character substring.
Why: in returns True if the first appears as a substring in the second — the characters must be contiguous and in order.
Notice the third case.
Why: 'nab' contains only letters that are present in 'banana', and it is still False, because they do not appear together in that order.
Connect to lesson 2b.
Why: This is the ordering point again: a string is a sequence, so containment is about consecutive positions rather than about which characters exist.
Figure (svg): The state of the program after each line of Worked example in on a substring rather than a character, drawn as a ladder with one rung per traced line
True, True, False. The operator tests for a contiguous substring, not for the presence of the individual characters.
Verify: Test a case where every character is present but not adjacent.
Why: 'bn' in 'banana' is False, even though both letters appear. Confirming that containment requires adjacency is worth doing once, because contains these letters and contains this substring are easy to conflate.
Trap
A student uses 'nab' in 'banana' to test whether the word is an anagram of nab, or contains those letters.
Read containment as being about the set of characters
Why: All three letters really are in banana, so the reading is not unreasonable.
The operator tests for a contiguous substring in order, so it returns False — and a test written this way silently answers a different question from the one intended.
in tests for a substring: contiguous, and in order.
Say it precisely: does this exact run of characters appear?
Why: Not are these characters present, which is a different question.
For the other question, test each character separately
Why: Which is what in_both does — one in test per character, in a loop.
The distinction matters because both questions are common and the syntax for one of them looks like it should answer the other. Being explicit about which you mean is the whole defence.
Discrimination
Ask whether the position is used.
Sort into buckets
For each task, which is the better tool?
Faded example
One operator, no comparison needed.
Fill in the blanks
filename = 'report.txt'
if '.txt' in filename:
print('a text file')
Why: in returns True or False directly, so it can be the whole condition with no comparison attached. Writing filename.find('.txt') != -1 would work identically and is longer; writing filename.find('.txt') alone would be a bug, since a match at index 0 gives 0, which counts as false. That last case is why in is preferred whenever the position is not needed.
Explain it
A choice beginners face constantly.
Discussion prompt
A classmate uses find for everything, including simple presence tests. Give them the rule, and one concrete bug that in would have prevented.
Hint: The bug involves a match at position zero.
Answer:
Say: use in when you only need to know whether something is there, and find when you need to know where.
The bug: if word.find('b') is a truthiness test, and 'b' is the first character, find returns 0 — which counts as false. So the test reports not found precisely when the target is at the front.
in cannot have that bug, because it returns a real boolean. That is a stronger argument than readability: choosing in eliminates an entire failure mode rather than merely looking nicer.
Section
Section 4
Concept
The relational operators work on strings, both for equality and for ordering — which is how words are put into alphabetical order.
if word == 'banana':
print('All right, bananas.')
if word < 'banana':
print('Your word comes before banana.')
elif word > 'banana':
print('Your word comes after banana.')
else:
print('All right, bananas.')| Operator | What it asks | Note |
|---|---|---|
| == | exact equality, character for character | True or False |
| < | comes before, alphabetically | for lowercase words |
| the chain | three cases, one of which always applies | lesson 5b's shape |
Python does not handle uppercase and lowercase letters the same way people do. All the uppercase letters come before all the lowercase letters, so the word Pineapple comes before banana — which is almost certainly not what a person sorting words would say.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 77-77
Picture it
Not A before a, then B before b. Every uppercase letter comes before every lowercase one.
Figure (svg): A row showing capital A through Z followed by lowercase a through z in comparison order
So Pineapple comes before banana, and Zebra comes before apple. The rule is consistent and it is not alphabetical order as people mean it.
Worked example
Two comparisons that a person would answer differently.
>>> 'Pineapple' < 'banana'
True
>>> 'Zebra' < 'apple'
True
>>> 'apple' < 'banana'
True| Comparison | Why | Result |
|---|---|---|
| 'Pineapple' vs 'banana' | P is uppercase, b is lowercase | all uppercase comes first |
| 'Zebra' vs 'apple' | Z is uppercase, a is lowercase | same reason |
| 'apple' vs 'banana' | both lowercase | ordinary alphabetical order |
Compare the first characters.
Why: String comparison works character by character from the left, and the first difference decides the answer.
Apply the case rule.
Why: All the uppercase letters come before all the lowercase letters, so any word starting with a capital comes before any word starting with a small letter.
Notice the third case is unsurprising.
Why: Within a single case, the ordering is exactly what you would expect. The surprise only arises when the cases are mixed.
Figure (svg): The state of the program after each line of Worked example the case surprise, drawn as a ladder with one rung per traced line
All three are True, and only the first two are surprising. Mixed-case comparison is not the alphabetical order people mean, and it is entirely consistent.
Verify: Sort a mixed-case list mentally and check the result.
Why: banana, Apple and cherry would sort as Apple, banana, cherry — which looks right by luck, since A is the earliest letter anyway. Trying Zebra, apple, Banana gives Banana, Zebra, apple, which is visibly not alphabetical. Finding an example where the difference shows is what makes the rule stick.
Prediction
The first differing character decides, and case matters.
Predict first
Is 'Zebra' < 'apple' True or False?
Correct: True — Z is an uppercase letter and a is lowercase, and all uppercase letters come before all lowercase letters.
Why: The comparison looks at the first characters, Z and a, and the case rule decides: every uppercase letter sorts before every lowercase one. This is consistent rather than random, and it is not the ordering a person means by alphabetical. Converting both with lower() first gives False, which is the human answer.
Worked example
Convert to one case before comparing.
>>> 'Pineapple'.lower() < 'banana'.lower()
False
>>> 'Pineapple'.lower()
'pineapple'| Step | What happens | Result |
|---|---|---|
| lower() | returns a new lowercase string | 'pineapple' |
| the comparison | both now lowercase | p comes after b |
| the result | False | which is what a person would say |
Convert both to a common case.
Why: A common way to address this problem is to convert strings to a standard format, such as all lowercase, before performing the comparison.
Compare the converted values.
Why: Now every letter is in the same case, so the ordering is ordinary alphabetical order.
Note what has not changed.
Why: lower returns a new string, so the originals are untouched — immutability again, and the conversion is for the comparison only.
Figure (svg): Two columns comparing a mixed-case comparison with the same comparison after converting both to lowercase
False, which is the answer a person sorting words would give. Normalising the case before comparing is the standard technique, and it costs one method call on each side.
Verify: Check that the conversion does not affect the stored values.
Why: The original strings are unchanged, because lower returns a new string. That is what makes this safe to do inside a comparison — you are normalising for the purpose of the test, not editing your data.
Trap
A program sorts a list of names typed by users and produces an order that looks scrambled.
Assume string comparison means alphabetical order
Why: It does, within a single case, and most test data is consistently capitalised.
Real input is not consistent. A single name typed in lowercase sorts after every capitalised one, which looks like a bug in the sorting rather than in the comparison.
Normalise before comparing whenever the data comes from people.
Convert to a standard format, such as all lowercase
Why: Which is the book's own advice, and one method call.
Decide whether to normalise for comparison only, or to store the normalised form
Why: Usually the first: display what the user typed, compare on the normalised version.
The book's closing joke about defending yourself against a man armed with a Pineapple is memorable for a reason — the surprise is specific enough that having a name for it makes it easy to spot.
Ranking
Uppercase first, then lowercase, alphabetically within each.
Put in order
Why: Both capitalised words come before both lowercase ones, and within each group ordinary alphabetical order applies. So Apple, Banana, apple, banana — which interleaves nothing and looks wrong to a person expecting Apple, apple, Banana, banana. That gap between the two orderings is exactly what case normalisation closes.
Faded example
One method call on each side.
Fill in the blanks
if word.lower() == 'banana':
print('All right, bananas.')
Why: Converting the input to lowercase before comparing means Banana, BANANA and banana all match. Note that the literal on the right is already lowercase, so only one conversion is needed — but converting both sides is the safer habit, since the literal might change later. The method returns a new string, so word itself is unchanged.
Counterexample
It fixes the case surprise. Find a case it does not fix.
Discussion prompt
Converting both strings to lowercase makes comparison match human expectations for English words. Give an example where it still does not, and say what that suggests.
Hint: Think about digits, spaces, or accented letters.
Answer:
Digits and punctuation sort before letters, so '2nd' comes before 'first' even after lowercasing — which is not how a person would order a list.
Accented letters sort after all unaccented ones, so a name beginning with an accented character lands at the end rather than beside its unaccented neighbour.
What that suggests is that character comparison is not the same thing as alphabetical order, and lowercasing closes only the most common gap. Real sorting of human-readable text needs rules that the comparison operators do not implement — which is worth knowing exists rather than solving here.
Section
Section 5
Concept
When you use indices to traverse the values in a sequence, it is tricky to get the beginning and end of the traversal right. The book's is_reverse function is supposed to return True when one word is the reverse of the other, and it contains two errors.
def is_reverse(word1, word2):
if len(word1) != len(word2):
return False
i = 0
j = len(word2)
while j > 0:
if word1[i] != word2[j]:
return False
i = i+1
j = j-1
return True| Part | Its role | Note |
|---|---|---|
| the guardian | different lengths cannot be reverses | lesson 6c's pattern |
| i | traverses word1 forward | from 0 |
| j | traverses word2 backward | from len(word2) |
| the bug | j starts one too high | IndexError |
The first if statement checks whether the words are the same length; if not, we can return False immediately. This is an example of the guardian pattern — for the rest of the function we can assume the words are the same length.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 77-78
Picture it
i walks forward through one word while j walks backward through the other.
Figure (svg): Two columns showing i moving forward through pots and j moving backward through stop
The right-hand column starts at 3, not 4. The bug is that the code starts j at len(word2), which is one past the end.
Worked example
The book's first move, and it is the whole technique.
while j > 0:
print(i, j)
if word1[i] != word2[j]:
return False
i = i+1
j = j-1| Stage | What you do | What you learn |
|---|---|---|
| the symptom | IndexError on the comparison line | which index? |
| the print | immediately before the failing line | shows both |
| the output | 0 4 | j is 4, and 'pots' has no index 4 |
Locate the failing line from the traceback.
Why: It names the comparison, which uses both indices — so the traceback alone cannot say which one is wrong.
Print the values immediately before it.
Why: For debugging this kind of error, the book's first move is to print the values of the indices immediately before the line where the error appears.
Read the output.
Why: The first time through the loop, the value of j is 4, which is out of range for a four-character string. The index of the last character is 3, so the initial value for j should be len(word2)-1.
Figure (svg): A traceback showing an IndexError on the comparison line, with the printed index values above it
j starts at 4 when the largest valid index is 3. One print statement identified which of the two indices was wrong and by how much.
Verify: Check that the fix makes the first pass valid.
Why: With j starting at len(word2)-1, the first pass has i = 0 and j = 3, which are both valid. This is lesson 8a's off-by-one at the end of a string — len is one past the last index — appearing in a place where two indices made it harder to see.
Prediction
The indices are printed before the failing comparison.
i = 0
j = len(word2)
while j > 0:
print(i, j)
if word1[i] != word2[j]:| Variable | Its initial value | Printed |
|---|---|---|
| i | initialized to 0 | 0 |
| j | initialized to len(word2) | 4 for a four-letter word |
| the print | runs before the comparison | 0 4 |
Predict first
For is_reverse('pots', 'stop'), what does the print statement show on the first pass?
Correct: 0 4 — i starts at 0 and j starts at len(word2), which is 4.
Why: That single line is the diagnosis: 4 is out of range for a four-character string, whose valid indices are 0 to 3. The traceback pointed at the comparison line, which uses both indices, so it could not say which one was wrong — and one print statement settles it immediately. This is why the book's first move is to print the indices rather than to reread the code.
Worked example
The fix works and the loop count is wrong. The book leaves this as an exercise.
>>> is_reverse('pots', 'stop')
0 3
1 2
2 1
True| Observation | What it is | Why it matters |
|---|---|---|
| the answer | True | correct for this input |
| the loop count | three passes | for four-character words |
| the suspicion | one pair was never compared | the condition stops early |
Notice the right answer.
Why: This time we get the right answer — 'pots' and 'stop' really are reverses.
Notice the loop count.
Why: It looks like the loop only ran three times, which is suspicious for words of four characters.
Draw the conclusion without fixing it.
Why: One pair of characters was never compared, so there are inputs for which this function would return True wrongly. The book invites you to draw a state diagram, run the program on paper, and find the second error.
Figure (svg): The state of the program after each line of Worked example the second bug, and why it is suspicious, drawn as a ladder with one rung per traced line
The loop runs three times for four-character words, so one comparison is skipped. The function gives the right answer here and would not for every input — which is exactly the kind of bug a correct-looking test hides.
Verify: Look for an input where the missing comparison would matter.
Why: The skipped pair is the one where j reaches 0, so a difference in that single position would go unnoticed. Constructing such an input is the exercise — and being able to say WHERE the gap is, from the loop count alone, is most of the work.
Trap
After fixing the IndexError, the function returns True for pots and stop, so it is declared working.
Treat a correct output as proof
Why: The test passed, and the previous version crashed, so progress is obvious.
The loop ran three times for four-character words. The right answer arrived despite a comparison being skipped, which means some other input will get the wrong answer.
Check the process, not only the answer.
Count the passes and compare with what you expect
Why: Four characters should mean four comparisons. Three is a discrepancy that demands an explanation.
Draw the state diagram and run it on paper
Why: The book's own advice: start with the diagram, changing i and j at each iteration, and find the second error.
This is the deepest debugging lesson in the chapter. A correct output from an incorrect process is the most dangerous state a program can be in, because every test that passes makes it more entrenched.
Error analysis
One has been fixed. Mark both, and say why only one produced an error.
Annotate
Fixing the loud bug revealed the quiet one, but only to somebody who counted the passes. That counting step is the whole technique.
Invariant
Step through and watch which pairs are compared.
Step through it
Which pair of characters is never compared, and what condition would include it?
The pair at i = 3 and j = 0. The condition needs to remain true when j is 0, so j >= 0 rather than j > 0 — which is exactly the boundary question from lesson 8a's backward traversal.
Explain it
The technique is one line long and it is worth passing on.
Discussion prompt
A classmate has an IndexError on a line that uses two indices, and cannot tell which one is wrong. Tell them what to do, and why the traceback alone was not enough.
Hint: The traceback names the line, not the value.
Answer:
Say: put a print of both indices on the line immediately before the failing one, and run it again. The first line of output tells you which index is out of range and by how much.
The traceback could not tell them because it names the line and never the values — which is lesson 3c's point that a traceback locates a failure without explaining it.
Then add the follow-up that catches the second bug: once it works, count the passes and check the count against the length. A right answer from too few passes means a comparison was skipped, and that is a bug waiting for a different input.
Comparison
Fill the blanks. They answer different questions and are easy to confuse.
Comparison matrix
| Expression | What it returns | Use it when |
|---|---|---|
| 'a' in word | True or False | you only need to know whether it is there |
| word.find('a') | a position, or -1 | you need to know where |
| word.find('a') != -1 | True or False | you have a position already and want a boolean too |
| word.find('a') used as a condition | a position, treated as truthy | never — a match at index 0 counts as false |
The last row is a bug rather than a technique. It is the clearest case of lesson 5a's truthiness warning, and in is the fix.
Pattern
Six steps, and the last two are what catch the silent bug.
Steps 5 and 6 are what separates this from ordinary debugging. Two of the three traversal failures are silent, so a correct output is not evidence that the traversal is correct.
Python documentation — Built-in Types Built-in Types
Check
The method returns a new string.
word = 'banana'
word.upper()
print(word)| Line | What happens | Output |
|---|---|---|
| word.upper() | returns 'BANANA' | and discards it |
| word | unchanged | strings are immutable |
| print(word) | the original | banana |
Check your understanding
What does this print?
Answer: B
Why: Strings are immutable, so upper cannot modify anything — it returns a new uppercase string, which is discarded because nothing assigns it. The original is untouched. Keeping the result requires word = word.upper(), and the absence of any error is what makes this mistake easy to miss.
Check
One returns a position and one returns a boolean.
>>> word = 'banana'
>>> word.find('b')
0| Expression | What it gives | Note |
|---|---|---|
| find('b') | 'b' is at index 0 | 0 |
| 0 as a condition | zero counts as false | the trap |
| 'b' in word | a genuine boolean | True |
Check your understanding
Why is if word.find('b'): a bug?
Answer: B
Why: find returns the position, and a match at the very start gives 0 — which is treated as false in a condition. So the test reports not found precisely when the character is at the front. The fixes are to compare explicitly against -1, or to use the in operator, which returns a real boolean and cannot have this bug.
Check
The first differing character decides, and case matters.
Check your understanding
Which of these is True in Python?
Answer: B
Why: All the uppercase letters come before all the lowercase ones, so any capitalised word sorts before any lowercase word regardless of the letters. Z before a is the clearest demonstration, and it is why comparing user input usually needs a lower() on both sides first.
Real world
Sorting that does not match human expectations is a familiar frustration.
Discussion prompt
Think of a list you have seen sorted in a way that looked wrong — files, contacts, songs. What ordering was the machine using, and what ordering did you expect?
Hint: Numbers in filenames are the classic case.
Answer:
File listings are the standard example: file10 sorts before file2, because the comparison is character by character and '1' comes before '2'.
Mixed-case contact lists show the same thing as the Pineapple example: everything capitalised sorts before everything lowercase, which looks scrambled.
The general lesson is that character comparison and human ordering are different things that usually agree. Where they disagree, somebody has to normalise — lowercasing, or padding numbers with zeros — and knowing that the machine's ordering is consistent rather than random is what makes the fix findable.
Commit first
Answer, then rate your confidence. This one has caught professionals.
Predict first
A program tests presence with if word.find(letter):. For which case does it give the wrong answer?
Correct: When the letter is at index 0 — find returns 0, which counts as false, so the test reports not found even though it was found first.
Why: This is lesson 5a's truthiness warning in its most damaging form. A perfectly valid answer, 0, is a falsy value, so the condition inverts the meaning for exactly one case — and that case is the very first position, which real data hits constantly. Note that the absent case works by accident: -1 is truthy, so the test reports found when nothing was found. The condition is wrong in both directions and right in the middle, which is the worst possible pattern for testing. The fixes are an explicit comparison with -1, or the in operator.
Explain it
The right-answer-wrong-process idea is the most valuable thing in this lesson.
Discussion prompt
A classmate has fixed an IndexError in a traversal and their function now returns the right answer for their test. Explain why that is not yet evidence of correctness, and give them the one check to run.
Hint: It involves counting.
Answer:
Say: count the passes and compare with the number of elements. Four characters should mean four comparisons, and three means one pair was never checked.
A skipped comparison can still produce the right answer, because the characters it would have compared might have matched anyway. So a passing test proves nothing about the pair that was never examined.
The book does exactly this: after fixing the first bug, is_reverse gives the right answer and the loop only ran three times, which is suspicious. Noticing that discrepancy is the whole of the second diagnosis — and it is the habit worth passing on, because the alternative is waiting for a user to find the input that breaks it.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: Method syntax is the same idea as function syntax rearranged, and it settles quickly — but the return-rather-than-modify point catches people repeatedly, so it is worth over-learning. The optional arguments to find are what make it possible to enumerate every occurrence, which chapter 9 uses. The in-versus-find choice has a genuine bug attached to getting it wrong, which makes it the highest-value item here. And the two-index debugging is a technique rather than a fact: the print statement is easy, and the pass-counting habit that catches the silent second bug is the part that takes practice.
Connect it up
One page, from memory.
Draw it
Draw two four-character words side by side, one forwards and one backwards, with the index pairs that should be compared joined by lines. Mark which pair the buggy loop skips and which index starts out of range. Then, beside it, write the four ways of asking about a substring — in, find, find compared with -1, and find used as a condition — and mark the one that is always a bug, with the input that exposes it.
Recap
Four pages, and chapter 8 is finished: a string is a sequence, and it carries its own operations.
| If you remember one thing | It is this |
|---|---|
| From methods | They return, they never modify. Assign the result or it is lost. |
| From find | It returns a position, and 0 is a valid one. Never use it as a bare condition. |
| From in | It answers is it there; find answers where. Choose by whether you use the position. |
| From comparison | All uppercase before all lowercase. Normalise before comparing anything a person typed. |
| From debugging | Print the indices, then count the passes. The second check finds the silent bug. |
Chapter 9 is the second case study: a set of word puzzles built on a real word list, where the search patterns of this chapter meet a file of a hundred thousand words and the exercises stop being about syntax.
Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-78 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.