8c String Methods, the in Operator, and String Comparison

This lesson introduces string methods and the dot notation that invokes them, meets optional arguments through the built-in find, adds the in operator, explains why uppercase letters sort before lowercase ones, and works a live diagnosis of an index traversal with two bugs in it.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 8c String Methods, the in Operator, and String Comparison

Title

Python · Chapter 8 — Strings

§8.8-8.11, pp. 75-78

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-78 — the pages these objectives are drawn from

3. Before we start: you have seen this syntax before

Warm-up

The dot arrives again, on a third kind of thing.

Discussion prompt

You have used a dot in two situations already: with a module in chapter 3, and with a turtle in chapter 4. Write down both, and say what the general rule was.

Hint: Lesson 3a gave a single reading that covers both.

Answer:

math.sqrt(2) reaches into a module, and bob.fd(100) asks a turtle to do something.

The rule was: look inside the thing on the left for the name on the right. It covered both without needing to know what an object is.

This lesson applies it to strings, which is where it becomes an everyday tool — because unlike turtles, you use strings in every program you write.

4. The one idea behind this lesson: a string carries its own operations

Concept

Strings provide methods that perform a variety of useful operations. A method is similar to a function — it takes arguments and returns a value — but the syntax is different: instead of upper(word), it is word.upper().

invocation — A method call. We say that we invoke a method on an object.

This form of dot notation specifies the name of the method and the name of the string to apply it to. The empty parentheses indicate that the method takes no arguments — and they are required, exactly as they were for a function.

Figure (svg): Two columns contrasting function syntax with method syntax for the same operation

Same idea — takes arguments, returns a value — with the subject moved to the front.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-76

5. Methods, and the dot notation that invokes them

Section

Section 1

6. The same idea, a different arrangement

Concept

A method is similar to a function: it takes arguments and returns a value. What differs is the syntax — the string the method applies to goes before the dot rather than inside the parentheses.

>>> word = 'banana'
>>> new_word = word.upper()
>>> new_word
'BANANA'
>>> word
'banana'
ExpressionWhat happensResult
word.upper()invoke upper on wordreturns a NEW string
new_wordthe returned value'BANANA'
wordunchangedstrings are immutable

A method call is called an invocation; in this case, we would say we are invoking upper on word. The method takes a string and returns a new string with all uppercase letters — and, because strings are immutable, the original is untouched.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-76

7. Picture it: the dot means what it always meant

Picture it

Three kinds of thing on the left, one rule.

Figure (svg): A diagram showing the dot used on a module, on a turtle and on a string, with the same meaning each time

The rule from lesson 3a, unchanged. Only the kind of thing on the left differs.

Chapter 15 explains what a string and a turtle have in common that a module does not. Until then, the reading covers every case you will meet.

8. Worked example: a method returns rather than modifies

Worked example

Immutability from the previous lesson, meeting method syntax.

>>> word = 'banana'
>>> word.upper()
'BANANA'
>>> word
'banana'
>>> word = word.upper()
>>> word
'BANANA'
LineWhat happensState
word.upper()returns a new stringthe original is unchanged
word alonestill lowercasenothing was modified
word = word.upper()the result is assignednow word refers to the new string

Invoke the method and look at the result.

Why: It returns a new string with all uppercase letters. At the prompt, the returned value is displayed.

Check the original.

Why: Unchanged. Strings are immutable, so no method can modify one — every string method returns a new string.

Assign the result if you want to keep it.

Why: word = word.upper() repoints the name at the new string, which is reassignment rather than modification.

Figure (svg): The state of the program after each line of Worked example a method returns rather than modifies, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The method returns a new uppercase string and leaves the original alone. Keeping the result requires an assignment, which is lesson 6a's fruitful-function rule applied to methods.

Verify: Call the method without assigning and check nothing changed.

Why: word.upper() on its own computes a new string and discards it — the same wasted-return-value mistake as calling math.sqrt on its own in lesson 6a. It produces no error, which is exactly why it is worth causing once deliberately.

9. Predict: what does word hold afterwards?

Prediction

The method returns; it does not modify.

word = 'banana'
word.upper()
print(word)
LineWhat happensOutput
word.upper()computes 'BANANA'and discards it
no assignmentthe result goes nowhereword is untouched
print(word)the originalbanana

Predict first

What does this print?

  • BANANA
  • banana
  • None
  • An error

Correct: banana — the method returned a new uppercase string, which was discarded, and the original is unchanged.

Why: Strings are immutable, so upper cannot modify anything; it produces a new string. Since nothing assigned that string, it is lost — the same wasted return value as calling math.sqrt on its own. The fix is word = word.upper(), and the absence of any error message is what makes this worth meeting deliberately.

10. Worked example: why the parentheses are needed

Worked example

Even with no arguments, they are required. Work out what they do.

>>> word = 'banana'
>>> word.upper
<built-in method upper of str object at 0x...>
>>> word.upper()
'BANANA'
ExpressionWhat it meansResult
word.uppernames the methodthe method object itself
word.upper()calls itthe uppercase string
the parenthesesthe call operatorrequired even when empty

Write the name without parentheses.

Why: This names the method without invoking it, exactly as math.sqrt without parentheses named the function in lesson 3a.

Add the parentheses.

Why: The empty parentheses indicate that this method takes no arguments — and they are what makes it a call.

Recall the general rule.

Why: The parentheses are the operator that means run this. Their emptiness says there is nothing to supply, not that they are optional.

Figure (svg): Two columns comparing naming a method with invoking it

Two characters, and the difference between naming a thing and running it.

Without parentheses you have named the method; with them you have invoked it. The rule is identical to the one for functions, which is one fewer thing to learn.

Verify: Compare with the same mistake on a function.

Why: math.sqrt without parentheses displays a function object, and word.upper without them displays a method object. Both are legal and neither computes anything — the same symptom, and the same fix.

11. Trap: expecting a method to change the string

Trap

The trap

A student writes word.upper() on a line by itself and then prints word, expecting it to be uppercase.

Read a method call as an instruction to the string

Why: The syntax reads like a command — word, uppercase yourself — which makes modification the natural expectation.

Strings are immutable, so nothing changed. The new string was created and immediately discarded, with no error to indicate it.

The fix

Every string method returns a new string and modifies nothing.

Assign the result

Why: word = word.upper(), which repoints the name at the new string.

Treat a string method call as an expression

Why: Its value is the point, exactly as with any fruitful function. A call whose result is discarded did nothing.

This is one of the two commonest string mistakes, and the other is the same mistake on the other side: assigning the result of a method that returns None. Both come from being unsure whether an operation modifies or returns.

12. Match each form to what it produces

Matching

Four expressions built from one method name.

Match the pairs

  • a. word.upper
  • b. word.upper()
  • c. word = word.upper()
  • d. print(word.upper())
  • r1. a method object — nothing is computed
  • r2. a new uppercase string, which is then discarded
  • r3. word now refers to the uppercase string
  • r4. the uppercase string is displayed; word is unchanged

Why: The pattern is entirely mechanical, and it is the same as for functions in lesson 3b: parentheses mean invoke, and what happens to the result depends on what surrounds the call. Note that only the third one changes what word refers to, and it does so by reassignment rather than by modifying any string.

13. Compare: function syntax and method syntax

Comparison

Fill the blanks. The ideas are identical; only the arrangement differs.

Comparison matrix

QuestionFunctionMethod
How is it written?len(word)word.upper()
Where does the string go?inside the parentheses, as an argumentbefore the dot
Does it return a value?yesyes — the same as a function
Are the parentheses required?yes, even with no argumentsyes, even with no arguments

Three of the four rows are identical. The one real difference is where the value being operated on is written, which is why the book says a method is similar to a function but the syntax is different.

14. Think it through: why have two syntaxes?

Socratic

len(word) and word.upper() do comparable things in different shapes.

Discussion prompt

Python has both len(word) and word.upper(). Suggest one reason a language might attach some operations to the value and leave others as free functions.

Hint: Which operations make sense for many different types?

Answer:

len works on strings, lists, dictionaries and more, so it belongs to no single type — making it a free function lets one name cover all of them.

upper only makes sense for strings, so attaching it to strings keeps it out of the way and makes it discoverable: everything a string can do is reachable through the dot.

There is a practical benefit too: two different types can each have a method called find with behaviour appropriate to that type, without colliding. Chapter 17 shows how that works, and it is the reason methods exist at all.

15. The built-in find, and optional arguments

Section

Section 2

16. More capable than the one you wrote

Concept

There is a string method named find that is remarkably similar to the function written in the previous lesson — and it is more general in three ways.

>>> word = 'banana'
>>> word.find('a')
1
>>> word.find('na')
2
>>> word.find('na', 3)
4
>>> 'bob'.find('b', 1, 2)
-1
CallWhat it doesResult
find('a')a single character1
find('na')a SUBSTRING, not just a character2
find('na', 3)start looking at index 34
find('b', 1, 2)search only indices 1 up to but not including 2-1

The method is more general than the function we wrote: it can find substrings, not just characters. By default it starts at the beginning, but it can take a second argument, the index where it should start — and a third, the index where it should stop. This is an example of optional arguments.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 76-76

17. Picture it: three ways the built-in is more general

Picture it

Each addition is a capability the hand-written version lacks.

Figure (svg): Two columns comparing the hand-written find function with the built-in string method

Same answer shape, three more capabilities.

Writing the simpler version first was not wasted: it is why you know what the built-in is doing, and why -1 means what it means.

18. Worked example: optional arguments

Worked example

The same method, called three ways. Each adds one argument.

>>> word = 'banana'
>>> word.find('na')
2
>>> word.find('na', 3)
4
>>> word.find('na', 3, 4)
-1
FormWhat it searchesResult
one argumentsearch the whole string from 0finds the first 'na' at 2
two argumentsstart at index 3finds the second 'na' at 4
three argumentssearch indices 3 up to but not including 4too narrow: -1

Call it with one argument.

Why: By default find starts at the beginning of the string, so it reports the first occurrence.

Add the start index.

Why: It can take a second argument, the index where it should start — which is how you find the second occurrence after finding the first.

Add the stop index.

Why: A third argument gives the index where it should stop, and searching up to but not including that index makes find consistent with the slice operator.

Figure (svg): The string banana with the two occurrences of na highlighted and the search region from index three marked

The first 'na' starts at index 2, the second at index 4. Starting the search at 3 skips the first.

2, then 4, then -1. Each optional argument narrows the region searched, and the third one narrows it so far that nothing is found.

Verify: Check the consistency claim against a slice.

Why: word.find('na', 3, 5) searches the same region as word[3:5], which is 'an' — and 'na' is not in it, so both agree there is nothing to find. Searching up to but not including the second index is exactly the half-open convention from lesson 8b, applied here so the two operations agree.

19. Predict: what does this return?

Prediction

The search starts partway along.

>>> 'banana'.find('a', 2)
StepWhat happensResult
start at index 2skip the 'a' at index 1search from 'nana'
first 'a' from thereat index 33
the resultan absolute indexnot relative to the start

Predict first

What does 'banana'.find('a', 2) return?

  • 1
  • 2
  • 3
  • 0

Correct: 3 — the first 'a' at or after index 2, reported as an index into the whole string.

Why: The start argument says where to begin looking, not where to count from: the returned index is still an offset into the original string. That distinction matters — if the result were relative to the start, chaining searches would need arithmetic on every call, and the enumerate-all-occurrences loop would be much harder to write.

20. Worked example: finding every occurrence

Worked example

The start argument exists for exactly this. Use it in a loop.

word = 'banana'
index = word.find('a')
while index != -1:
    print(index)
    index = word.find('a', index + 1)
CallWhat it searchesResult
find('a')the first occurrence1
find('a', 2)search from just after it3
find('a', 4)again5
find('a', 6)nothing left-1, so the loop ends

Find the first occurrence.

Why: An ordinary call, giving index 1.

Search again from just past it.

Why: index + 1 rather than index, otherwise the same occurrence would be found forever — which is the classic infinite loop in this pattern.

Stop when the sentinel comes back.

Why: -1 means nothing further was found, so the loop condition tests for it.

Figure (svg): The state of the program after each line of Worked example finding every occurrence, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

1, 3 and 5 — every position of an 'a'. The start argument turns a single search into a way of enumerating all occurrences.

Verify: Try it with index rather than index + 1 and predict what happens.

Why: It finds the same occurrence forever, because searching from index finds the thing at index. That is an infinite loop caused by an update that does not move — lesson 7a's requirement that the body must change what the condition depends on, in a form specific to searching.

21. Trap: assuming find returns a boolean

Trap

The trap

A student writes if word.find('a'): to test whether the letter is present.

Treat a search result as a yes-or-no answer

Why: The question being asked really is yes-or-no, so a truthy test feels natural.

find returns a position. When the letter is at index 0 the result is 0, which counts as false — so the test reports not found for exactly the case where it was found first.

The fix

Compare the result with -1, or use the in operator.

Test explicitly against the sentinel

Why: if word.find('a') != -1: says what it means and cannot be confused by a zero.

Or use in, when the position does not matter

Why: if 'a' in word: is shorter and returns a genuine boolean.

This is lesson 5a's truthiness warning in its most damaging form: a valid answer of 0 counting as false. The book's advice to avoid relying on truthiness unless you know what you are doing is exactly about cases like this.

22. Discriminate: which call finds it?

Discrimination

The optional arguments narrow the search region.

Sort into buckets

For word = 'banana', which of these calls return a position rather than -1?

returns a position
word.find('a'); word.find('a', 2); word.find('ban')
returns -1
word.find('z'); word.find('na', 5); word.find('a', 1, 1)
found
Each searches a region that contains the target. Note that the third finds a three-character substring, which the hand-written function could not do at all.
not
One target is absent entirely. One starts too late — there is no complete 'na' at or after index 5. And one has a stop index equal to its start, so the region searched is empty, exactly as the slice word[1:1] is.

23. Complete it: find the second occurrence

Faded example

Use the position of the first as the starting point for the second.

Fill in the blanks

word = 'banana'
first = word.find('a')
second = word.find('a', first + 1)

Why: Searching from first + 1 skips past the occurrence already found. Searching from first itself would find the same one again, which is the classic infinite loop when this pattern goes in a while loop. The plus one is the update that moves the search forward, and it is the same requirement as lesson 7a's rule about the body changing what the condition depends on.

24. Explain it yourself: why was writing your own find worthwhile?

Explain it to yourself

The built-in is strictly better. Say what the exercise bought.

Discussion prompt

Python already had a find method when you wrote your own in the previous lesson. Name two things writing it gave you that using the built-in would not have.

Hint: One is about the pattern; one is about the -1.

Answer:

It taught the search pattern: traverse, return early on a match, return the sentinel after the loop. That pattern applies to lists, files and anything else you traverse, and no amount of calling the built-in would have shown it.

It also explains the -1. Having written the not-found return yourself, you know why the sentinel is negative and why it must be checked rather than used directly.

The general point is worth keeping: reimplementing something that already exists is a way of understanding it, not a waste. The book does it deliberately here, and again in chapter 11 with dictionaries.

25. The in operator

Section

Section 3

26. A boolean operator that reads like English

Concept

The word in is a boolean operator that takes two strings and returns True if the first appears as a substring in the second.

>>> 'a' in 'banana'
True
>>> 'seed' in 'banana'
False
>>> 'nan' in 'banana'
True
ExpressionWhat it asksResult
'a' in 'banana'a single characterTrue
'seed' in 'banana'not presentFalse
'nan' in 'banana'a substring of any lengthTrue

With well-chosen variable names, Python sometimes reads like English. The book's example is a loop you can read as: for each letter in the first word, if the letter appears in the second word, print the letter.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 76-77

27. Picture it: in versus find

Picture it

Same question, two answers, and one of them is a genuine boolean.

Figure (svg): Two columns comparing the in operator with the find method for testing presence

in answers is it there; find answers where is it.

Choosing between them is the same question as lesson 8b's search-versus-counter: what does the answer need to be? If a position is never used, in says so more clearly.

28. Worked example: the in_both function

Worked example

The book's example, and it reads almost as English.

def in_both(word1, word2):
    for letter in word1:
        if letter in word2:
            print(letter)
Letter from word1Is it in word2?Output
'a' from apples'a' in 'oranges'? yesprint a
'p'not in orangesskip
'e'yesprint e
's'yesprint s

Read the loop aloud.

Why: For each letter in the first word, if the letter appears in the second word, print the letter. The code and the sentence are nearly identical.

Notice the two different uses of the word in.

Why: The first is the for loop's keyword, meaning iterate over; the second is the boolean operator, meaning appears within. Same word, different jobs.

Trace the book's example.

Why: in_both('apples', 'oranges') prints a, e and s — the letters the two words have in common, in the order they appear in the first.

Figure (svg): A diagram showing each letter of the first word being tested for membership in the second and printed if present

The traversal comes from the for loop; the test comes from the in operator.

a, e and s. The function traverses one word and tests each letter against the other, which is the counter pattern with printing instead of counting.

Verify: Check what happens with a repeated letter.

Why: in_both('aaa', 'ab') prints a three times, because the traversal visits each character and does not track which have already been reported. That is worth noticing: the function reports occurrences rather than distinct letters, which may or may not be what was wanted.

29. Predict: is this True?

Prediction

Contiguous, and in order.

Predict first

What is 'ana' in 'banana'?

  • True
  • False
  • 1
  • An error

Correct: True — 'ana' appears starting at index 1, and again at index 3.

Why: The characters a, n, a appear consecutively and in that order, so the operator returns True. Note that it returns a boolean rather than a position, unlike find — so it can be used directly in a condition without any comparison. If you answered 1, you were thinking of find, which is the operator that reports where.

30. Worked example: in on a substring rather than a character

Worked example

It is not limited to single characters, which is easy to miss.

>>> 'nan' in 'banana'
True
>>> 'ban' in 'banana'
True
>>> 'nab' in 'banana'
False
SubstringWhyResult
'nan'appears at index 2True
'ban'appears at index 0True
'nab'those letters are all present, in that ORDER they are notFalse

Test a multi-character substring.

Why: in returns True if the first appears as a substring in the second — the characters must be contiguous and in order.

Notice the third case.

Why: 'nab' contains only letters that are present in 'banana', and it is still False, because they do not appear together in that order.

Connect to lesson 2b.

Why: This is the ordering point again: a string is a sequence, so containment is about consecutive positions rather than about which characters exist.

Figure (svg): The state of the program after each line of Worked example in on a substring rather than a character, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

True, True, False. The operator tests for a contiguous substring, not for the presence of the individual characters.

Verify: Test a case where every character is present but not adjacent.

Why: 'bn' in 'banana' is False, even though both letters appear. Confirming that containment requires adjacency is worth doing once, because contains these letters and contains this substring are easy to conflate.

31. Trap: reading in as *contains these characters*

Trap

The trap

A student uses 'nab' in 'banana' to test whether the word is an anagram of nab, or contains those letters.

Read containment as being about the set of characters

Why: All three letters really are in banana, so the reading is not unreasonable.

The operator tests for a contiguous substring in order, so it returns False — and a test written this way silently answers a different question from the one intended.

The fix

in tests for a substring: contiguous, and in order.

Say it precisely: does this exact run of characters appear?

Why: Not are these characters present, which is a different question.

For the other question, test each character separately

Why: Which is what in_both does — one in test per character, in a loop.

The distinction matters because both questions are common and the syntax for one of them looks like it should answer the other. Being explicit about which you mean is the whole defence.

32. Discriminate: in or find?

Discrimination

Ask whether the position is used.

Sort into buckets

For each task, which is the better tool?

the in operator
test whether a word contains a vowel; decide whether to print a warning about a missing digit; check whether a filename ends in a known extension
the find method
report where the first vowel appears; split a string at the first space; find every position of a letter
in
Each of these needs a yes-or-no answer and never uses a position. in returns a genuine boolean and reads as the question being asked.
find
Each of these needs the position: to report it, to slice at it, or to continue searching from it. Only find supplies that, and its optional start argument is what makes the last one possible.

33. Complete it: test for a substring

Faded example

One operator, no comparison needed.

Fill in the blanks

filename = 'report.txt'
if '.txt' in filename:
print('a text file')

Why: in returns True or False directly, so it can be the whole condition with no comparison attached. Writing filename.find('.txt') != -1 would work identically and is longer; writing filename.find('.txt') alone would be a bug, since a match at index 0 gives 0, which counts as false. That last case is why in is preferred whenever the position is not needed.

34. Explain it: when should I use in rather than find?

Explain it

A choice beginners face constantly.

Discussion prompt

A classmate uses find for everything, including simple presence tests. Give them the rule, and one concrete bug that in would have prevented.

Hint: The bug involves a match at position zero.

Answer:

Say: use in when you only need to know whether something is there, and find when you need to know where.

The bug: if word.find('b') is a truthiness test, and 'b' is the first character, find returns 0 — which counts as false. So the test reports not found precisely when the target is at the front.

in cannot have that bug, because it returns a real boolean. That is a stronger argument than readability: choosing in eliminates an entire failure mode rather than merely looking nicer.

35. String comparison, and the case surprise

Section

Section 4

36. The relational operators work on strings

Concept

The relational operators work on strings, both for equality and for ordering — which is how words are put into alphabetical order.

if word == 'banana':
    print('All right, bananas.')

if word < 'banana':
    print('Your word comes before banana.')
elif word > 'banana':
    print('Your word comes after banana.')
else:
    print('All right, bananas.')
OperatorWhat it asksNote
==exact equality, character for characterTrue or False
<comes before, alphabeticallyfor lowercase words
the chainthree cases, one of which always applieslesson 5b's shape

Python does not handle uppercase and lowercase letters the same way people do. All the uppercase letters come before all the lowercase letters, so the word Pineapple comes before banana — which is almost certainly not what a person sorting words would say.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 77-77

37. Picture it: all uppercase before all lowercase

Picture it

Not A before a, then B before b. Every uppercase letter comes before every lowercase one.

Figure (svg): A row showing capital A through Z followed by lowercase a through z in comparison order

The highlighted boundary is where all the uppercase letters end and all the lowercase ones begin.

So Pineapple comes before banana, and Zebra comes before apple. The rule is consistent and it is not alphabetical order as people mean it.

38. Worked example: the case surprise

Worked example

Two comparisons that a person would answer differently.

>>> 'Pineapple' < 'banana'
True
>>> 'Zebra' < 'apple'
True
>>> 'apple' < 'banana'
True
ComparisonWhyResult
'Pineapple' vs 'banana'P is uppercase, b is lowercaseall uppercase comes first
'Zebra' vs 'apple'Z is uppercase, a is lowercasesame reason
'apple' vs 'banana'both lowercaseordinary alphabetical order

Compare the first characters.

Why: String comparison works character by character from the left, and the first difference decides the answer.

Apply the case rule.

Why: All the uppercase letters come before all the lowercase letters, so any word starting with a capital comes before any word starting with a small letter.

Notice the third case is unsurprising.

Why: Within a single case, the ordering is exactly what you would expect. The surprise only arises when the cases are mixed.

Figure (svg): The state of the program after each line of Worked example the case surprise, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

All three are True, and only the first two are surprising. Mixed-case comparison is not the alphabetical order people mean, and it is entirely consistent.

Verify: Sort a mixed-case list mentally and check the result.

Why: banana, Apple and cherry would sort as Apple, banana, cherry — which looks right by luck, since A is the earliest letter anyway. Trying Zebra, apple, Banana gives Banana, Zebra, apple, which is visibly not alphabetical. Finding an example where the difference shows is what makes the rule stick.

39. Predict: which comes first?

Prediction

The first differing character decides, and case matters.

Predict first

Is 'Zebra' < 'apple' True or False?

  • True
  • False
  • They are equal
  • An error

Correct: True — Z is an uppercase letter and a is lowercase, and all uppercase letters come before all lowercase letters.

Why: The comparison looks at the first characters, Z and a, and the case rule decides: every uppercase letter sorts before every lowercase one. This is consistent rather than random, and it is not the ordering a person means by alphabetical. Converting both with lower() first gives False, which is the human answer.

40. Worked example: the standard fix

Worked example

Convert to one case before comparing.

>>> 'Pineapple'.lower() < 'banana'.lower()
False
>>> 'Pineapple'.lower()
'pineapple'
StepWhat happensResult
lower()returns a new lowercase string'pineapple'
the comparisonboth now lowercasep comes after b
the resultFalsewhich is what a person would say

Convert both to a common case.

Why: A common way to address this problem is to convert strings to a standard format, such as all lowercase, before performing the comparison.

Compare the converted values.

Why: Now every letter is in the same case, so the ordering is ordinary alphabetical order.

Note what has not changed.

Why: lower returns a new string, so the originals are untouched — immutability again, and the conversion is for the comparison only.

Figure (svg): Two columns comparing a mixed-case comparison with the same comparison after converting both to lowercase

One method call on each side, and the comparison means what you meant.

False, which is the answer a person sorting words would give. Normalising the case before comparing is the standard technique, and it costs one method call on each side.

Verify: Check that the conversion does not affect the stored values.

Why: The original strings are unchanged, because lower returns a new string. That is what makes this safe to do inside a comparison — you are normalising for the purpose of the test, not editing your data.

41. Trap: sorting user input without normalising case

Trap

The trap

A program sorts a list of names typed by users and produces an order that looks scrambled.

Assume string comparison means alphabetical order

Why: It does, within a single case, and most test data is consistently capitalised.

Real input is not consistent. A single name typed in lowercase sorts after every capitalised one, which looks like a bug in the sorting rather than in the comparison.

The fix

Normalise before comparing whenever the data comes from people.

Convert to a standard format, such as all lowercase

Why: Which is the book's own advice, and one method call.

Decide whether to normalise for comparison only, or to store the normalised form

Why: Usually the first: display what the user typed, compare on the normalised version.

The book's closing joke about defending yourself against a man armed with a Pineapple is memorable for a reason — the surprise is specific enough that having a name for it makes it easy to spot.

42. Rank: put these in Python's comparison order

Ranking

Uppercase first, then lowercase, alphabetically within each.

Put in order

  1. 'Apple'
  2. 'Banana'
  3. 'apple'
  4. 'banana'

Why: Both capitalised words come before both lowercase ones, and within each group ordinary alphabetical order applies. So Apple, Banana, apple, banana — which interleaves nothing and looks wrong to a person expecting Apple, apple, Banana, banana. That gap between the two orderings is exactly what case normalisation closes.

43. Complete it: compare without regard to case

Faded example

One method call on each side.

Fill in the blanks

if word.lower() == 'banana':
print('All right, bananas.')

Why: Converting the input to lowercase before comparing means Banana, BANANA and banana all match. Note that the literal on the right is already lowercase, so only one conversion is needed — but converting both sides is the safer habit, since the literal might change later. The method returns a new string, so word itself is unchanged.

44. Find the counterexample: is lowercasing always the right fix?

Counterexample

It fixes the case surprise. Find a case it does not fix.

Discussion prompt

Converting both strings to lowercase makes comparison match human expectations for English words. Give an example where it still does not, and say what that suggests.

Hint: Think about digits, spaces, or accented letters.

Answer:

Digits and punctuation sort before letters, so '2nd' comes before 'first' even after lowercasing — which is not how a person would order a list.

Accented letters sort after all unaccented ones, so a name beginning with an accented character lands at the end rather than beside its unaccented neighbour.

What that suggests is that character comparison is not the same thing as alphabetical order, and lowercasing closes only the most common gap. Real sorting of human-readable text needs rules that the comparison operators do not implement — which is worth knowing exists rather than solving here.

45. Debugging an index traversal

Section

Section 5

46. Two bugs, found by printing the indices

Concept

When you use indices to traverse the values in a sequence, it is tricky to get the beginning and end of the traversal right. The book's is_reverse function is supposed to return True when one word is the reverse of the other, and it contains two errors.

def is_reverse(word1, word2):
    if len(word1) != len(word2):
        return False
    i = 0
    j = len(word2)
    while j > 0:
        if word1[i] != word2[j]:
            return False
        i = i+1
        j = j-1
    return True
PartIts roleNote
the guardiandifferent lengths cannot be reverseslesson 6c's pattern
itraverses word1 forwardfrom 0
jtraverses word2 backwardfrom len(word2)
the bugj starts one too highIndexError

The first if statement checks whether the words are the same length; if not, we can return False immediately. This is an example of the guardian pattern — for the rest of the function we can assume the words are the same length.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 77-78

47. Picture it: two indices moving in opposite directions

Picture it

i walks forward through one word while j walks backward through the other.

Figure (svg): Two columns showing i moving forward through pots and j moving backward through stop

Paired correctly, the characters match at every step — which is what makes them reverses.

The right-hand column starts at 3, not 4. The bug is that the code starts j at len(word2), which is one past the end.

48. Worked example: printing the indices

Worked example

The book's first move, and it is the whole technique.

    while j > 0:
        print(i, j)
        if word1[i] != word2[j]:
            return False
        i = i+1
        j = j-1
StageWhat you doWhat you learn
the symptomIndexError on the comparison linewhich index?
the printimmediately before the failing lineshows both
the output0 4j is 4, and 'pots' has no index 4

Locate the failing line from the traceback.

Why: It names the comparison, which uses both indices — so the traceback alone cannot say which one is wrong.

Print the values immediately before it.

Why: For debugging this kind of error, the book's first move is to print the values of the indices immediately before the line where the error appears.

Read the output.

Why: The first time through the loop, the value of j is 4, which is out of range for a four-character string. The index of the last character is 3, so the initial value for j should be len(word2)-1.

Figure (svg): A traceback showing an IndexError on the comparison line, with the printed index values above it

j starts at 4 when the largest valid index is 3. One print statement identified which of the two indices was wrong and by how much.

Verify: Check that the fix makes the first pass valid.

Why: With j starting at len(word2)-1, the first pass has i = 0 and j = 3, which are both valid. This is lesson 8a's off-by-one at the end of a string — len is one past the last index — appearing in a place where two indices made it harder to see.

49. Predict: what does the print show first?

Prediction

The indices are printed before the failing comparison.

i = 0
j = len(word2)
while j > 0:
    print(i, j)
    if word1[i] != word2[j]:
VariableIts initial valuePrinted
iinitialized to 00
jinitialized to len(word2)4 for a four-letter word
the printruns before the comparison0 4

Predict first

For is_reverse('pots', 'stop'), what does the print statement show on the first pass?

  • 0 3
  • 0 4
  • 1 4
  • 0 0

Correct: 0 4 — i starts at 0 and j starts at len(word2), which is 4.

Why: That single line is the diagnosis: 4 is out of range for a four-character string, whose valid indices are 0 to 3. The traceback pointed at the comparison line, which uses both indices, so it could not say which one was wrong — and one print statement settles it immediately. This is why the book's first move is to print the indices rather than to reread the code.

50. Worked example: the second bug, and why it is suspicious

Worked example

The fix works and the loop count is wrong. The book leaves this as an exercise.

>>> is_reverse('pots', 'stop')
0 3
1 2
2 1
True
ObservationWhat it isWhy it matters
the answerTruecorrect for this input
the loop countthree passesfor four-character words
the suspicionone pair was never comparedthe condition stops early

Notice the right answer.

Why: This time we get the right answer — 'pots' and 'stop' really are reverses.

Notice the loop count.

Why: It looks like the loop only ran three times, which is suspicious for words of four characters.

Draw the conclusion without fixing it.

Why: One pair of characters was never compared, so there are inputs for which this function would return True wrongly. The book invites you to draw a state diagram, run the program on paper, and find the second error.

Figure (svg): The state of the program after each line of Worked example the second bug, and why it is suspicious, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The loop runs three times for four-character words, so one comparison is skipped. The function gives the right answer here and would not for every input — which is exactly the kind of bug a correct-looking test hides.

Verify: Look for an input where the missing comparison would matter.

Why: The skipped pair is the one where j reaches 0, so a difference in that single position would go unnoticed. Constructing such an input is the exercise — and being able to say WHERE the gap is, from the loop count alone, is most of the work.

51. Trap: accepting a right answer as evidence of correctness

Trap

The trap

After fixing the IndexError, the function returns True for pots and stop, so it is declared working.

Treat a correct output as proof

Why: The test passed, and the previous version crashed, so progress is obvious.

The loop ran three times for four-character words. The right answer arrived despite a comparison being skipped, which means some other input will get the wrong answer.

The fix

Check the process, not only the answer.

Count the passes and compare with what you expect

Why: Four characters should mean four comparisons. Three is a discrepancy that demands an explanation.

Draw the state diagram and run it on paper

Why: The book's own advice: start with the diagram, changing i and j at each iteration, and find the second error.

This is the deepest debugging lesson in the chapter. A correct output from an incorrect process is the most dangerous state a program can be in, because every test that passes makes it more entrenched.

52. Error analysis: is_reverse, both bugs

Error analysis

One has been fixed. Mark both, and say why only one produced an error.

Annotate

  • The first bug is on the second line: j is initialized to len(word2), which is one past the last valid index. The first comparison raises an IndexError.
  • That bug is loud — it stops the program on the first pass, which is why it was found first.
  • The second bug is in the condition. With j counting down and the loop running while j > 0, the pass where j is 0 never happens.
  • So the character at index 0 of word2 is never compared with anything, and the loop runs one time fewer than there are characters.
  • That bug is silent: it produces a correct answer for pots and stop, and a wrong one for any pair differing only in that position.
  • The two bugs are at the two ends of the traversal, which is exactly where lesson 8a said traversal errors live — and only one of them announces itself.

Fixing the loud bug revealed the quiet one, but only to somebody who counted the passes. That counting step is the whole technique.

53. Watch the indices: is_reverse after the first fix

Invariant

Step through and watch which pairs are compared.

Step through it

Which pair of characters is never compared, and what condition would include it?

  1. The first pass compares the first character of word1 with the last of word2. Both are 'p'.
  2. The indices move toward each other. Second character against second-to-last.
  3. Third pass. Three comparisons so far, and the loop condition is about to fail.
  4. j reaches 0, and the condition j > 0 is false — so this pair is never compared, even though both indices are valid.

The pair at i = 3 and j = 0. The condition needs to remain true when j is 0, so j >= 0 rather than j > 0 — which is exactly the boundary question from lesson 8a's backward traversal.

54. Explain it: how do you debug an index bug?

Explain it

The technique is one line long and it is worth passing on.

Discussion prompt

A classmate has an IndexError on a line that uses two indices, and cannot tell which one is wrong. Tell them what to do, and why the traceback alone was not enough.

Hint: The traceback names the line, not the value.

Answer:

Say: put a print of both indices on the line immediately before the failing one, and run it again. The first line of output tells you which index is out of range and by how much.

The traceback could not tell them because it names the line and never the values — which is lesson 3c's point that a traceback locates a failure without explaining it.

Then add the follow-up that catches the second bug: once it works, count the passes and check the count against the length. A right answer from too few passes means a comparison was skipped, and that is a bug waiting for a different input.

55. Compare: four ways to ask about a substring

Comparison

Fill the blanks. They answer different questions and are easy to confuse.

Comparison matrix

ExpressionWhat it returnsUse it when
'a' in wordTrue or Falseyou only need to know whether it is there
word.find('a')a position, or -1you need to know where
word.find('a') != -1True or Falseyou have a position already and want a boolean too
word.find('a') used as a conditiona position, treated as truthynever — a match at index 0 counts as false

The last row is a bug rather than a technique. It is the clearest case of lesson 5a's truthiness warning, and in is the fix.

56. The procedure: debugging a traversal that uses indices

Pattern

Six steps, and the last two are what catch the silent bug.

  1. Read the traceback for the failing line, and note that it names the line rather than the value.
  2. Add a print of every index immediately before that line, and run again.
  3. Compare each printed index against the valid range: 0 to len minus one.
  4. Fix the initialization or the condition, whichever the values implicate.
  5. Count the passes, and compare the count with the number of elements you expected to process.
  6. If the count is short, find which pair was skipped — a right answer from too few passes is a bug that has not shown itself yet.

Steps 5 and 6 are what separates this from ordinary debugging. Two of the three traversal failures are silent, so a correct output is not evidence that the traversal is correct.

Python documentation — Built-in Types Built-in Types

57. Check yourself 1 of 3: methods

Check

The method returns a new string.

word = 'banana'
word.upper()
print(word)
LineWhat happensOutput
word.upper()returns 'BANANA'and discards it
wordunchangedstrings are immutable
print(word)the originalbanana

Check your understanding

What does this print?

  • A. BANANA
  • B. banana (correct)
  • C. None
  • D. An error

Answer: B

Why: Strings are immutable, so upper cannot modify anything — it returns a new uppercase string, which is discarded because nothing assigns it. The original is untouched. Keeping the result requires word = word.upper(), and the absence of any error is what makes this mistake easy to miss.

Why A tempts people
This would require the method to modify the string in place, which immutability forbids. Every string method returns a new string.
Why C tempts people
The method returns a real string, not None. It is discarded rather than being nothing.
Why D tempts people
The call is perfectly legal and does exactly what it should. The mistake is in not using the result.

58. Check yourself 2 of 3: find versus in

Check

One returns a position and one returns a boolean.

>>> word = 'banana'
>>> word.find('b')
0
ExpressionWhat it givesNote
find('b')'b' is at index 00
0 as a conditionzero counts as falsethe trap
'b' in worda genuine booleanTrue

Check your understanding

Why is if word.find('b'): a bug?

  • A. Because find returns a string rather than a number
  • B. Because a match at index 0 returns 0, which counts as false (correct)
  • C. Because find requires three arguments
  • D. Because find raises an error when the character is absent

Answer: B

Why: find returns the position, and a match at the very start gives 0 — which is treated as false in a condition. So the test reports not found precisely when the character is at the front. The fixes are to compare explicitly against -1, or to use the in operator, which returns a real boolean and cannot have this bug.

Why A tempts people
find returns an integer index, which is exactly why the truthiness problem arises: 0 is a legitimate answer and a falsy value.
Why C tempts people
The second and third arguments are optional. A single argument is the normal form.
Why D tempts people
find returns -1 rather than raising. That sentinel is itself a source of bugs when used unchecked as a slice index, but it is not an error.

59. Check yourself 3 of 3: string comparison

Check

The first differing character decides, and case matters.

Check your understanding

Which of these is True in Python?

  • A. 'apple' < 'Apple'
  • B. 'Zebra' < 'apple' (correct)
  • C. 'banana' == 'Banana'
  • D. 'b' < 'a'

Answer: B

Why: All the uppercase letters come before all the lowercase ones, so any capitalised word sorts before any lowercase word regardless of the letters. Z before a is the clearest demonstration, and it is why comparing user input usually needs a lower() on both sides first.

Why A tempts people
This is backwards: 'Apple' comes first, because A is uppercase. The comparison as written is False.
Why C tempts people
Equality is exact, character for character, so a difference in case makes two strings unequal.
Why D tempts people
Within a single case, ordering is ordinary alphabetical order, so 'a' comes before 'b' and this is False.

60. Where this shows up outside this course

Real world

Sorting that does not match human expectations is a familiar frustration.

Discussion prompt

Think of a list you have seen sorted in a way that looked wrong — files, contacts, songs. What ordering was the machine using, and what ordering did you expect?

Hint: Numbers in filenames are the classic case.

Answer:

File listings are the standard example: file10 sorts before file2, because the comparison is character by character and '1' comes before '2'.

Mixed-case contact lists show the same thing as the Pineapple example: everything capitalised sorts before everything lowercase, which looks scrambled.

The general lesson is that character comparison and human ordering are different things that usually agree. Where they disagree, somebody has to normalise — lowercasing, or padding numbers with zeros — and knowing that the machine's ordering is consistent rather than random is what makes the fix findable.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence. This one has caught professionals.

Predict first

A program tests presence with if word.find(letter):. For which case does it give the wrong answer?

  • When the letter is absent
  • When the letter is at index 0
  • When the letter appears more than once
  • It is always correct

Correct: When the letter is at index 0 — find returns 0, which counts as false, so the test reports not found even though it was found first.

Why: This is lesson 5a's truthiness warning in its most damaging form. A perfectly valid answer, 0, is a falsy value, so the condition inverts the meaning for exactly one case — and that case is the very first position, which real data hits constantly. Note that the absent case works by accident: -1 is truthy, so the test reports found when nothing was found. The condition is wrong in both directions and right in the middle, which is the worst possible pattern for testing. The fixes are an explicit comparison with -1, or the in operator.

62. Explain it to someone else

Explain it

The right-answer-wrong-process idea is the most valuable thing in this lesson.

Discussion prompt

A classmate has fixed an IndexError in a traversal and their function now returns the right answer for their test. Explain why that is not yet evidence of correctness, and give them the one check to run.

Hint: It involves counting.

Answer:

Say: count the passes and compare with the number of elements. Four characters should mean four comparisons, and three means one pair was never checked.

A skipped comparison can still produce the right answer, because the characters it would have compared might have matched anyway. So a passing test proves nothing about the pair that was never examined.

The book does exactly this: after fixing the first bug, is_reverse gives the right answer and the loop only ran three times, which is suspicious. Noticing that discrepancy is the whole of the second diagnosis — and it is the habit worth passing on, because the alternative is waiting for a user to find the input that breaks it.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • Method syntax, and the fact that string methods return rather than modify
  • The built-in find, and its optional start and stop arguments
  • The in operator, and when to prefer it over find
  • Debugging a two-index traversal by printing the indices

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: Method syntax is the same idea as function syntax rearranged, and it settles quickly — but the return-rather-than-modify point catches people repeatedly, so it is worth over-learning. The optional arguments to find are what make it possible to enumerate every occurrence, which chapter 9 uses. The in-versus-find choice has a genuine bug attached to getting it wrong, which makes it the highest-value item here. And the two-index debugging is a technique rather than a fact: the print statement is easy, and the pass-counting habit that catches the silent second bug is the part that takes practice.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw two four-character words side by side, one forwards and one backwards, with the index pairs that should be compared joined by lines. Mark which pair the buggy loop skips and which index starts out of range. Then, beside it, write the four ways of asking about a substring — in, find, find compared with -1, and find used as a condition — and mark the one that is always a bug, with the input that exposes it.

65. What you can do now

Recap

Four pages, and chapter 8 is finished: a string is a sequence, and it carries its own operations.

If you remember one thingIt is this
From methodsThey return, they never modify. Assign the result or it is lost.
From findIt returns a position, and 0 is a valid one. Never use it as a bare condition.
From inIt answers is it there; find answers where. Choose by whether you use the position.
From comparisonAll uppercase before all lowercase. Normalise before comparing anything a person typed.
From debuggingPrint the indices, then count the passes. The second check finds the silent bug.

Chapter 9 is the second case study: a set of word puzzles built on a real word list, where the search patterns of this chapter meet a file of a hundred thousand words and the exercises stop being about syntax.

Think Python, 2nd edition — Allen B. Downey §8.8-8.11, pp. 75-78 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §8.8-8.11, pp. 75-78
  2. Python documentation — Built-in Types
  3. Python documentation — Expressions

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108