This lesson introduces the slice operator and the between-the-characters picture that explains it, discovers that strings cannot be changed at all, and names the two computational patterns most string processing is built from: the search and the counter.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 8 — Strings
§8.4-8.7, pp. 73-75
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 73-75 — the pages these objectives are drawn from
Warm-up
You can take one. Now try to take several.
Discussion prompt
With only the bracket operator from the last lesson, how would you get the first three characters of 'banana' as a single string? Write it out, and say what is unsatisfying about it.
Hint: One character at a time, joined together.
Answer:
fruit[0] + fruit[1] + fruit[2] works and gives 'ban'.
What is unsatisfying is that it hard-codes the number three. Taking the first n characters would need a loop, for something that ought to be one operation.
A segment of a string is called a slice, and Python has an operator for exactly this. It is written with a colon, and the whole of this section is about the one thing people find surprising about it.
Concept
The slice operator returns the part of the string from one index to another, including the first and excluding the last. That is counterintuitive as stated — and it becomes obvious if you imagine the indices pointing between the characters rather than at them.
slice — A part of a string specified by a range of indices.
This is the book's own device, and figure 8.1 draws it. Under that reading a slice is not characters n to m but everything between position n and position m, which is a length rather than a list of positions.
Figure (svg): The string banana with index markers drawn between the characters, from zero to six
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 73-73 — figure 8.1, slice indices
Section
Section 1
Concept
A segment of a string is called a slice, and selecting one is similar to selecting a character. The operator takes two indices separated by a colon.
>>> s = 'Monty Python'
>>> s[0:5]
'Monty'
>>> s[6:12]
'Python'| Expression | What it takes | Result |
|---|---|---|
| s[0:5] | from boundary 0 to boundary 5 | five characters |
| s[6:12] | from boundary 6 to boundary 12 | six characters |
| the rule | m minus n characters | no adjustment needed |
The operator returns the part of the string from the n-eth character to the m-eth character, including the first and excluding the last. This behaviour is counterintuitive, but it might help to imagine the indices pointing between the characters.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 73-73
Picture it
Six characters have seven boundaries, numbered 0 to 6.
Figure (svg): The string banana with boundary markers numbered zero through six drawn between and around the characters
Under this reading, s[0:5] taking five characters needs no explanation at all: 5 minus 0 is 5. The half-open convention stops being a rule and becomes arithmetic.
Worked example
The clean arithmetic is the whole reason for the convention.
>>> s = 'Monty Python'
>>> len(s[0:5])
5
>>> len(s[6:12])
6
>>> len(s[2:2])
0| Slice | The subtraction | Length |
|---|---|---|
| s[0:5] | 5 - 0 | 5 characters |
| s[6:12] | 12 - 6 | 6 characters |
| s[2:2] | 2 - 2 | 0 characters |
Compute the length as a subtraction.
Why: The number of characters in s[n:m] is m minus n, with no correction term at all.
Check it on the third case.
Why: Two minus two is zero, and indeed the slice is empty. The arithmetic keeps working at the boundary.
Notice what an inclusive convention would cost.
Why: If the end were included, the length would be m minus n plus one, and every calculation involving slice lengths would carry that plus one — which is exactly the kind of adjustment that produces off-by-one errors.
Figure (svg): Two columns comparing the half-open slice convention with a hypothetical inclusive one
The length of a slice is always the difference of its indices. That is the payoff of excluding the end, and it is why the convention is worth the initial surprise.
Verify: Check that adjacent slices join up without overlap or gap.
Why: s[0:5] and s[5:12] together give the whole string, with no character repeated and none missing. Under an inclusive convention the second slice would have to start at 6, and forgetting that is a classic bug. Adjacent slices sharing a boundary index is a direct consequence of the half-open rule.
Prediction
Boundaries 1 to 4, which is three characters.
>>> s = 'banana'
>>> s[1:4]| Marker | Where it is | What is between |
|---|---|---|
| boundary 1 | just after the b | start here |
| boundary 4 | just after the second n | stop here |
| between them | a, n, a | 'ana' |
Predict first
What is s[1:4] when s is 'banana'?
Correct: 'ana' — three characters, because 4 minus 1 is 3.
Why: The slice starts at index 1, which is the second character, and stops before index 4. The subtraction gives the length immediately: three characters. If you answered 'anan', you included the character at index 4, which is the half-open convention being read as inclusive.
Worked example
Leaving out an index means all the way to the end.
>>> fruit = 'banana'
>>> fruit[:3]
'ban'
>>> fruit[3:]
'ana'
>>> fruit[:]
'banana'| Slice | What it means | Result |
|---|---|---|
| fruit[:3] | from the beginning to boundary 3 | 'ban' |
| fruit[3:] | from boundary 3 to the end | 'ana' |
| fruit[:] | from the beginning to the end | the whole string |
Omit the first index.
Why: If you omit the first index, before the colon, the slice starts at the beginning of the string.
Omit the second.
Why: If you omit the second index, the slice goes to the end of the string.
Omit both.
Why: The book asks what fruit[:] means and invites you to try it. It gives the whole string — from the beginning to the end.
Figure (svg): The state of the program after each line of Worked example the omitted-index forms, drawn as a ladder with one rung per traced line
'ban', 'ana', and the whole string. The omitted forms are shorthand for 0 and len, and they are used constantly because they need no arithmetic.
Verify: Check that the two halves reassemble.
Why: fruit[:3] + fruit[3:] gives 'banana' exactly. Splitting at any boundary and rejoining always reproduces the original, which is another consequence of the half-open convention and a useful sanity check.
Trap
A student writes s[0:5] expecting six characters, reasoning that indices 0 through 5 is six positions.
Read the indices as naming characters
Why: Which is how the bracket operator works with a single index, so the reading carries over naturally.
The slice gives five characters, and every slice in the program is one character short of what was intended — silently.
Read the indices as boundaries between characters, not as characters.
Use the book's picture
Why: Imagine the indices pointing between the characters, as in figure 8.1. Then s[0:5] is everything between boundary 0 and boundary 5.
Check with the subtraction
Why: The length is m minus n. If you wanted six characters starting at 0, the slice is s[0:6].
The convention is genuinely counterintuitive at first and the book says so. What makes it worth learning is that every piece of slice arithmetic comes out without a correction term.
Matching
For fruit = 'banana'.
Match the pairs
Why: The first two split the string at boundary 3 and together reproduce it. The third omits both indices and gives the whole string. The fourth takes two characters, because 4 minus 2 is 2 — the subtraction gives the length in every case, which is the quickest way to check a slice before running it.
Faded example
Two ways to write it. Give the one that needs no length.
Fill in the blanks
word = 'banana'
last_three = word[-3:] # should be 'ana'
Why: A negative start index counts back from the end, so word[-3:] takes the last three characters whatever the string's length. The positive equivalent is word[len(word)-3:], which needs the length and a subtraction. Negative indices work in slices exactly as they do in single-character selection, and for the same reason: they save an arithmetic step where the arithmetic could go wrong.
Edge cases
The book states the answer. Predict it, then explain it.
Discussion prompt
What does fruit[3:3] produce, and what about fruit[4:2]? Explain both using the boundary picture rather than as special cases.
Hint: How much is between boundary 3 and boundary 3?
Answer:
Both give the empty string. If the first index is greater than or equal to the second, the result is an empty string, represented by two quotation marks.
The boundary picture explains it without a special rule: there is nothing between boundary 3 and boundary 3, and nothing between boundary 4 and boundary 2 going forwards.
An empty string contains no characters and has length 0, but other than that it is the same as any other string — it can be concatenated, compared and traversed, and traversing it does nothing, which is exactly right.
Section
Section 2
Concept
It is tempting to use the bracket operator on the left side of an assignment, with the intention of changing a character in a string. Python refuses.
immutable — A property of a value that cannot be modified after it is created.
>>> greeting = 'Hello, world!'
>>> greeting[0] = 'J'
TypeError: 'str' object does not support item assignment| Line | What happens | Note |
|---|---|---|
| greeting[0] = 'J' | an attempt to change one character | TypeError |
| the reason | strings are immutable | you cannot change an existing string |
| the alternative | create a new string | a variation on the original |
The object in this case is the string and the item is the character you tried to assign. For now, an object is the same thing as a value — the book refines that definition in chapter 10.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 74-74
Picture it
The original is never touched. A new string is made from pieces of it.
Figure (svg): Two columns contrasting an attempt to modify a string in place with building a new string from a slice
Notice that this is the same move as lesson 2b's concatenation: nothing is modified, a new value is produced. Immutability is why that was always true.
Worked example
One character replaced, without modifying anything.
>>> greeting = 'Hello, world!'
>>> new_greeting = 'J' + greeting[1:]
>>> new_greeting
'Jello, world!'
>>> greeting
'Hello, world!'| Expression | What it produces | Result |
|---|---|---|
| greeting[1:] | everything from index 1 onward | 'ello, world!' |
| 'J' + ... | concatenate a new first character | a new string |
| greeting | unchanged | still 'Hello, world!' |
Take everything except the character being replaced.
Why: greeting[1:] is a slice from index 1 to the end, which drops the first character.
Concatenate the replacement onto the front.
Why: This example concatenates a new first letter onto a slice of greeting.
Check the original.
Why: It has no effect on the original string. Two strings now exist, and greeting still refers to the first.
Figure (svg): The state of the program after each line of Worked example building the variation, drawn as a ladder with one rung per traced line
A new string, 'Jello, world!', with the original untouched. Replacing a character means building a new string from the parts you want to keep.
Verify: Check that both strings still exist.
Why: greeting is unchanged and new_greeting holds the variation, so nothing was destroyed. That is the defining property: an immutable value cannot be modified, so an apparent modification must have produced something new.
Prediction
The bracket operator on the left of an assignment.
s = 'hello'
s[0] = 'H'| Part | What it means | Result |
|---|---|---|
| s[0] | on the LEFT of an assignment | an attempt to modify |
| strings are immutable | modification is not supported | TypeError |
| the message | str object does not support item assignment | names the problem |
Predict first
What happens when this runs?
Correct: A TypeError saying that a str object does not support item assignment.
Why: Strings are immutable, so there is no way to change an existing one. Note that this is a TypeError rather than a SyntaxError: the line is perfectly well-formed, and it is the type of the thing being assigned to that makes the operation unavailable. Lists in chapter 10 accept exactly this syntax, which is why the error is about the type rather than about the brackets.
Worked example
The same technique, with two slices rather than one.
>>> word = 'banana'
>>> new_word = word[:2] + 'X' + word[3:]
>>> new_word
'baXana'| Part | What it takes | Result |
|---|---|---|
| word[:2] | everything before index 2 | 'ba' |
| 'X' | the replacement | one character |
| word[3:] | everything after index 2 | 'ana' |
Take the part before the character.
Why: word[:2] stops at boundary 2, so it excludes the character at index 2 — which is the one being replaced.
Supply the replacement.
Why: Concatenated in the middle.
Take the part after.
Why: word[3:] starts at boundary 3, so it skips the replaced character and keeps everything after it.
Figure (svg): The string banana with the character at index two highlighted and the two surrounding slices marked
'baXana'. The two slices are word[:i] and word[i+1:], which between them cover everything except the character at index i.
Verify: Check the boundaries by reassembling without the replacement.
Why: word[:2] + word[2:] gives 'banana' back, so the character at index 2 is exactly the one dropped when the second slice starts at 3. Testing the reassembly first is a good way to get the two indices right before adding the replacement.
Trap
A student calls a string method that returns a modified string and does not assign the result, expecting the original to have changed.
Assume an operation on a value modifies it
Why: Which is true for many things and is exactly what immutability rules out for strings.
The method returns a new string and the original is untouched, so the program continues with the unmodified value and nothing indicates why.
Every string operation produces a new string. Assign it, or it is lost.
Read every string operation as producing a value
Why: Which is lesson 6a's fruitful-function point: the result has to be used, or the call was pointless.
Assign the result, usually back to the same name
Why: word = word.upper() rather than word.upper() on its own.
This becomes the standard shape for string work: build a new string from pieces of the old one and assign it. Chapter 10's lists behave differently, and the contrast is one of the most important in the book.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C confuses two different things. The VALUE cannot be changed; the NAME can be repointed at any time. greeting = 'goodbye' is perfectly legal and creates no conflict with immutability, because it changes what the name refers to rather than changing any string. Lesson 7a's state diagram is the picture that keeps these apart: immutability is about the box on the right of the arrow, and reassignment moves the arrow.
Faded example
Keep everything except the last character, then add the replacement.
Fill in the blanks
word = 'banana'
new = word[:-1] + 'S' # should be 'bananS'
Why: word[:-1] takes everything up to but not including the last character, which is exactly what needs keeping. The positive equivalent is word[:len(word)-1], which needs the length. Note how the half-open convention helps here: the slice stops at the boundary before the last character, so no plus-or-minus-one adjustment is needed beyond the -1 itself.
Socratic
Lists are mutable. Ask what strings gain by not being.
Discussion prompt
Python's lists can be changed in place and its strings cannot. Suggest one advantage of immutability, and one thing it makes more awkward.
Hint: Think about two names referring to the same string.
Answer:
The advantage is safety from surprise: if two names refer to the same string, nothing either of them does can change what the other sees. Chapter 10 shows what happens when that guarantee is absent, and the section is called Aliasing for a reason.
It also lets Python share strings freely behind the scenes, since no one can modify a shared value.
What it makes awkward is building a string up piece by piece, since every step creates a new string. For long strings that is a real cost, and chapter 10's lists are the usual answer — build a list, then join it.
Section
Section 3
Concept
This pattern of computation — traversing a sequence and returning when we find what we are looking for — is called a search. The book's example is the inverse of the bracket operator.
def find(word, letter):
index = 0
while index < len(word):
if word[index] == letter:
return index
index = index + 1
return -1| Part | Its role | Note |
|---|---|---|
| the traversal | index from 0 to len-1 | the standard pattern |
| the test | is this the character? | inside the loop |
| return index | found: stop immediately | the first return |
| return -1 | the loop finished without finding | not found |
In a sense find is the inverse of the bracket operator: instead of taking an index and extracting the corresponding character, it takes a character and finds the index where that character appears. This is the first example of a return statement inside a loop.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 74-75
Picture it
A search has two exits, and each returns something different.
Figure (svg): A flow chart showing a search loop with an early return when the target is found and a return of minus one after the loop
The two exits are what make this a pattern rather than an ordinary traversal. An ordinary traversal always visits everything; a search stops as soon as it has an answer.
Worked example
It stops as soon as it finds the character. Count the passes.
>>> find('banana', 'n')
2| Pass | The test | What happens |
|---|---|---|
| index 0 | 'b' == 'n'? no | advance |
| index 1 | 'a' == 'n'? no | advance |
| index 2 | 'n' == 'n'? yes | return 2 immediately |
| indices 3-5 | never examined | the function has already returned |
Traverse from the start.
Why: The standard index traversal from the previous lesson, unchanged.
Return as soon as the test succeeds.
Why: If the character matches, the function breaks out of the loop and returns immediately — which is what the return statement does, from lesson 6a.
Notice what is skipped.
Why: Indices 3, 4 and 5 are never examined. A search does the minimum work needed to answer the question.
Figure (svg): The state of the program after each line of Worked example tracing a successful search, drawn as a ladder with one rung per traced line
2, after three passes. The second 'n' at index 4 is never reached, because find returns the position of the FIRST occurrence.
Verify: Search for a character that appears twice and check which position comes back.
Why: find('banana', 'a') gives 1, not 3 or 5. That the first occurrence wins is a consequence of returning immediately, and it is worth knowing because a search that needed the LAST occurrence would have to be written differently.
Prediction
The character appears more than once.
>>> find('banana', 'a')| Pass | What happens | Note |
|---|---|---|
| index 0 | 'b', no match | advance |
| index 1 | 'a', match | return 1 |
| indices 3 and 5 | also 'a' | never reached |
Predict first
What does find('banana', 'a') return?
Correct: 1 — the first occurrence, because the function returns immediately on finding a match.
Why: The a's are at indices 1, 3 and 5, and the search stops at the first. This is a consequence of the early return rather than a special rule, and it is worth being explicit about: a search returns the FIRST match. Finding the last one would need a different structure — traversing the whole string and remembering the most recent match.
Worked example
The second exit, and why it is after the loop rather than inside it.
>>> find('banana', 'z')
-1| Stage | What happens | Note |
|---|---|---|
| indices 0 to 5 | no match at any | the loop completes |
| after the loop | return -1 | the not-found value |
| why after | you only know it is absent once you have checked all | the last possible moment |
Let the loop finish.
Why: If the character does not appear in the string, the program exits the loop normally.
Return the sentinel.
Why: It returns -1, which is a value no valid index can be — so the caller can distinguish it from a real answer.
Notice why it cannot be inside the loop.
Why: You cannot conclude a character is absent until every position has been checked, so the not-found return must come after the traversal, not during it.
Figure (svg): Two columns comparing a successful search that returns early with an unsuccessful one that returns after the loop
-1, returned after the loop has examined every character without a match. The placement is forced: absence can only be established once the search is complete.
Verify: Check that -1 could not be a genuine answer.
Why: Valid indices are 0 and above, so -1 is unambiguous as a not-found marker — which is what makes it a usable sentinel. Compare with lesson 6c's boundary probe, where None was a bad sentinel precisely because it could also be a real result.
Trap
A student adds an else to the if, returning -1 when the current character does not match.
Treat each character as decisive
Why: It feels symmetric: match means found, no match means not found.
Now the function returns -1 after examining only the first character. It reports not found for every string whose first character is not the target, including strings that contain it.
A single non-match proves nothing; only an exhausted loop does.
Return the found value from inside the loop
Why: One match is enough to answer yes.
Return the not-found value after the loop
Why: Because you only know a character is absent once you have checked every position.
This asymmetry is the essence of the search pattern, and it recurs everywhere: proving something exists takes one example, and proving it does not takes an exhaustive check.
Discrimination
Ask what evidence each conclusion needs.
Sort into buckets
For a search, where does each of these belong?
Faded example
The book's exercise: let the caller say where to start looking.
Fill in the blanks
def find(word, letter, start):
index = start
while index < len(word):
if word[index] == letter:
return index
index = index + 1
return -1
Why: The only change is the initialization: instead of always beginning at 0, the traversal begins wherever the caller says. Everything else — the condition, the test, the two returns — is unchanged, which is a good sign that the search pattern is separable from where it starts. This three-parameter version is exactly what the next lesson's counting exercise uses to find repeated occurrences.
Explain it
The asymmetry is the interesting part.
Discussion prompt
A classmate wants to add an else that returns -1 when a character does not match. Explain why the two returns cannot both go inside the loop.
Hint: What does one non-matching character actually prove?
Answer:
Say: one match proves the character is there, but one non-match proves nothing at all — the character could be at any later position.
So the found-return can go inside, because a single piece of evidence settles it. The not-found return has to wait until every position has been checked, which is after the loop.
Give them the test that shows the bug: with the else in place, find('banana', 'n') returns -1, because the first character is a b. One test with the target anywhere but position zero exposes it immediately.
Section
Section 4
Concept
The second pattern of computation in this chapter is a counter: a variable initialized to zero and incremented each time something is found. Unlike a search, it always examines the whole sequence.
word = 'banana'
count = 0
for letter in word:
if letter == 'a':
count = count + 1
print(count)| Line | Its role | Note |
|---|---|---|
| count = 0 | initialized before the loop | the starting tally |
| for letter in word | traverse every character | no early exit |
| if letter == 'a' | the condition | decides whether to count |
| count = count + 1 | the update | only when the condition holds |
The variable count is initialized to 0 and then incremented each time an a is found. When the loop exits, count contains the result — the total number of a's.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 75-75
Picture it
Both traverse. Only one of them can leave early.
Figure (svg): Two columns comparing the search pattern with the counter pattern on when each stops and what each returns
That is why a counter uses a for loop and a search usually uses an index: the counter has no reason to stop early, and the search's early return is its whole point.
Worked example
Six characters, three matches. Watch the counter.
word = 'banana'
count = 0
for letter in word:
if letter == 'a':
count = count + 1| Character | Does it match? | count after |
|---|---|---|
| b | no match | count 0 |
| a | match | count 1 |
| n | no match | count 1 |
| a, n, a | two more matches | count 3 |
Initialize outside the loop.
Why: Lesson 7a's rule. Putting count = 0 inside the body would reset it every pass and the result would always be 0 or 1.
Traverse the whole string.
Why: Unlike a search, a counter has no reason to stop early: the answer depends on every character.
Increment conditionally.
Why: The update happens only when the condition holds, which is what distinguishes a counter from a length.
Figure (svg): The string banana with the three a characters highlighted as the ones that increment the counter
3. The counter examines all six characters and increments on three of them, and the answer is only complete once the loop has finished.
Verify: Check the result against a case with no matches and one with all matches.
Why: Counting 'z' in banana gives 0 and counting 'a' in 'aaa' gives 3, which equals the length. Both extremes behaving sensibly is a good check — a counter that can never reach 0 or can never reach len usually has its condition in the wrong place.
Prediction
The return is in the wrong place.
def count(word, letter):
total = 0
for c in word:
if c == letter:
total = total + 1
return total| Stage | What happens | Result |
|---|---|---|
| first character | 'b', no match | total stays 0 |
| the return | runs on the first pass regardless | returns 0 |
| the rest | never examined | the loop ends immediately |
Predict first
What does count('banana', 'a') return with this version?
Correct: 0 — the return is inside the loop and runs on the first pass, before any 'a' has been seen.
Why: The return is indented to the loop body but not to the if, so it runs on every pass — and the first pass examines 'b', which does not match, leaving total at 0. Note how the symptom depends on the first character: count('apple', 'a') would return 1 and look almost right, which is what makes this bug survive casual testing.
Worked example
The book's exercise, and it is lesson 4a's two process moves.
def count(word, letter):
total = 0
for c in word:
if c == letter:
total = total + 1
return total| Move | What changed | Benefit |
|---|---|---|
| encapsulation | wrap the working code in a function | it gains a name |
| generalization | the word and the letter become parameters | it works on anything |
| the return | hand the total back | rather than printing it |
Encapsulate the working code.
Why: Exactly lesson 4a's first process move: wrap code that already works, and give it a name.
Generalize with two parameters.
Why: The string and the letter both become parameters, so the function counts anything in anything.
Return rather than print.
Why: Lesson 6a's point: a function that prints its answer cannot be used inside a larger computation.
Figure (svg): The state of the program after each line of Worked example encapsulating and generalizing it, drawn as a ladder with one rung per traced line
A general counting function. Three moves from three earlier chapters — encapsulate, generalize, return — applied to six lines of working code.
Verify: Check that the counter variable was renamed and every use updated.
Why: The local is called total and the parameter is letter, and no occurrence of the original count or 'a' remains. An incomplete rename is the classic encapsulation bug from lesson 4a, and testing with a different letter — count('banana', 'n') giving 2 — catches it.
Trap
A student writes return total inside the loop, next to the increment.
Copy the shape of the search pattern
Why: Both traverse and both produce a number, so the structures look interchangeable.
The function returns after the first character, so the total is 0 or 1 — never the actual count. The two patterns differ in exactly this respect, and confusing them produces a plausible-looking wrong answer.
A count is only complete when the traversal is.
Put the return after the loop
Why: Every character must be examined before the total means anything.
Ask what evidence the answer needs
Why: A search needs one example; a count needs all of them. That question chooses between the two patterns.
The characteristic symptom is a count that is always 0 or 1, or one that equals 1 whenever the target appears at all. Either is a return that has escaped into the loop.
Discrimination
Ask whether one example settles the question.
Sort into buckets
For each task, which pattern applies?
Comparison
Fill the blanks. The differences all follow from one question.
Comparison matrix
| Question | Search | Counter |
|---|---|---|
| Can it stop early? | yes, at the first match | no — every element affects the answer |
| Where does it return? | inside the loop when found, after it when not | after the loop, always |
| What does it return? | a position, or -1 | a total |
| Which loop form? | usually while with an index, since the position is the answer | usually for, since only the characters matter |
The bottom row follows from the third: a search returns a position, so it needs one; a counter does not, so it can use the simpler loop.
Real world
They are not about strings.
Discussion prompt
Describe a search and a count you have performed on something physical — a shelf, a document, a list of names. What told you when to stop in each case?
Hint: One of them lets you stop early.
Answer:
Looking for a particular book on a shelf is a search: you stop the moment you find it, and you only conclude it is absent after checking every shelf.
Counting how many books are by one author is a counter: you have to look at all of them, and you cannot know the answer partway.
The transferable insight is the asymmetry: existence takes one example and absence takes an exhaustive check. That is why a failed search is more work than a successful one, in code and on a shelf alike.
Section
Section 5
Concept
Real string processing combines these. A typical function traverses, tests, and builds a new string from slices — because it cannot modify the original.
def remove_letter(word, letter):
result = ''
for c in word:
if c != letter:
result = result + c
return result| Line | Its role | Which idea it uses |
|---|---|---|
| result = '' | an accumulator, initialized empty | the counter pattern's shape |
| for c in word | traverse everything | no early exit |
| if c != letter | the condition | decides what to keep |
| result = result + c | build a NEW string | because strings are immutable |
The shape is the counter pattern with a string accumulator instead of a number. Nothing is modified — every pass creates a new string and points result at it, which is what immutability requires.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 74-75
Picture it
Same shape as a counter, with concatenation instead of addition.
Figure (svg): Two columns comparing a numeric counter with a string accumulator, showing that the shape is identical
Recognising that these are the same pattern is worth more than either one alone: an accumulator can hold a number, a string, or in chapter 10 a list, and the shape never changes.
Worked example
An accumulator, filled in the opposite order.
def reverse(word):
result = ''
for c in word:
result = c + result
return result| Pass | What is computed | result after |
|---|---|---|
| c = 'b' | 'b' + '' | 'b' |
| c = 'a' | 'a' + 'b' | 'ab' |
| c = 'n' | 'n' + 'ab' | 'nab' |
| ... | each new character goes in FRONT | eventually 'ananab' |
Notice the order of the concatenation.
Why: c + result puts the new character in front rather than behind, which is the only difference from a copying loop.
Trace two passes to see the effect.
Why: After 'b' and then 'a', result is 'ab' — the characters are accumulating in reverse order relative to the traversal.
Notice what is not needed.
Why: No index, no length, and no slicing. The reversal comes entirely from which side the concatenation happens on.
Figure (svg): A ladder showing the accumulator growing from empty to the reversed string, one character at a time
'ananab'. Putting each character in front of the accumulator rather than behind it reverses the string, with no index arithmetic at all.
Verify: Check it against a slice-based reversal.
Why: Python can also reverse with word[::-1], a slice with a step, which the book does not cover here. Both give 'ananab', and agreeing with an independent method is a stronger check than tracing — which is lesson 3a's verification advice applied to a string.
Prediction
An accumulator with a conditional.
result = ''
for c in 'banana':
if c != 'a':
result = result + c
print(result)| Character | Kept or skipped | result |
|---|---|---|
| b | not 'a', keep it | 'b' |
| a | skipped | 'b' |
| n | kept | 'bn' |
| a, n, a | one kept, two skipped | 'bnn' |
Predict first
What does this print?
Correct: 'bnn' — every character except the a's, in their original order.
Why: The condition keeps characters that are not 'a', so the three a's are skipped and the b and two n's survive. Note that the original string is untouched throughout: each pass creates a new string and points result at it, because strings are immutable and there is no other way to do it.
Worked example
Find a position, then use it to cut.
def before_first(word, letter):
i = find(word, letter)
if i == -1:
return word
return word[:i]| Line | What it does | Which idea |
|---|---|---|
| find | returns a position or -1 | the search pattern |
| i == -1 | not found: return the whole string | the guardian |
| word[:i] | everything before that position | the slice |
Use the search to get a position.
Why: This is why find returns an index rather than True or False: the position is what the next step needs.
Handle the not-found case first.
Why: A guardian, from lesson 6c. Without it, word[:-1] would silently drop the last character — a wrong answer rather than an error.
Slice up to the position.
Why: word[:i] takes everything before index i, which under the boundary reading is everything before that character.
Figure (svg): The state of the program after each line of Worked example combining a search with a slice, drawn as a ladder with one rung per traced line
The part of the word before the first occurrence, or the whole word if the letter is absent. Three ideas from this lesson, composed.
Verify: Test the not-found case specifically.
Why: before_first('banana', 'z') returns 'banana'. Without the guardian it would return word[:-1], which is 'banan' — a plausible-looking wrong answer. That -1 is a valid index in Python is exactly why the sentinel needs checking rather than being used directly.
Trap
A program calls find and uses the result as a slice index without checking for -1.
Assume the target is present
Why: In testing it always was, and the code is shorter without the check.
When it is absent, find returns -1 — which is a perfectly valid negative index. The slice silently means up to the last character instead of raising an error, so the wrong answer travels onward.
Check the sentinel before using the result.
Test for -1 immediately after the search
Why: One if statement, exactly as lesson 6c's guardian pattern prescribes.
Decide what the not-found case should do
Why: Return the whole string, return an empty one, or report a problem — but decide, rather than letting -1 be interpreted as an index.
This is a particularly nasty case because the sentinel is a legal index. In a language without negative indexing it would crash; in Python it quietly means something else, which is worse.
Sorting
Four ideas from this lesson, appearing in real code.
Sort into buckets
For each line, which idea from this lesson is it using?
Reverse engineer
The input and output are given. Supply the test.
Fill in the blanks
# 'banana' became 'aaa'
result = ''
for c in 'banana':
if c == 'a':
result = result + c
Why: Keeping only the a's means the condition must be equality rather than inequality. The rest of the structure is the standard accumulator: initialize empty, traverse everything, concatenate conditionally. Reading backwards from output to condition is a genuinely useful skill — it is how you work out what a piece of code you did not write is filtering for.
Explain it
The reason is immutability, and it is worth making explicit.
Discussion prompt
A classmate asks why you cannot just modify the string as you go, adding and removing characters. Explain what forces the accumulator pattern, and what the equivalent would look like if strings could be changed.
Hint: The answer is one word, and then its consequence.
Answer:
Say: strings are immutable, so there is no operation that changes one. Every apparent modification produces a new string.
So the pattern is forced: start with an empty string, and on each pass create a new string that is the old one plus something. The name result is repointed each time, which is reassignment rather than modification.
If strings were mutable you could append in place, which is exactly what lists do in chapter 10 — and that chapter's first surprise is that mutability brings problems of its own, which is why the two types exist side by side.
Comparison
Fill the blanks. One question decides between them.
Comparison matrix
| Question | Search | Counter |
|---|---|---|
| What is it asking? | where is it, or is it there? | how many are there? |
| How much does it examine? | as little as possible — it stops at the first match | all of it, always |
| How many exits? | two: found and not found | one, after the loop |
| What settles the answer? | a single example | every element |
The bottom row is the one to remember. Existence takes one example; a total takes all of them — and that determines everything else about the structure.
Pattern
Six steps, and the first is the one that chooses the shape.
Step 5 is the one that produces the nastiest bug in this lesson, because the failure is silent: an unchecked -1 slices from the end rather than raising an error.
Python documentation — An Informal Introduction to Python An Informal Introduction to Python
Check
The length is the difference of the indices.
>>> s = 'banana'
>>> s[2:5]| Marker | Where | Note |
|---|---|---|
| start | boundary 2 | just after 'ba' |
| stop | boundary 5 | just before the last 'a' |
| length | 5 - 2 | 3 characters |
Check your understanding
What is s[2:5] when s is 'banana'?
Answer: B
Why: The slice includes the character at index 2 and excludes the one at index 5, giving three characters: n, a, n. The subtraction is the fastest check — 5 minus 2 is 3, so any answer with a different number of characters is wrong before you even work out which ones.
Check
The error is about the type, not the syntax.
Check your understanding
Why does s[0] = 'H' fail for a string?
Answer: B
Why: Strings are immutable, so there is no way to change an existing one. The message says a str object does not support item assignment — it is the type that refuses, not the syntax. Lists in chapter 10 accept exactly this line, which is what makes the distinction worth learning now.
Check
The two returns are in different places for a reason.
def find(word, letter):
index = 0
while index < len(word):
if word[index] == letter:
return index
index = index + 1
return -1| Case | Where the return is | Why |
|---|---|---|
| found | return from inside the loop | one match is enough |
| not found | return after the loop | needs an exhaustive check |
| why | absence cannot be shown from one position | the asymmetry |
Check your understanding
Why is return -1 placed after the loop rather than inside it?
Answer: B
Why: A single non-matching character proves nothing — the letter could be at any later position. Only an exhausted loop establishes absence, so the not-found return must come after the traversal. A single match, by contrast, settles the question immediately, which is why that return can be inside.
Real world
Half-open intervals are everywhere once you notice them.
Discussion prompt
Find a real-world range that includes its start and excludes its end — an opening time, a booking, a page range. Why is that convention used, and what goes wrong when people assume the end is included?
Hint: Think about two consecutive bookings for the same room.
Answer:
Room bookings are the clearest: a 9-to-10 slot and a 10-to-11 slot do not overlap, because the first excludes its end. If both included their ends, both would own 10 o'clock.
That is exactly the property of adjacent slices: s[0:5] and s[5:12] join with no overlap and no gap, sharing the boundary index.
What goes wrong when people assume inclusion is double-counting — the same hour billed twice, or the same character appearing in two slices. The half-open convention exists to make adjacency work, and the arithmetic falling out cleanly is the same fact seen from the other side.
Commit first
Answer, then rate your confidence. This one has a silent failure mode.
Predict first
A program calls find, which returns -1 when the letter is absent, and uses the result directly as word[:i]. What happens when the letter is absent?
Correct: It returns the word with its last character removed, because word[:-1] is a perfectly valid slice meaning everything except the last character.
Why: This is the nastiest bug in the lesson. In a language without negative indexing, -1 would be out of range and the program would crash — which would at least tell you something. In Python it is a legal index counting from the end, so the slice silently means something quite different from what was intended, and a plausible-looking wrong string travels onward. The fix is a guardian: check for -1 before using the result, and decide explicitly what the not-found case should do.
Explain it
The half-open slice is the thing beginners most want justified.
Discussion prompt
A classmate is annoyed that s[0:5] gives five characters rather than six. Explain the convention using the book's picture, and give them one concrete thing it makes easier.
Hint: The picture is about boundaries, not characters.
Answer:
Say: imagine the index markers pointing BETWEEN the characters rather than at them. Then s[0:5] is everything between boundary 0 and boundary 5, which is five characters — no rule to remember, just the gap between two marks.
What it makes easier: the length of a slice is the difference of its indices, with no plus one anywhere. And two adjacent slices share a boundary, so s[:5] and s[5:] join perfectly with nothing repeated or missing.
Then give them the everyday parallel: a 9-to-10 booking and a 10-to-11 booking do not both own 10 o'clock. Half-open intervals are how ranges are made to fit together, in code and out of it.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: Slicing settles once the boundary picture clicks, and until it does every slice is a guess — so it is worth drawing figure 8.1 by hand. Immutability is a single fact with a large consequence, and it becomes vivid in chapter 10 when lists turn out not to share it. The search pattern's asymmetry is a genuinely general idea — existence takes one example, absence takes all of them — and it will recur for the rest of your programming life. The counter is the easiest of the four and the one whose bug is quietest, so the habit of initializing outside the loop is worth over-learning.
Connect it up
One page, from memory.
Draw it
Draw a six-character string with the boundary indices marked between the characters, and use it to show three slices: one from the start, one to the end, and one in the middle. Then, beside it, write the search pattern and the counter pattern side by side as skeletons, and mark on each where its return statement goes. Finally, write one sentence saying what forces a string-building loop to use an accumulator rather than modifying the original.
Recap
Three pages, and most of what string processing consists of.
| If you remember one thing | It is this |
|---|---|
| From slices | The indices point between the characters. The length is the difference. |
| From immutability | Nothing modifies a string. Every operation produces a new one. |
| From the search | One match proves presence; only an exhausted loop proves absence. |
| From the counter | Initialize outside the loop, update inside, return after. |
| From the sentinel | -1 is a valid index. An unchecked not-found result fails silently. |
The next lesson finishes the chapter with the methods a string already provides — including a find far more capable than the one you wrote — the in operator, and what happens when you compare two strings with a less-than sign.
Think Python, 2nd edition — Allen B. Downey §8.4-8.7, pp. 73-75 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.