This lesson reclassifies the string as a sequence of characters, introduces the bracket operator and zero-based indexing, meets the IndexError at the end of a string, and gives the two ways to traverse a string one character at a time.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 8 — Strings
§8.1-8.3, pp. 71-72
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-72 — the pages these objectives are drawn from
Warm-up
You have treated a string as one value since lesson 1a. Question that.
Discussion prompt
The value 'banana' has a type, str, and behaves as a single value in an assignment. But it obviously has parts. Name two operations you would want on those parts that nothing you know so far can do.
Hint: Getting one letter out, and going through them one at a time.
Answer:
Getting a single character out of it, and processing every character in turn. Neither is possible with what you have — concatenation and repetition treat a string as a lump.
The reason is that you have been treating a string as an atom, like an int. It is not.
Strings are not like integers, floats and booleans. A string is a sequence, which means it is an ordered collection of other values — and this chapter is what follows from that.
Concept
A string is a sequence of characters. That means it has parts, the parts are in a definite order, and every part has a position you can name — which is what the bracket operator is for.
sequence — An ordered collection of values, in which each value is identified by an integer index.
The word ordered is doing work. A string is not merely a bag of characters: 'banana' and 'nabana' contain the same letters and are different strings, exactly as lesson 2b's concatenation was not commutative.
Figure (svg): The string banana drawn as six boxes with the indices zero to five beneath them
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-71
Section
Section 1
Concept
You can access the characters one at a time with the bracket operator. The expression in brackets is called an index, and it indicates which character in the sequence you want.
>>> fruit = 'banana'
>>> letter = fruit[1]
>>> letter
'a'| Expression | What it does | Result |
|---|---|---|
| fruit[1] | select character number 1 | 'a', not 'b' |
| fruit[0] | select character number 0 | 'b' |
| why | the index is an OFFSET from the beginning | the first has offset zero |
For most people the first letter of banana is b, not a. But for computer scientists the index is an offset from the beginning of the string, and the offset of the first letter is zero. So b is the zero-eth letter of banana, a is the one-eth, and n is the two-eth.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-71
Picture it
The index answers how far from the start, not which one in order.
Figure (svg): The string banana with each character labelled by its offset from the beginning
Once you read the index as a distance rather than a position in a queue, zero-based counting stops being arbitrary — the first item really is zero steps along.
Worked example
Anything with an integer value can go in the brackets.
>>> fruit = 'banana'
>>> i = 1
>>> fruit[i]
'a'
>>> fruit[i+1]
'n'| Expression | What happens | Result |
|---|---|---|
| fruit[i] | i is looked up: 1 | 'a' |
| fruit[i+1] | the expression is evaluated: 2 | 'n' |
| the rule | any integer-valued expression works | composition again |
Use a variable as an index.
Why: As an index you can use an expression that contains variables and operators — the same composition rule from lesson 3a.
Use an arithmetic expression.
Why: The expression inside the brackets is fully evaluated before the character is selected, exactly as an argument is evaluated before a call.
Note the requirement.
Why: The value of the index has to be an integer. A float raises a TypeError, even one like 1.5 whose meaning seems obvious.
Figure (svg): The state of the program after each line of Worked example the index is an expression, drawn as a ladder with one rung per traced line
Both work, giving 'a' and 'n'. Anything that evaluates to an integer can be an index, which is what makes loops over indices possible.
Verify: Try a float index and read the error.
Why: fruit[1.5] gives a TypeError saying string indices must be integers. The message names the requirement directly — there is no character halfway between two characters, so the operation has no meaning for a fractional index.
Prediction
Count offsets, starting from zero.
>>> fruit = 'banana'
>>> fruit[2]| Index | Character | Which one it is |
|---|---|---|
| offset 0 | b | the first character |
| offset 1 | a | the second |
| offset 2 | n | the third |
Predict first
What is fruit[2] when fruit is 'banana'?
Correct: 'n' — offsets 0, 1 and 2 are b, a and n, so index 2 is the third character.
Why: The book's own phrasing helps here: n is the two-eth letter of banana. Note that a single index always produces a single character, never a group — getting two characters needs a slice, which is the next lesson. If you answered 'a', you were reading the index as a count rather than an offset.
Worked example
Compare the two conventions on the arithmetic they produce.
# zero-based: the index is an offset
# first character: fruit[0]
# n-th character: fruit[n-1]... no: fruit[n] is the (n+1)th
# slice of first 3: fruit[0:3]
# one-based would need:
# slice of first 3: fruit[1:3] -- or 1:4? ambiguous| Convention | What an index means | Consequence |
|---|---|---|
| zero-based | index = distance from the start | arithmetic is clean |
| length and index | the last index is len - 1 | one subtraction |
| slicing | n:m is m minus n characters | no adjustment needed |
Read the index as a distance.
Why: The offset of the first letter is zero because it is at the beginning — no steps along.
Notice what that buys.
Why: The number of characters between two indices is simply the difference, with no adjustment. That becomes important in the next lesson, where slicing relies on it.
Notice what it costs.
Why: The last character's index is len minus one rather than len, which is the source of the most common error in this chapter.
Figure (svg): Two columns comparing zero-based indexing with a hypothetical one-based scheme on the arithmetic each produces
Zero-based indexing makes distances and slice lengths come out as simple subtractions, at the cost of a len-minus-one adjustment at the end. It is a trade, and the arithmetic is why it is the near-universal choice.
Verify: Count the characters in fruit[1:4] and compare with 4 minus 1.
Why: Three characters, and 4 - 1 is 3. Under a one-based convention that subtraction would need a correction, which is exactly the sort of adjustment that produces off-by-one errors.
Trap
A student reads fruit[1] as the first letter and is surprised to get 'a'.
Import the everyday meaning of first
Why: In ordinary counting the first thing is number one, and the index looks like a count.
Every index is then off by one, and the error is systematic rather than occasional — it affects every access in the program.
Read the index as an offset: how far along, not which one.
Say the character at offset 1 rather than the first character
Why: Offset 1 is one step from the beginning, which is the second character.
Use the book's ordinals if it helps
Why: b is the zero-eth letter, a is the one-eth, n is the two-eth. Deliberately odd, and it fixes the reading.
The reframing is worth the effort because it also explains slicing, negative indices and the len-minus-one rule, all of which follow from distance rather than from position.
Sorting
The index must be an integer, and it must be in range.
Sort into buckets
For fruit = 'banana', sort each expression.
Faded example
Offsets start at zero.
Fill in the blanks
word = 'python'
third = word[2] # should be 't'
Why: The characters are at offsets 0 (p), 1 (y) and 2 (t), so the third character is at index 2. The general rule follows: the n-th character, counting from one as people do, is at index n minus one. That single subtraction is where zero-based indexing costs you something, and it is worth stating explicitly rather than rediscovering each time.
Socratic
1.5 has an obvious meaning to a person. Say why it does not to Python.
Discussion prompt
fruit[1.5] raises a TypeError. What would it even mean, and why is refusing better than rounding?
Hint: There is nothing between two adjacent characters.
Answer:
It has no meaning: there is no character halfway between two characters, so no value could be returned that is not an invention.
Rounding would be a guess — and a silent one. fruit[1.5] would give either 'a' or 'n' depending on the rounding rule, and neither is what anybody meant.
This is the same design decision as refusing to add a str to an int in lesson 1b: where there is no obviously correct answer, Python declines rather than guessing. A TypeError at the point of the mistake is much cheaper than a wrong character travelling onward.
Section
Section 2
Concept
len is a built-in function that returns the number of characters in a string. Using it to reach the last character is where nearly everybody gets caught once.
>>> fruit = 'banana'
>>> len(fruit)
6
>>> last = fruit[len(fruit)]
IndexError: string index out of range| Expression | What happens | Result |
|---|---|---|
| len(fruit) | the number of characters | 6 |
| fruit[6] | there is no character at offset 6 | IndexError |
| fruit[5] | the valid offsets are 0 to 5 | 'a', the last character |
The reason for the IndexError is that there is no letter in banana with the index 6. Since we started counting at zero, the six letters are numbered 0 to 5. To get the last character you have to subtract 1 from the length — or use a negative index.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-72
Picture it
The count and the last index differ by one, always.
Figure (svg): The string banana with positive indices zero to five above and negative indices minus one to minus six below
Negative indices count backward from the end: fruit[-1] is the last letter, fruit[-2] the second to last, and so on. They exist precisely so you do not have to write len minus one.
Worked example
One arithmetic, one direct. Both are used constantly.
>>> fruit = 'banana'
>>> length = len(fruit)
>>> fruit[length-1]
'a'
>>> fruit[-1]
'a'| Expression | How it locates the character | Result |
|---|---|---|
| len(fruit) | 6 | the count |
| fruit[6-1] | index 5 | the last character |
| fruit[-1] | counting backward from the end | the same character |
Compute the last index from the length.
Why: Since counting started at zero, the six letters are numbered 0 to 5, so the last index is length minus one.
Use a negative index instead.
Why: The expression fruit[-1] yields the last letter, fruit[-2] the second to last, and so on.
Notice why negative indices exist.
Why: They remove the subtraction, and with it the commonest place to make an off-by-one error.
Figure (svg): The state of the program after each line of Worked example two ways to reach the last character, drawn as a ladder with one rung per traced line
Both give 'a'. The negative index is shorter and safer; the length-minus-one form is what you need when the index is being computed for other reasons anyway.
Verify: Check the relationship between the two schemes.
Why: fruit[-1] and fruit[len(fruit)-1] name the same character, and in general fruit[-n] is fruit[len(fruit)-n]. That correspondence is worth knowing, because it means a negative index is not a separate feature but a shorthand.
Prediction
Six characters, and two indexing schemes.
Predict first
For fruit = 'banana', which of these raises an IndexError?
Correct: fruit[6] — the valid forward indices are 0 to 5, and 6 is one past the end.
Why: A six-character string has offsets 0 through 5 forwards and -1 through -6 backwards, so fruit[-6] is the first character and is perfectly valid. The one that fails is the index equal to the length, which is exactly the off-by-one the section is about. Note that fruit[-7] would also fail, for the same reason at the other end.
Worked example
The message names the problem precisely. Extract what it tells you.
>>> fruit = 'banana'
>>> fruit[6]
IndexError: string index out of range| Part | What it tells you | Next step |
|---|---|---|
| the error name | IndexError | the index was the problem |
| the message | string index out of range | not the type, the value |
| what to check | len versus the index used | 6 versus valid 0 to 5 |
Read the error kind.
Why: IndexError, not TypeError. The index was of the right type and the wrong value, which is the ValueError-versus-TypeError distinction from lesson 6c applied here.
Read the message.
Why: Out of range — so the fix is arithmetic rather than conversion.
Find the arithmetic.
Why: Compare the index used against len minus one. Almost every IndexError in this chapter is a missing subtraction of one.
Figure (svg): A panel showing an IndexError with the valid index range and the index that was used
The error says the index was an integer that named no character. The fix is nearly always to subtract one, or to use a negative index instead.
Verify: Check what a negative index too large in magnitude does.
Why: fruit[-7] also raises an IndexError, because there are only six characters to count back through. Negative indices remove one off-by-one and do not remove the range check — they still have to name a character that exists.
Trap
A student writes s[len(s)] to reach the last character, reasoning that a six-character string should have a sixth character.
Treat the length as the last position
Why: It is the count, and counts and last positions coincide under one-based numbering.
Under zero-based numbering they never coincide. len(s) is always one past the end, so this always raises an IndexError — which is at least loud.
The last index is len minus one, and there is a shorter way to say it.
Write s[-1] when you want the last character
Why: It needs no arithmetic and cannot be off by one.
Write s[len(s)-1] when you already have the length for another reason
Why: For instance inside a loop that counts down.
It is worth noticing that this error is loud: an index one past the end raises immediately. The quiet version is a loop condition that stops one short, which produces a wrong answer instead — and that is the subject of the next section.
Matching
Forwards and backwards name the same characters.
Match the pairs
Why: Each character has two names, one positive and one negative, and they differ by the length: index i forwards is index i minus len backwards. Knowing that the two schemes coincide is what makes negative indices feel like a shorthand rather than a separate mechanism.
Fill the middle
Two ways to write it. Give the negative one.
Fill in the blanks
word = 'python'
second_last = word[-2] # should be 'o'
Why: fruit[-1] is the last and fruit[-2] the second to last, so -2 gives 'o' from 'python'. The equivalent positive form is word[len(word)-2], which is longer and involves a subtraction that can go wrong. Negative indices exist for exactly this: reaching things relative to the end without computing the length.
Explain it to yourself
State it in terms of offsets rather than as a rule to memorise.
Discussion prompt
Explain, using the word offset, why a string of length 6 has no character at index 6 — without saying because indexing starts at zero.
Hint: How far along is the last character?
Answer:
The last character is five steps from the beginning, because the first is zero steps along. Six steps would take you past the end entirely.
So len counts characters and an index measures distance, and those are different quantities that happen to be related by one.
Framing it this way makes the fix obvious rather than memorised: the largest distance you can travel and still land on a character is one less than the number of characters, which is len minus one.
Section
Section 3
Concept
A lot of computations involve processing a string one character at a time. Often they start at the beginning, select each character in turn, do something to it, and continue until the end. This pattern of processing is called a traversal.
traversal — Processing each element of a sequence in turn, from one end to the other.
index = 0
while index < len(fruit):
letter = fruit[index]
print(letter)
index = index + 1| Line | Its role | Note |
|---|---|---|
| index = 0 | initialization | start at the beginning |
| index < len(fruit) | the condition | stops one past the last index |
| fruit[index] | select the current character | the work |
| index = index + 1 | the update | move along |
The loop condition is index less than len(fruit), so when index is equal to the length the condition is false and the body does not run. The last character accessed is the one with the index len minus one, which is the last character in the string.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-72
Picture it
One pass per character, with the index taking each valid value in turn.
Figure (svg): The string banana with the index shown moving across each character in turn
Notice the four parts of a loop from lesson 7a, all present: initialize the index, test it against the length, do the work, update the index.
Worked example
One character in the condition decides whether the loop is right.
# correct
while index < len(fruit):
# one too far
while index <= len(fruit):| Condition | The final pass | Outcome |
|---|---|---|
| index < len | last pass has index = len - 1 | the last character |
| index <= len | last pass has index = len | IndexError |
| the rule | the largest valid index is len - 1 | so the test is strict |
Work out the last pass under the correct condition.
Why: The condition fails when index equals len, so the last pass has index equal to len minus one — which is the last valid index.
Work out the last pass under the wrong one.
Why: With less-than-or-equal, the loop runs one more time with index equal to len, and fruit[len] raises an IndexError.
Notice which mistake is louder.
Why: This one crashes, which is helpful. The quieter mistake is a condition that stops one short, which silently skips the last character.
Figure (svg): Two columns comparing a traversal condition using less-than with one using less-than-or-equal
Strict less-than is correct: the loop must stop when index reaches len, because len is one past the last valid index. Using less-than-or-equal runs one pass too many and raises an IndexError.
Verify: Count the passes and compare with the length.
Why: The loop runs exactly len times, once per character. Whenever a traversal produces the wrong number of passes, the condition is where to look — and comparing the pass count with the length settles it in one step.
Prediction
Count them, then compare with the length.
fruit = 'banana'
index = 0
while index < len(fruit):
print(fruit[index])
index = index + 1| Index values | The condition | Result |
|---|---|---|
| index = 0 to 5 | the condition is true | six passes |
| index = 6 | 6 < 6 is false | the loop ends |
| output | one line per character | six lines |
Predict first
How many lines does this print?
Correct: Six — one per character, with the index taking every value from 0 to 5.
Why: The condition is true for indices 0 through 5 and false at 6, so the loop runs exactly len times. That equality — passes equals length — is the check to run on any traversal: five lines would mean the condition stops one short, and seven would mean an IndexError.
Worked example
The book sets this as an exercise. The changes are all in the loop's three parts.
index = len(fruit) - 1
while index >= 0:
print(fruit[index])
index = index - 1| Part | What changes | Value |
|---|---|---|
| initialization | start at the last valid index | len - 1 |
| condition | keep going while the index is still valid | index >= 0 |
| update | move toward the beginning | index - 1 |
Change where it starts.
Why: The last valid index is len minus one, which is where a backward traversal begins.
Change the condition.
Why: It must remain true for index 0, since that is a real character, so the test is greater-than-or-equal rather than strictly greater.
Change the direction of the update.
Why: Decrement rather than increment, so the index moves toward the condition failing.
Figure (svg): The state of the program after each line of Worked example traversing backwards, drawn as a ladder with one rung per traced line
All three parts change: the initialization, the condition, and the direction of the update. The body is untouched, which is a good sign that the traversal pattern is genuinely separable from the work.
Verify: Check the boundary at both ends.
Why: The first pass uses index len-1 and the last uses index 0, so both end characters are visited — and index -1 is never used, because the condition stops first. Checking both ends is the standard way to validate a traversal, since off-by-one errors live at exactly those two places.
Trap
A student writes the condition as index < len(fruit) - 1, reasoning that the last valid index is len minus one.
Confuse the last valid INDEX with the stopping point
Why: The last valid index really is len minus one, so putting it in the condition feels right.
The condition is false when index equals len minus one, so the loop stops before processing the last character. The program runs perfectly and silently produces an answer that is one item short.
The condition names the first index that is NOT to be processed.
Stop at len, not at len minus one
Why: index < len is true for every valid index including the last, and false for the first invalid one.
Check by counting passes
Why: A traversal of a six-character string should run six times. Five means the condition stops one short.
This is the quiet version of the error. Going one too far raises an IndexError immediately; stopping one short produces a wrong answer with no message at all, which makes it much more expensive.
Invariant
Step through and track both the index and the character it selects.
Step through it
At which index does the loop stop, and why is no character accessed at that index?
It stops at 6, which is len — one past the last valid index. Nothing is accessed there because the condition is tested BEFORE the body, which is lesson 7a's flow of execution doing exactly the job it exists for.
Faded example
The four parts of a loop, with two blanks.
Fill in the blanks
index = 0
while index < len(word):
print(word[index])
index = index + 1
Why: A forward traversal starts at index 0, the first valid offset, and increments toward the condition failing at len. Getting either blank wrong has a characteristic symptom: starting at 1 silently skips the first character, and decrementing produces an immediate IndexError at index -1... except that -1 is a valid index in Python, so it would silently traverse backwards forever. That second case is worth noticing.
Edge cases
A traversal of nothing should do nothing. Check that it does.
Discussion prompt
What does the while-loop traversal do when fruit is the empty string? Trace it, and say which property of the while statement makes the answer correct.
Hint: What is len of the empty string?
Answer:
len is 0, so the condition 0 < 0 is false on the very first test and the body never runs. Nothing is printed and no error occurs.
That is exactly right: traversing an empty sequence should do nothing at all.
The property that makes it work is that a while loop tests before the first pass, which is lesson 7a's socratic probe answered concretely. A loop that tested afterwards would try to access fruit[0] of an empty string and raise an IndexError — so the empty case is handled for free rather than needing a special check.
Section
Section 4
Concept
Another way to write a traversal is with a for loop. Each time through the loop, the next character in the string is assigned to the variable, and the loop continues until no characters are left.
for letter in fruit:
print(letter)| Pass | What letter holds | Output |
|---|---|---|
| pass 1 | letter is 'b' | print b |
| pass 2 | letter is 'a' | print a |
| ... | each character in turn | ... |
| after the last | no characters left | the loop ends |
Notice what has disappeared: there is no index, no initialization, no condition and no update. The loop variable holds the character itself rather than its position, and the loop ends when the string does.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-73
Picture it
Five lines become two, and every off-by-one opportunity disappears.
Figure (svg): Two columns comparing a while traversal using an index with a for traversal over the characters
Use the for version unless you actually need the index — for instance to look at two positions at once, which is what lesson 9b is about.
Worked example
The book's example, and it combines a for loop with concatenation.
prefixes = 'JKLMNOPQ'
suffix = 'ack'
for letter in prefixes:
print(letter + suffix)| Pass | What is computed | Output |
|---|---|---|
| letter = 'J' | 'J' + 'ack' | Jack |
| letter = 'K' | 'K' + 'ack' | Kack |
| ... | one per prefix | ... |
| letter = 'Q' | 'Q' + 'ack' | Qack |
Notice the loop variable holds a character.
Why: Each time through the loop, the next character in the string is assigned to letter. It is a one-character string, which can be concatenated like any other.
Notice the concatenation.
Why: letter + suffix builds a new string on every pass. Nothing is modified — a new string is created each time.
Notice the count.
Why: Eight prefixes, eight lines of output. The loop runs once per character with no index arithmetic anywhere.
Figure (svg): A diagram showing each character of the prefix string being concatenated with the suffix to produce a name
Eight names, from Jack to Qack, one per character of the prefix string. The whole traversal is two lines.
Verify: Check the output against the book's list of ducklings.
Why: It gives Oack and Quack as Oack and Qack, which the book notes are misspelled — the real names are Ouack and Quack. That the program is correct and the OUTPUT is wrong is a nice small example of a semantic error: the code does exactly what it says and not what was wanted.
Prediction
Not a number.
for letter in 'abc':
print(letter)| Pass | What letter holds | Output |
|---|---|---|
| pass 1 | letter is 'a' | a |
| pass 2 | letter is 'b' | b |
| pass 3 | letter is 'c' | c |
Predict first
What does this print?
Correct: a, b, c — each on its own line, because the loop variable holds each character in turn.
Why: The for loop over a string assigns the next character to the variable on each pass, not the next index. That is the essential difference from the for-in-range form of chapter 4, and it is why the loop needs no length and no arithmetic. Each value is a one-character string, so it can be printed, concatenated or compared like any other string.
Worked example
The for loop hides the index. Sometimes that is a loss.
# for: you have the character, not its position
for letter in word:
print(letter)
# while: you have both
index = 0
while index < len(word):
print(index, word[index])
index = index + 1| Form | What you have access to | Verdict |
|---|---|---|
| for | the character only | simpler, safer |
| while | the character and its position | more to get wrong |
| choose | by whether you need the position | usually you do not |
Ask what the body needs.
Why: If it only needs each character, the for loop gives exactly that and nothing to get wrong.
Notice what the for loop does not give.
Why: The position. If you need to report where something was found, or to compare two positions, the index has to come from somewhere.
Choose accordingly.
Why: Prefer for; use while with an index when the position is genuinely part of the problem.
Figure (svg): The state of the program after each line of Worked example when you still need the index, drawn as a ladder with one rung per traced line
The for loop is shorter and safer and hides the index. Use it unless the position itself is needed — which is less often than beginners expect.
Verify: Look ahead to a case that needs the index.
Why: Lesson 8b's find function must return the position where a character was found, so it cannot use a plain for loop over characters. That is the standard reason to reach for an index, and it is worth knowing that it exists rather than assuming for always works.
Trap
A student writes for i in word and then uses word[i], expecting i to be a position.
Carry over the for-in-range form from chapter 4
Why: There the loop variable really was a number, so the habit is understandable.
Here i holds a character, so word[i] tries to index a string with a string and raises a TypeError — string indices must be integers.
The for loop over a string gives you the characters, not their positions.
Name the variable for what it holds
Why: letter or char, not i. The name is the cheapest defence against this confusion.
If you need positions, loop over a range instead
Why: for i in range(len(word)) gives indices, and word[i] then works — which is the bridge between the two forms.
The TypeError here is informative: string indices must be integers is the same message as fruit[1.5], and it tells you the bracket got something that was not a position.
Discrimination
Ask whether the body needs the position.
Sort into buckets
For each task, which traversal form is appropriate?
Translation
Match each part of the while version to what replaces it.
Match the pairs
Why: Three of the four parts vanish entirely and the fourth becomes a single line. What is really happening is that the for loop is a specialised construct for the traversal pattern, with the bookkeeping built in — which is why it cannot go out of range and the while version can.
Explain it
A question every beginner asks once and then never again.
Discussion prompt
A classmate asks whether they should use a for loop or a while loop to go through a string. Give them a one-sentence rule and one example of the exception.
Hint: The rule is about what the body needs.
Answer:
Say: use a for loop unless the body needs to know WHERE it is, in which case use an index.
The example: a function that reports the position of a character has to know the position, so a plain for loop over characters cannot do it.
Then add the reason it matters, which is not style: the for loop cannot go out of range, because it never computes an index. Choosing it eliminates a whole category of bug, which is a stronger argument than brevity.
Section
Section 5
Concept
A traversal has the same shape whatever it does: visit each character in turn, do something with it, and stop at the end. The work varies; the pattern does not — and neither do the three ways it fails.
# the pattern
for letter in word:
# do something with letter
# the same pattern with an index
index = 0
while index < len(word):
letter = word[index]
# do something with letter
index = index + 1| Failure | What causes it | Symptom |
|---|---|---|
| off by one at the start | index starts at 1 | the first character is missed |
| off by one at the end | condition uses len - 1 | the last character is missed |
| one too far | condition uses <= len | IndexError |
Two of those three failures are silent, and both are at the ends. That is why checking the first and last character is the standard test for a traversal — the middle almost never goes wrong.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-73
Picture it
The middle of a traversal is nearly always right.
Figure (svg): The string banana with the first and last characters highlighted as the places where traversal errors occur
Which gives a two-item test for any traversal: does it process the first character, and does it process the last? If both are yes, the middle almost certainly is too.
Worked example
The pattern with a counter in it, which lesson 8b formalises.
word = 'banana'
count = 0
for letter in word:
if letter == 'a':
count = count + 1
print(count)| Character | The test | count after |
|---|---|---|
| b | not 'a' | count stays 0 |
| a | matches | count becomes 1 |
| n | not 'a' | count stays 1 |
| a, n, a | two more matches | count becomes 3 |
Initialize the counter before the loop.
Why: Lesson 7a's rule: initialize outside, update inside. Putting count = 0 in the body would reset it on every pass.
Traverse with a for loop.
Why: The body only needs the characters, not their positions, so the for form applies.
Update conditionally.
Why: The counter is incremented only when the character matches, which is what makes it a count rather than a length.
Figure (svg): The state of the program after each line of Worked example writing a traversal that counts, drawn as a ladder with one rung per traced line
3 — there are three a's in banana. The traversal pattern supplies the loop, and the work inside it is a conditional increment.
Verify: Check against len, which counts every character.
Why: len('banana') is 6 and the count is 3, so the condition is genuinely filtering. If a counter ever equals the length, the condition is always true and is doing nothing — which is a quick check on any counting loop.
Error analysis
Each has one fault. Name it and say whether it is loud or silent.
Annotate
The loud failure is the lucky one. A traversal that quietly processes one character too few can survive for a long time.
Worked example
Two checks find nearly every traversal bug.
# does it process the FIRST character?
# count the a's in 'abc' -> should be 1
# does it process the LAST character?
# count the a's in 'cba' -> should be 1
# does it handle the empty case?
# count the a's in '' -> should be 0| Test input | What it exercises | What it catches |
|---|---|---|
| 'abc' | the target is first | catches a start-at-1 bug |
| 'cba' | the target is last | catches a stop-early bug |
| '' | nothing at all | catches a special-case bug |
Design a test where the answer depends on the first character.
Why: If the traversal starts at index 1, this test fails and every other test might pass.
Design one where it depends on the last.
Why: If the condition stops one short, only this test fails.
Add the empty case.
Why: It costs nothing and catches any code that assumes at least one character exists.
Figure (svg): Two columns listing three test inputs against the traversal bug each one detects
Three tests, each targeting one of the three failure modes, and each one input long. Between them they cover the places where traversals actually go wrong.
Verify: Ask why a test on 'banana' alone is insufficient.
Why: It has a's in the middle as well as at the end, so a traversal that skipped the first character would still find two and look plausible. A test only discriminates if the answer changes when the bug is present, which is the same argument as lesson 5b's chain-versus-separate-ifs.
Trap
A traversal is tested on 'banana' and works, so it is declared correct.
Test on the example from the book
Why: It is the string in front of you and it gives a plausible answer.
banana has a's in the middle, so a traversal that skips the first or last character still finds some and looks right. The test cannot distinguish a correct traversal from two broken ones.
Choose inputs where the answer changes if the bug is present.
Put the thing being looked for at the very start, and then at the very end
Why: Two inputs, each of which fails for exactly one of the two silent bugs.
Add the empty string
Why: It costs one test and catches everything that assumes a first character exists.
This is lesson 6a's advice about choosing a test case, applied to sequences: it is useful to know the right answer, and it is more useful for the right answer to be different when the code is wrong.
Prediction
The condition stops one short.
word = 'abc'
index = 0
while index < len(word) - 1:
print(word[index])
index = index + 1| Value | What it is | Consequence |
|---|---|---|
| len(word) - 1 | 3 - 1 | 2 |
| indices used | 0 and 1 | a and b |
| index 2 | 2 < 2 is false | c is never printed |
Predict first
What does this print?
Correct: a and b — the last character is missed, because the condition stops before index 2.
Why: The condition names the first index NOT to process, and here it names 2, which is a valid character. The confusion comes from the last valid index being len minus one: that is the last index to USE, not the place to stop. The correct condition is index < len(word), which is false only at the first invalid index.
Comparison
Fill the blanks. Both are correct; they differ in what they expose.
Comparison matrix
| Question | for letter in word | while with an index |
|---|---|---|
| What does the loop variable hold? | the character | the index; the character needs word[index] |
| Can it go out of range? | no — it never computes an index | yes, if the condition is wrong |
| How many lines? | two | five |
| When do you need it? | almost always | when the position itself is part of the problem |
The second row is the strongest argument. A construct that cannot make a whole class of error is worth preferring even when both would work.
Real world
They are not a Python problem.
Discussion prompt
Think of an off-by-one error outside programming — a fencepost, a schedule, a count of days. Describe it, and say what makes this family of mistake so persistent.
Hint: How many posts does a fence with ten panels have?
Answer:
Eleven, not ten — the classic fencepost problem. Counting the gaps and counting the posts give answers that differ by one, and it is easy to answer the wrong question.
Dates do it too: from Monday to Friday is four days or five days depending on whether you count the endpoints.
What makes them persistent is that both answers are correct answers to slightly different questions, so neither looks obviously wrong. The programming defence is the same as the everyday one: state exactly which endpoints are included, and check the ends rather than the middle.
Comparison
Fill the blanks. All three name characters; they differ in where they count from.
Comparison matrix
| Form | Where it counts from | The last character |
|---|---|---|
| positive index | the beginning, from 0 | s[len(s)-1] |
| negative index | the end, from -1 | s[-1] |
| for loop | no counting at all | the last pass, automatically |
The third row is why the for loop is preferred: a form that does no counting cannot count wrongly.
Pattern
Six steps, and the last two are the tests that matter.
Step 5 is the one that finds real bugs. Two of the three traversal failures are silent and both are at the ends, so a test whose answer depends on a middle character proves nothing.
Python documentation — An Informal Introduction to Python An Informal Introduction to Python
Check
Offsets, counted from zero.
>>> word = 'python'
>>> word[1]| Index | Character | Note |
|---|---|---|
| offset 0 | p | the first character |
| offset 1 | y | the second |
| the rule | the index is a distance | not a count |
Check your understanding
What is word[1] when word is 'python'?
Answer: B
Why: The index is an offset from the beginning, and the offset of the first letter is zero — so index 1 is one step along, which is the second character. A single index always produces exactly one character; getting two would need a slice.
Check
The count and the last index differ by one.
>>> s = 'abc'
>>> len(s)
3
>>> s[3]
IndexError: string index out of range| Expression | What happens | Result |
|---|---|---|
| len | 3 characters | indices 0, 1, 2 |
| s[3] | one past the end | IndexError |
| s[2] or s[-1] | the last character | 'c' |
Check your understanding
Which expression gives the last character of s?
Answer: B
Why: Negative indices count backward from the end, so s[-1] is the last character. The equivalent positive form is s[len(s)-1], which needs a subtraction and is where the off-by-one usually creeps in.
Check
The loop variable holds a character.
for c in 'abc':
print(c)| Pass | What c holds | Output |
|---|---|---|
| pass 1 | c is 'a' | a |
| pass 2 | c is 'b' | b |
| pass 3 | c is 'c' | c |
Check your understanding
In a for loop over a string, what does the loop variable hold on each pass?
Answer: B
Why: Each time through the loop, the next character in the string is assigned to the variable, and the loop continues until no characters are left. That is what distinguishes it from the for-in-range form of chapter 4, where the variable really does hold a number.
Real world
Sequences and their indices are a shape you already use.
Discussion prompt
Find something in daily life that is an ordered collection with positions — a queue, a playlist, a street of houses. Does it number from zero or one, and where does that convention cause confusion?
Hint: Building floors are the classic case.
Answer:
Building floors are the best example: the floor at street level is the ground floor in Britain and the first floor in America, so the second floor means two different things — an off-by-one built into a language.
Playlists and page numbers count from one; distances and offsets count from zero. The pattern is that positions in a queue start at one and distances start at zero, which is exactly the distinction the book draws.
So Python is not being perverse. It has chosen the offset interpretation, and once you read s[0] as zero steps from the start it agrees with how distances work everywhere else.
Commit first
Answer, then rate your confidence. This is the off-by-one the whole chapter turns on.
Predict first
For a string s of length 6, which of these is the correct condition for a traversal that visits every character exactly once?
Correct: while index < len(s): — and also the last option, which is equivalent, though the second is the idiomatic form.
Why: The condition must be true for every valid index, which is 0 to 5, and false at the first invalid one, which is 6. Strict less-than against len does exactly that. Less-than-or-equal against len runs one pass too many and raises an IndexError; strict less-than against len minus one stops one short and silently skips the last character. The fourth option is arithmetically the same as the second and is written that way far less often, because the subtraction is an extra opportunity to be wrong.
Explain it
Zero-based indexing is the thing beginners most want explained.
Discussion prompt
A classmate is annoyed that s[1] is not the first character. Explain the convention in a way that makes it feel inevitable rather than arbitrary, and give them one thing it makes easier.
Hint: The word is offset.
Answer:
Say: the index is not a count, it is a distance from the beginning — and the first character is zero steps along, so its index is 0.
What it makes easier: the number of characters between two indices is just the difference, with no adjustment. s[2:5] holds 5 minus 2 characters, which is why slicing comes out clean.
Then be honest about the cost: the last character is at len minus one rather than at len, which is where nearly every off-by-one in this chapter comes from. It is a trade, not a free win, and knowing which side of the trade you are on is what stops the errors.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The offset reading is worth over-learning now, because slicing in the next lesson depends on it entirely and negative indices only make sense from it. The len-minus-one rule stops surprising you after two or three IndexErrors, which are loud and therefore cheap. The traversal condition is the one that produces silent bugs, so the habit of testing both ends is the durable takeaway. And the for loop is the form you will use nearly always — if its behaviour is unclear, it is worth writing both versions side by side once.
Connect it up
One page, from memory.
Draw it
Draw a six-character string as boxes, with three rows of labels beneath: the positive indices, the negative indices, and the character each names. Mark the two indices that raise an IndexError. Then, beside it, write the while traversal and the for traversal of the same string, and draw an arrow from each part of the while version to the thing in the for version that replaced it — marking clearly the three parts that disappeared entirely.
Recap
Two pages, and a type you have used since lesson 1a has come apart into pieces.
| If you remember one thing | It is this |
|---|---|
| From indexing | The index is a distance, not a count. The first character is zero steps along. |
| From len | The count and the last index always differ by one. |
| From negative indices | s[-1] is the last character, with no arithmetic to get wrong. |
| From traversal | Two of the three failures are silent, and both are at the ends. Test both ends. |
| From the for loop | It holds characters, not positions — and it cannot go out of range. |
The next lesson takes several characters at once: slices, the reason a slice excludes its end, the discovery that a string cannot be changed at all, and the two patterns — searching and counting — that most string processing is built from.
Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-72 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.