8a A String Is a Sequence: Indexing, len, and Traversal

This lesson reclassifies the string as a sequence of characters, introduces the bracket operator and zero-based indexing, meets the IndexError at the end of a string, and gives the two ways to traverse a string one character at a time.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 8a A String Is a Sequence: Indexing, len, and Traversal

Title

Python · Chapter 8 — Strings

§8.1-8.3, pp. 71-72

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-72 — the pages these objectives are drawn from

3. Before we start: how many things is a string?

Warm-up

You have treated a string as one value since lesson 1a. Question that.

Discussion prompt

The value 'banana' has a type, str, and behaves as a single value in an assignment. But it obviously has parts. Name two operations you would want on those parts that nothing you know so far can do.

Hint: Getting one letter out, and going through them one at a time.

Answer:

Getting a single character out of it, and processing every character in turn. Neither is possible with what you have — concatenation and repetition treat a string as a lump.

The reason is that you have been treating a string as an atom, like an int. It is not.

Strings are not like integers, floats and booleans. A string is a sequence, which means it is an ordered collection of other values — and this chapter is what follows from that.

4. The one idea behind this chapter: a string is an ordered collection

Concept

A string is a sequence of characters. That means it has parts, the parts are in a definite order, and every part has a position you can name — which is what the bracket operator is for.

sequence — An ordered collection of values, in which each value is identified by an integer index.

The word ordered is doing work. A string is not merely a bag of characters: 'banana' and 'nabana' contain the same letters and are different strings, exactly as lesson 2b's concatenation was not commutative.

Figure (svg): The string banana drawn as six boxes with the indices zero to five beneath them

Six characters, six positions, numbered from zero.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-71

5. The bracket operator, and why counting starts at zero

Section

Section 1

6. Selecting one character

Concept

You can access the characters one at a time with the bracket operator. The expression in brackets is called an index, and it indicates which character in the sequence you want.

>>> fruit = 'banana'
>>> letter = fruit[1]
>>> letter
'a'
ExpressionWhat it doesResult
fruit[1]select character number 1'a', not 'b'
fruit[0]select character number 0'b'
whythe index is an OFFSET from the beginningthe first has offset zero

For most people the first letter of banana is b, not a. But for computer scientists the index is an offset from the beginning of the string, and the offset of the first letter is zero. So b is the zero-eth letter of banana, a is the one-eth, and n is the two-eth.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-71

7. Picture it: offsets, not counts

Picture it

The index answers how far from the start, not which one in order.

Figure (svg): The string banana with each character labelled by its offset from the beginning

The first character is zero steps from the beginning, which is why its index is 0.

Once you read the index as a distance rather than a position in a queue, zero-based counting stops being arbitrary — the first item really is zero steps along.

8. Worked example: the index is an expression

Worked example

Anything with an integer value can go in the brackets.

>>> fruit = 'banana'
>>> i = 1
>>> fruit[i]
'a'
>>> fruit[i+1]
'n'
ExpressionWhat happensResult
fruit[i]i is looked up: 1'a'
fruit[i+1]the expression is evaluated: 2'n'
the ruleany integer-valued expression workscomposition again

Use a variable as an index.

Why: As an index you can use an expression that contains variables and operators — the same composition rule from lesson 3a.

Use an arithmetic expression.

Why: The expression inside the brackets is fully evaluated before the character is selected, exactly as an argument is evaluated before a call.

Note the requirement.

Why: The value of the index has to be an integer. A float raises a TypeError, even one like 1.5 whose meaning seems obvious.

Figure (svg): The state of the program after each line of Worked example the index is an expression, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Both work, giving 'a' and 'n'. Anything that evaluates to an integer can be an index, which is what makes loops over indices possible.

Verify: Try a float index and read the error.

Why: fruit[1.5] gives a TypeError saying string indices must be integers. The message names the requirement directly — there is no character halfway between two characters, so the operation has no meaning for a fractional index.

9. Predict: what is fruit[2]?

Prediction

Count offsets, starting from zero.

>>> fruit = 'banana'
>>> fruit[2]
IndexCharacterWhich one it is
offset 0bthe first character
offset 1athe second
offset 2nthe third

Predict first

What is fruit[2] when fruit is 'banana'?

  • 'b'
  • 'a'
  • 'n'
  • 'an'

Correct: 'n' — offsets 0, 1 and 2 are b, a and n, so index 2 is the third character.

Why: The book's own phrasing helps here: n is the two-eth letter of banana. Note that a single index always produces a single character, never a group — getting two characters needs a slice, which is the next lesson. If you answered 'a', you were reading the index as a count rather than an offset.

10. Worked example: why zero-based is not arbitrary

Worked example

Compare the two conventions on the arithmetic they produce.

# zero-based: the index is an offset
#   first character:  fruit[0]
#   n-th character:   fruit[n-1]... no: fruit[n] is the (n+1)th
#   slice of first 3: fruit[0:3]

# one-based would need:
#   slice of first 3: fruit[1:3]  -- or 1:4? ambiguous
ConventionWhat an index meansConsequence
zero-basedindex = distance from the startarithmetic is clean
length and indexthe last index is len - 1one subtraction
slicingn:m is m minus n charactersno adjustment needed

Read the index as a distance.

Why: The offset of the first letter is zero because it is at the beginning — no steps along.

Notice what that buys.

Why: The number of characters between two indices is simply the difference, with no adjustment. That becomes important in the next lesson, where slicing relies on it.

Notice what it costs.

Why: The last character's index is len minus one rather than len, which is the source of the most common error in this chapter.

Figure (svg): Two columns comparing zero-based indexing with a hypothetical one-based scheme on the arithmetic each produces

Clean arithmetic in the middle, one adjustment at the end.

Zero-based indexing makes distances and slice lengths come out as simple subtractions, at the cost of a len-minus-one adjustment at the end. It is a trade, and the arithmetic is why it is the near-universal choice.

Verify: Count the characters in fruit[1:4] and compare with 4 minus 1.

Why: Three characters, and 4 - 1 is 3. Under a one-based convention that subtraction would need a correction, which is exactly the sort of adjustment that produces off-by-one errors.

11. Trap: reading the index as a position in a queue

Trap

The trap

A student reads fruit[1] as the first letter and is surprised to get 'a'.

Import the everyday meaning of first

Why: In ordinary counting the first thing is number one, and the index looks like a count.

Every index is then off by one, and the error is systematic rather than occasional — it affects every access in the program.

The fix

Read the index as an offset: how far along, not which one.

Say the character at offset 1 rather than the first character

Why: Offset 1 is one step from the beginning, which is the second character.

Use the book's ordinals if it helps

Why: b is the zero-eth letter, a is the one-eth, n is the two-eth. Deliberately odd, and it fixes the reading.

The reframing is worth the effort because it also explains slicing, negative indices and the len-minus-one rule, all of which follow from distance rather than from position.

12. Sort: legal or an error?

Sorting

The index must be an integer, and it must be in range.

Sort into buckets

For fruit = 'banana', sort each expression.

gives a character
fruit[0]; fruit[5]; fruit[2+1]
raises an error
fruit[1.5]; fruit[6]; fruit['1']
ok
Each index is an integer between 0 and 5, which are the valid offsets for a six-character string. An arithmetic expression is fine, because it is evaluated to an integer before the selection happens.
err
One index is a float and one is a string, and the value of the index has to be an integer — both give a TypeError. The third is an integer and out of range, giving an IndexError, which is the subject of the next section.

13. Complete it: select the third character

Faded example

Offsets start at zero.

Fill in the blanks

word = 'python'
third = word[2] # should be 't'

Why: The characters are at offsets 0 (p), 1 (y) and 2 (t), so the third character is at index 2. The general rule follows: the n-th character, counting from one as people do, is at index n minus one. That single subtraction is where zero-based indexing costs you something, and it is worth stating explicitly rather than rediscovering each time.

14. Think it through: why must an index be an integer?

Socratic

1.5 has an obvious meaning to a person. Say why it does not to Python.

Discussion prompt

fruit[1.5] raises a TypeError. What would it even mean, and why is refusing better than rounding?

Hint: There is nothing between two adjacent characters.

Answer:

It has no meaning: there is no character halfway between two characters, so no value could be returned that is not an invention.

Rounding would be a guess — and a silent one. fruit[1.5] would give either 'a' or 'n' depending on the rounding rule, and neither is what anybody meant.

This is the same design decision as refusing to add a str to an int in lesson 1b: where there is no obviously correct answer, Python declines rather than guessing. A TypeError at the point of the mistake is much cheaper than a wrong character travelling onward.

15. len, and the character at the end

Section

Section 2

16. The most common off-by-one in the book

Concept

len is a built-in function that returns the number of characters in a string. Using it to reach the last character is where nearly everybody gets caught once.

>>> fruit = 'banana'
>>> len(fruit)
6
>>> last = fruit[len(fruit)]
IndexError: string index out of range
ExpressionWhat happensResult
len(fruit)the number of characters6
fruit[6]there is no character at offset 6IndexError
fruit[5]the valid offsets are 0 to 5'a', the last character

The reason for the IndexError is that there is no letter in banana with the index 6. Since we started counting at zero, the six letters are numbered 0 to 5. To get the last character you have to subtract 1 from the length — or use a negative index.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-72

17. Picture it: six characters, offsets 0 to 5

Picture it

The count and the last index differ by one, always.

Figure (svg): The string banana with positive indices zero to five above and negative indices minus one to minus six below

The last character is at index 5 counting forwards, and at index -1 counting backwards.

Negative indices count backward from the end: fruit[-1] is the last letter, fruit[-2] the second to last, and so on. They exist precisely so you do not have to write len minus one.

18. Worked example: two ways to reach the last character

Worked example

One arithmetic, one direct. Both are used constantly.

>>> fruit = 'banana'
>>> length = len(fruit)
>>> fruit[length-1]
'a'
>>> fruit[-1]
'a'
ExpressionHow it locates the characterResult
len(fruit)6the count
fruit[6-1]index 5the last character
fruit[-1]counting backward from the endthe same character

Compute the last index from the length.

Why: Since counting started at zero, the six letters are numbered 0 to 5, so the last index is length minus one.

Use a negative index instead.

Why: The expression fruit[-1] yields the last letter, fruit[-2] the second to last, and so on.

Notice why negative indices exist.

Why: They remove the subtraction, and with it the commonest place to make an off-by-one error.

Figure (svg): The state of the program after each line of Worked example two ways to reach the last character, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Both give 'a'. The negative index is shorter and safer; the length-minus-one form is what you need when the index is being computed for other reasons anyway.

Verify: Check the relationship between the two schemes.

Why: fruit[-1] and fruit[len(fruit)-1] name the same character, and in general fruit[-n] is fruit[len(fruit)-n]. That correspondence is worth knowing, because it means a negative index is not a separate feature but a shorthand.

19. Predict: which index raises an error?

Prediction

Six characters, and two indexing schemes.

Predict first

For fruit = 'banana', which of these raises an IndexError?

  • fruit[5]
  • fruit[-6]
  • fruit[6]
  • fruit[0]

Correct: fruit[6] — the valid forward indices are 0 to 5, and 6 is one past the end.

Why: A six-character string has offsets 0 through 5 forwards and -1 through -6 backwards, so fruit[-6] is the first character and is perfectly valid. The one that fails is the index equal to the length, which is exactly the off-by-one the section is about. Note that fruit[-7] would also fail, for the same reason at the other end.

20. Worked example: reading an IndexError

Worked example

The message names the problem precisely. Extract what it tells you.

>>> fruit = 'banana'
>>> fruit[6]
IndexError: string index out of range
PartWhat it tells youNext step
the error nameIndexErrorthe index was the problem
the messagestring index out of rangenot the type, the value
what to checklen versus the index used6 versus valid 0 to 5

Read the error kind.

Why: IndexError, not TypeError. The index was of the right type and the wrong value, which is the ValueError-versus-TypeError distinction from lesson 6c applied here.

Read the message.

Why: Out of range — so the fix is arithmetic rather than conversion.

Find the arithmetic.

Why: Compare the index used against len minus one. Almost every IndexError in this chapter is a missing subtraction of one.

Figure (svg): A panel showing an IndexError with the valid index range and the index that was used

The error says the index was an integer that named no character. The fix is nearly always to subtract one, or to use a negative index instead.

Verify: Check what a negative index too large in magnitude does.

Why: fruit[-7] also raises an IndexError, because there are only six characters to count back through. Negative indices remove one off-by-one and do not remove the range check — they still have to name a character that exists.

21. Trap: using len as an index

Trap

The trap

A student writes s[len(s)] to reach the last character, reasoning that a six-character string should have a sixth character.

Treat the length as the last position

Why: It is the count, and counts and last positions coincide under one-based numbering.

Under zero-based numbering they never coincide. len(s) is always one past the end, so this always raises an IndexError — which is at least loud.

The fix

The last index is len minus one, and there is a shorter way to say it.

Write s[-1] when you want the last character

Why: It needs no arithmetic and cannot be off by one.

Write s[len(s)-1] when you already have the length for another reason

Why: For instance inside a loop that counts down.

It is worth noticing that this error is loud: an index one past the end raises immediately. The quiet version is a loop condition that stops one short, which produces a wrong answer instead — and that is the subject of the next section.

22. Match each index to the character it selects

Matching

Forwards and backwards name the same characters.

Match the pairs

  • a. fruit[0]
  • b. fruit[-1]
  • c. fruit[len(fruit)-1]
  • d. fruit[-6]
  • r1. the first character, 'b'
  • r2. the last character, 'a'
  • r3. the last character, 'a' — the same one
  • r4. the first character, 'b' — the same one

Why: Each character has two names, one positive and one negative, and they differ by the length: index i forwards is index i minus len backwards. Knowing that the two schemes coincide is what makes negative indices feel like a shorthand rather than a separate mechanism.

23. Fill the middle: reach the second-to-last character

Fill the middle

Two ways to write it. Give the negative one.

Fill in the blanks

word = 'python'
second_last = word[-2] # should be 'o'

Why: fruit[-1] is the last and fruit[-2] the second to last, so -2 gives 'o' from 'python'. The equivalent positive form is word[len(word)-2], which is longer and involves a subtraction that can go wrong. Negative indices exist for exactly this: reaching things relative to the end without computing the length.

24. Explain it yourself: why is len not a valid index?

Explain it to yourself

State it in terms of offsets rather than as a rule to memorise.

Discussion prompt

Explain, using the word offset, why a string of length 6 has no character at index 6 — without saying because indexing starts at zero.

Hint: How far along is the last character?

Answer:

The last character is five steps from the beginning, because the first is zero steps along. Six steps would take you past the end entirely.

So len counts characters and an index measures distance, and those are different quantities that happen to be related by one.

Framing it this way makes the fix obvious rather than memorised: the largest distance you can travel and still land on a character is one less than the number of characters, which is len minus one.

25. Traversal with a while loop

Section

Section 3

26. Processing a string one character at a time

Concept

A lot of computations involve processing a string one character at a time. Often they start at the beginning, select each character in turn, do something to it, and continue until the end. This pattern of processing is called a traversal.

traversal — Processing each element of a sequence in turn, from one end to the other.

index = 0
while index < len(fruit):
    letter = fruit[index]
    print(letter)
    index = index + 1
LineIts roleNote
index = 0initializationstart at the beginning
index < len(fruit)the conditionstops one past the last index
fruit[index]select the current characterthe work
index = index + 1the updatemove along

The loop condition is index less than len(fruit), so when index is equal to the length the condition is false and the body does not run. The last character accessed is the one with the index len minus one, which is the last character in the string.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-72

27. Picture it: the index walking along

Picture it

One pass per character, with the index taking each valid value in turn.

Figure (svg): The string banana with the index shown moving across each character in turn

The condition index < len stops the loop exactly when index reaches 6, which is one past the end.

Notice the four parts of a loop from lesson 7a, all present: initialize the index, test it against the length, do the work, update the index.

28. Worked example: why the condition uses less-than

Worked example

One character in the condition decides whether the loop is right.

# correct
while index < len(fruit):

# one too far
while index <= len(fruit):
ConditionThe final passOutcome
index < lenlast pass has index = len - 1the last character
index <= lenlast pass has index = lenIndexError
the rulethe largest valid index is len - 1so the test is strict

Work out the last pass under the correct condition.

Why: The condition fails when index equals len, so the last pass has index equal to len minus one — which is the last valid index.

Work out the last pass under the wrong one.

Why: With less-than-or-equal, the loop runs one more time with index equal to len, and fruit[len] raises an IndexError.

Notice which mistake is louder.

Why: This one crashes, which is helpful. The quieter mistake is a condition that stops one short, which silently skips the last character.

Figure (svg): Two columns comparing a traversal condition using less-than with one using less-than-or-equal

One character of difference, and one extra pass that cannot succeed.

Strict less-than is correct: the loop must stop when index reaches len, because len is one past the last valid index. Using less-than-or-equal runs one pass too many and raises an IndexError.

Verify: Count the passes and compare with the length.

Why: The loop runs exactly len times, once per character. Whenever a traversal produces the wrong number of passes, the condition is where to look — and comparing the pass count with the length settles it in one step.

29. Predict: how many passes?

Prediction

Count them, then compare with the length.

fruit = 'banana'
index = 0
while index < len(fruit):
    print(fruit[index])
    index = index + 1
Index valuesThe conditionResult
index = 0 to 5the condition is truesix passes
index = 66 < 6 is falsethe loop ends
outputone line per charactersix lines

Predict first

How many lines does this print?

  • Five
  • Six
  • Seven
  • Zero

Correct: Six — one per character, with the index taking every value from 0 to 5.

Why: The condition is true for indices 0 through 5 and false at 6, so the loop runs exactly len times. That equality — passes equals length — is the check to run on any traversal: five lines would mean the condition stops one short, and seven would mean an IndexError.

30. Worked example: traversing backwards

Worked example

The book sets this as an exercise. The changes are all in the loop's three parts.

index = len(fruit) - 1
while index >= 0:
    print(fruit[index])
    index = index - 1
PartWhat changesValue
initializationstart at the last valid indexlen - 1
conditionkeep going while the index is still validindex >= 0
updatemove toward the beginningindex - 1

Change where it starts.

Why: The last valid index is len minus one, which is where a backward traversal begins.

Change the condition.

Why: It must remain true for index 0, since that is a real character, so the test is greater-than-or-equal rather than strictly greater.

Change the direction of the update.

Why: Decrement rather than increment, so the index moves toward the condition failing.

Figure (svg): The state of the program after each line of Worked example traversing backwards, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

All three parts change: the initialization, the condition, and the direction of the update. The body is untouched, which is a good sign that the traversal pattern is genuinely separable from the work.

Verify: Check the boundary at both ends.

Why: The first pass uses index len-1 and the last uses index 0, so both end characters are visited — and index -1 is never used, because the condition stops first. Checking both ends is the standard way to validate a traversal, since off-by-one errors live at exactly those two places.

31. Trap: a traversal that misses the last character

Trap

The trap

A student writes the condition as index < len(fruit) - 1, reasoning that the last valid index is len minus one.

Confuse the last valid INDEX with the stopping point

Why: The last valid index really is len minus one, so putting it in the condition feels right.

The condition is false when index equals len minus one, so the loop stops before processing the last character. The program runs perfectly and silently produces an answer that is one item short.

The fix

The condition names the first index that is NOT to be processed.

Stop at len, not at len minus one

Why: index < len is true for every valid index including the last, and false for the first invalid one.

Check by counting passes

Why: A traversal of a six-character string should run six times. Five means the condition stops one short.

This is the quiet version of the error. Going one too far raises an IndexError immediately; stopping one short produces a wrong answer with no message at all, which makes it much more expensive.

32. Watch the traversal: index and character

Invariant

Step through and track both the index and the character it selects.

Step through it

At which index does the loop stop, and why is no character accessed at that index?

  1. The first pass. index is 0, which selects the character at offset zero.
  2. The index has been incremented once and selects the second character.
  3. Two passes later. Note that 'a' appears at several indices — the index identifies a position, not a value.
  4. The last pass, at index 5, which is len minus one and therefore the last character.
  5. The index reaches 6 and the condition fails. No character is selected at this index, which is why the loop stops here rather than one pass later.

It stops at 6, which is len — one past the last valid index. Nothing is accessed there because the condition is tested BEFORE the body, which is lesson 7a's flow of execution doing exactly the job it exists for.

33. Complete it: traverse a string forwards

Faded example

The four parts of a loop, with two blanks.

Fill in the blanks

index = 0
while index < len(word):
print(word[index])
index = index + 1

Why: A forward traversal starts at index 0, the first valid offset, and increments toward the condition failing at len. Getting either blank wrong has a characteristic symptom: starting at 1 silently skips the first character, and decrementing produces an immediate IndexError at index -1... except that -1 is a valid index in Python, so it would silently traverse backwards forever. That second case is worth noticing.

34. Push the boundary: what about the empty string?

Edge cases

A traversal of nothing should do nothing. Check that it does.

Discussion prompt

What does the while-loop traversal do when fruit is the empty string? Trace it, and say which property of the while statement makes the answer correct.

Hint: What is len of the empty string?

Answer:

len is 0, so the condition 0 < 0 is false on the very first test and the body never runs. Nothing is printed and no error occurs.

That is exactly right: traversing an empty sequence should do nothing at all.

The property that makes it work is that a while loop tests before the first pass, which is lesson 7a's socratic probe answered concretely. A loop that tested afterwards would try to access fruit[0] of an empty string and raise an IndexError — so the empty case is handled for free rather than needing a special check.

35. Traversal with a for loop

Section

Section 4

36. The shorter way, and the one you will use

Concept

Another way to write a traversal is with a for loop. Each time through the loop, the next character in the string is assigned to the variable, and the loop continues until no characters are left.

for letter in fruit:
    print(letter)
PassWhat letter holdsOutput
pass 1letter is 'b'print b
pass 2letter is 'a'print a
...each character in turn...
after the lastno characters leftthe loop ends

Notice what has disappeared: there is no index, no initialization, no condition and no update. The loop variable holds the character itself rather than its position, and the loop ends when the string does.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-73

37. Picture it: two traversals, same output

Picture it

Five lines become two, and every off-by-one opportunity disappears.

Figure (svg): Two columns comparing a while traversal using an index with a for traversal over the characters

The for version cannot go out of range, because it never computes an index.

Use the for version unless you actually need the index — for instance to look at two positions at once, which is what lesson 9b is about.

38. Worked example: the abecedarian series

Worked example

The book's example, and it combines a for loop with concatenation.

prefixes = 'JKLMNOPQ'
suffix = 'ack'

for letter in prefixes:
    print(letter + suffix)
PassWhat is computedOutput
letter = 'J''J' + 'ack'Jack
letter = 'K''K' + 'ack'Kack
...one per prefix...
letter = 'Q''Q' + 'ack'Qack

Notice the loop variable holds a character.

Why: Each time through the loop, the next character in the string is assigned to letter. It is a one-character string, which can be concatenated like any other.

Notice the concatenation.

Why: letter + suffix builds a new string on every pass. Nothing is modified — a new string is created each time.

Notice the count.

Why: Eight prefixes, eight lines of output. The loop runs once per character with no index arithmetic anywhere.

Figure (svg): A diagram showing each character of the prefix string being concatenated with the suffix to produce a name

A new string is built on every pass. The originals are untouched.

Eight names, from Jack to Qack, one per character of the prefix string. The whole traversal is two lines.

Verify: Check the output against the book's list of ducklings.

Why: It gives Oack and Quack as Oack and Qack, which the book notes are misspelled — the real names are Ouack and Quack. That the program is correct and the OUTPUT is wrong is a nice small example of a semantic error: the code does exactly what it says and not what was wanted.

39. Predict: what does the loop variable hold?

Prediction

Not a number.

for letter in 'abc':
    print(letter)
PassWhat letter holdsOutput
pass 1letter is 'a'a
pass 2letter is 'b'b
pass 3letter is 'c'c

Predict first

What does this print?

  • 0, 1, 2
  • a, b, c
  • abc three times
  • Nothing

Correct: a, b, c — each on its own line, because the loop variable holds each character in turn.

Why: The for loop over a string assigns the next character to the variable on each pass, not the next index. That is the essential difference from the for-in-range form of chapter 4, and it is why the loop needs no length and no arithmetic. Each value is a one-character string, so it can be printed, concatenated or compared like any other string.

40. Worked example: when you still need the index

Worked example

The for loop hides the index. Sometimes that is a loss.

# for: you have the character, not its position
for letter in word:
    print(letter)

# while: you have both
index = 0
while index < len(word):
    print(index, word[index])
    index = index + 1
FormWhat you have access toVerdict
forthe character onlysimpler, safer
whilethe character and its positionmore to get wrong
chooseby whether you need the positionusually you do not

Ask what the body needs.

Why: If it only needs each character, the for loop gives exactly that and nothing to get wrong.

Notice what the for loop does not give.

Why: The position. If you need to report where something was found, or to compare two positions, the index has to come from somewhere.

Choose accordingly.

Why: Prefer for; use while with an index when the position is genuinely part of the problem.

Figure (svg): The state of the program after each line of Worked example when you still need the index, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The for loop is shorter and safer and hides the index. Use it unless the position itself is needed — which is less often than beginners expect.

Verify: Look ahead to a case that needs the index.

Why: Lesson 8b's find function must return the position where a character was found, so it cannot use a plain for loop over characters. That is the standard reason to reach for an index, and it is worth knowing that it exists rather than assuming for always works.

41. Trap: expecting the loop variable to be an index

Trap

The trap

A student writes for i in word and then uses word[i], expecting i to be a position.

Carry over the for-in-range form from chapter 4

Why: There the loop variable really was a number, so the habit is understandable.

Here i holds a character, so word[i] tries to index a string with a string and raises a TypeError — string indices must be integers.

The fix

The for loop over a string gives you the characters, not their positions.

Name the variable for what it holds

Why: letter or char, not i. The name is the cheapest defence against this confusion.

If you need positions, loop over a range instead

Why: for i in range(len(word)) gives indices, and word[i] then works — which is the bridge between the two forms.

The TypeError here is informative: string indices must be integers is the same message as fruit[1.5], and it tells you the bracket got something that was not a position.

42. Discriminate: for loop or while with an index?

Discrimination

Ask whether the body needs the position.

Sort into buckets

For each task, which traversal form is appropriate?

a for loop is enough
print every character on its own line; count how many times 'a' appears; build a new string with each character doubled
an index is needed
report the position of the first vowel; compare the character at position i with the one at position len-1-i; print each character together with its index
forl
Each of these needs only the characters themselves. Printing, counting and building a new string never refer to where a character sits.
idx
Each of these needs a position: to report it, to compare two positions with each other, or to display it. The character alone cannot supply that.

43. Translate: while traversal to for traversal

Translation

Match each part of the while version to what replaces it.

Match the pairs

  • a. index = 0
  • b. while index < len(word):
  • c. letter = word[index]
  • d. index = index + 1
  • r1. handled automatically: the loop starts at the first character
  • r2. for letter in word:
  • r3. handled automatically: letter is assigned each pass
  • r4. handled automatically: the loop advances by itself

Why: Three of the four parts vanish entirely and the fourth becomes a single line. What is really happening is that the for loop is a specialised construct for the traversal pattern, with the bookkeeping built in — which is why it cannot go out of range and the while version can.

44. Explain it: which loop should I use?

Explain it

A question every beginner asks once and then never again.

Discussion prompt

A classmate asks whether they should use a for loop or a while loop to go through a string. Give them a one-sentence rule and one example of the exception.

Hint: The rule is about what the body needs.

Answer:

Say: use a for loop unless the body needs to know WHERE it is, in which case use an index.

The example: a function that reports the position of a character has to know the position, so a plain for loop over characters cannot do it.

Then add the reason it matters, which is not style: the for loop cannot go out of range, because it never computes an index. Choosing it eliminates a whole category of bug, which is a stronger argument than brevity.

45. Putting it together: reading and writing traversals

Section

Section 5

46. The pattern, and the three ways it goes wrong

Concept

A traversal has the same shape whatever it does: visit each character in turn, do something with it, and stop at the end. The work varies; the pattern does not — and neither do the three ways it fails.

# the pattern
for letter in word:
    # do something with letter

# the same pattern with an index
index = 0
while index < len(word):
    letter = word[index]
    # do something with letter
    index = index + 1
FailureWhat causes itSymptom
off by one at the startindex starts at 1the first character is missed
off by one at the endcondition uses len - 1the last character is missed
one too farcondition uses <= lenIndexError

Two of those three failures are silent, and both are at the ends. That is why checking the first and last character is the standard test for a traversal — the middle almost never goes wrong.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 72-73

47. Picture it: the three failures all live at the ends

Picture it

The middle of a traversal is nearly always right.

Figure (svg): The string banana with the first and last characters highlighted as the places where traversal errors occur

The highlighted ends are where off-by-one errors live. The middle takes care of itself.

Which gives a two-item test for any traversal: does it process the first character, and does it process the last? If both are yes, the middle almost certainly is too.

48. Worked example: writing a traversal that counts

Worked example

The pattern with a counter in it, which lesson 8b formalises.

word = 'banana'
count = 0
for letter in word:
    if letter == 'a':
        count = count + 1
print(count)
CharacterThe testcount after
bnot 'a'count stays 0
amatchescount becomes 1
nnot 'a'count stays 1
a, n, atwo more matchescount becomes 3

Initialize the counter before the loop.

Why: Lesson 7a's rule: initialize outside, update inside. Putting count = 0 in the body would reset it on every pass.

Traverse with a for loop.

Why: The body only needs the characters, not their positions, so the for form applies.

Update conditionally.

Why: The counter is incremented only when the character matches, which is what makes it a count rather than a length.

Figure (svg): The state of the program after each line of Worked example writing a traversal that counts, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

3 — there are three a's in banana. The traversal pattern supplies the loop, and the work inside it is a conditional increment.

Verify: Check against len, which counts every character.

Why: len('banana') is 6 and the count is 3, so the condition is genuinely filtering. If a counter ever equals the length, the condition is always true and is doing nothing — which is a quick check on any counting loop.

49. Error analysis: three broken traversals

Error analysis

Each has one fault. Name it and say whether it is loud or silent.

Annotate

  • The initialization is 1 rather than 0, so the traversal begins at the second character.
  • The first character is never printed, and no error occurs. This is a silent bug — the program runs perfectly and produces output that is one item short.
  • A second version with the condition written as index <= len(word) would run one pass too many and raise an IndexError on the last pass. That one is loud, and therefore cheaper.
  • A third version with the condition index < len(word) - 1 stops one short and silently omits the last character — the mirror image of the first fault.
  • Two of the three failures are silent, and both are at the ends of the string. That is why the standard test is to check the first and last characters.
  • The fix for this one is a single character: index = 0.

The loud failure is the lucky one. A traversal that quietly processes one character too few can survive for a long time.

50. Worked example: testing a traversal at both ends

Worked example

Two checks find nearly every traversal bug.

# does it process the FIRST character?
#   count the a's in 'abc'   -> should be 1
# does it process the LAST character?
#   count the a's in 'cba'   -> should be 1
# does it handle the empty case?
#   count the a's in ''      -> should be 0
Test inputWhat it exercisesWhat it catches
'abc'the target is firstcatches a start-at-1 bug
'cba'the target is lastcatches a stop-early bug
''nothing at allcatches a special-case bug

Design a test where the answer depends on the first character.

Why: If the traversal starts at index 1, this test fails and every other test might pass.

Design one where it depends on the last.

Why: If the condition stops one short, only this test fails.

Add the empty case.

Why: It costs nothing and catches any code that assumes at least one character exists.

Figure (svg): Two columns listing three test inputs against the traversal bug each one detects

Three one-word tests, three distinct bugs.

Three tests, each targeting one of the three failure modes, and each one input long. Between them they cover the places where traversals actually go wrong.

Verify: Ask why a test on 'banana' alone is insufficient.

Why: It has a's in the middle as well as at the end, so a traversal that skipped the first character would still find two and look plausible. A test only discriminates if the answer changes when the bug is present, which is the same argument as lesson 5b's chain-versus-separate-ifs.

51. Trap: testing only on a convenient string

Trap

The trap

A traversal is tested on 'banana' and works, so it is declared correct.

Test on the example from the book

Why: It is the string in front of you and it gives a plausible answer.

banana has a's in the middle, so a traversal that skips the first or last character still finds some and looks right. The test cannot distinguish a correct traversal from two broken ones.

The fix

Choose inputs where the answer changes if the bug is present.

Put the thing being looked for at the very start, and then at the very end

Why: Two inputs, each of which fails for exactly one of the two silent bugs.

Add the empty string

Why: It costs one test and catches everything that assumes a first character exists.

This is lesson 6a's advice about choosing a test case, applied to sequences: it is useful to know the right answer, and it is more useful for the right answer to be different when the code is wrong.

52. Predict: which character is missed?

Prediction

The condition stops one short.

word = 'abc'
index = 0
while index < len(word) - 1:
    print(word[index])
    index = index + 1
ValueWhat it isConsequence
len(word) - 13 - 12
indices used0 and 1a and b
index 22 < 2 is falsec is never printed

Predict first

What does this print?

  • a, b and c — every character
  • a and b — the last character is missed
  • b and c — the first character is missed
  • nothing at all

Correct: a and b — the last character is missed, because the condition stops before index 2.

Why: The condition names the first index NOT to process, and here it names 2, which is a valid character. The confusion comes from the last valid index being len minus one: that is the last index to USE, not the place to stop. The correct condition is index < len(word), which is false only at the first invalid index.

53. Compare: the two traversal forms

Comparison

Fill the blanks. Both are correct; they differ in what they expose.

Comparison matrix

Questionfor letter in wordwhile with an index
What does the loop variable hold?the characterthe index; the character needs word[index]
Can it go out of range?no — it never computes an indexyes, if the condition is wrong
How many lines?twofive
When do you need it?almost alwayswhen the position itself is part of the problem

The second row is the strongest argument. A construct that cannot make a whole class of error is worth preferring even when both would work.

54. Where off-by-one errors come from

Real world

They are not a Python problem.

Discussion prompt

Think of an off-by-one error outside programming — a fencepost, a schedule, a count of days. Describe it, and say what makes this family of mistake so persistent.

Hint: How many posts does a fence with ten panels have?

Answer:

Eleven, not ten — the classic fencepost problem. Counting the gaps and counting the posts give answers that differ by one, and it is easy to answer the wrong question.

Dates do it too: from Monday to Friday is four days or five days depending on whether you count the endpoints.

What makes them persistent is that both answers are correct answers to slightly different questions, so neither looks obviously wrong. The programming defence is the same as the everyday one: state exactly which endpoints are included, and check the ends rather than the middle.

55. Compare: three ways to reach a character

Comparison

Fill the blanks. All three name characters; they differ in where they count from.

Comparison matrix

FormWhere it counts fromThe last character
positive indexthe beginning, from 0s[len(s)-1]
negative indexthe end, from -1s[-1]
for loopno counting at allthe last pass, automatically

The third row is why the for loop is preferred: a form that does no counting cannot count wrongly.

56. The procedure: writing a traversal

Pattern

Six steps, and the last two are the tests that matter.

  1. Ask whether the body needs the position of each character, or only the character itself.
  2. If only the character, write for letter in word — and stop, because there is nothing else to get wrong.
  3. If the position is needed, initialize the index to 0 and write the condition as index < len(word).
  4. Do the work, then update the index by one in the direction the condition requires.
  5. Test with the target at the very start of the string, and again at the very end.
  6. Test with the empty string, which should do nothing at all.

Step 5 is the one that finds real bugs. Two of the three traversal failures are silent and both are at the ends, so a test whose answer depends on a middle character proves nothing.

Python documentation — An Informal Introduction to Python An Informal Introduction to Python

57. Check yourself 1 of 3: indexing

Check

Offsets, counted from zero.

>>> word = 'python'
>>> word[1]
IndexCharacterNote
offset 0pthe first character
offset 1ythe second
the rulethe index is a distancenot a count

Check your understanding

What is word[1] when word is 'python'?

  • A. 'p'
  • B. 'y' (correct)
  • C. 'py'
  • D. An error

Answer: B

Why: The index is an offset from the beginning, and the offset of the first letter is zero — so index 1 is one step along, which is the second character. A single index always produces exactly one character; getting two would need a slice.

Why A tempts people
This is word[0]. Reading index 1 as the first character is the systematic off-by-one this section exists to prevent.
Why C tempts people
A single index selects one character. Two characters would need the slice syntax from the next lesson.
Why D tempts people
Index 1 is well within range for a six-character string, whose valid indices are 0 to 5.

58. Check yourself 2 of 3: the end of a string

Check

The count and the last index differ by one.

>>> s = 'abc'
>>> len(s)
3
>>> s[3]
IndexError: string index out of range
ExpressionWhat happensResult
len3 charactersindices 0, 1, 2
s[3]one past the endIndexError
s[2] or s[-1]the last character'c'

Check your understanding

Which expression gives the last character of s?

  • A. s[len(s)]
  • B. s[-1] (correct)
  • C. s[0]
  • D. len(s)

Answer: B

Why: Negative indices count backward from the end, so s[-1] is the last character. The equivalent positive form is s[len(s)-1], which needs a subtraction and is where the off-by-one usually creeps in.

Why A tempts people
This is one past the end and raises an IndexError. len is the count of characters, and the last index is always one less.
Why C tempts people
This is the first character, not the last.
Why D tempts people
This is a number, the length — not a character at all.

59. Check yourself 3 of 3: for loops over strings

Check

The loop variable holds a character.

for c in 'abc':
    print(c)
PassWhat c holdsOutput
pass 1c is 'a'a
pass 2c is 'b'b
pass 3c is 'c'c

Check your understanding

In a for loop over a string, what does the loop variable hold on each pass?

  • A. The index of the next character
  • B. The next character itself (correct)
  • C. The whole string
  • D. The number of characters remaining

Answer: B

Why: Each time through the loop, the next character in the string is assigned to the variable, and the loop continues until no characters are left. That is what distinguishes it from the for-in-range form of chapter 4, where the variable really does hold a number.

Why A tempts people
This is the for-in-range form, where the loop variable is an integer. Over a string directly, the variable holds characters — and using one as an index raises a TypeError.
Why C tempts people
The whole string is what the loop iterates over, not what the variable holds on any single pass.
Why D tempts people
Nothing in the loop tracks a remaining count. The loop simply ends when there are no characters left.

60. Where this shows up outside this course

Real world

Sequences and their indices are a shape you already use.

Discussion prompt

Find something in daily life that is an ordered collection with positions — a queue, a playlist, a street of houses. Does it number from zero or one, and where does that convention cause confusion?

Hint: Building floors are the classic case.

Answer:

Building floors are the best example: the floor at street level is the ground floor in Britain and the first floor in America, so the second floor means two different things — an off-by-one built into a language.

Playlists and page numbers count from one; distances and offsets count from zero. The pattern is that positions in a queue start at one and distances start at zero, which is exactly the distinction the book draws.

So Python is not being perverse. It has chosen the offset interpretation, and once you read s[0] as zero steps from the start it agrees with how distances work everywhere else.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence. This is the off-by-one the whole chapter turns on.

Predict first

For a string s of length 6, which of these is the correct condition for a traversal that visits every character exactly once?

  • while index <= len(s):
  • while index < len(s):
  • while index < len(s) - 1:
  • while index <= len(s) - 1:

Correct: while index < len(s): — and also the last option, which is equivalent, though the second is the idiomatic form.

Why: The condition must be true for every valid index, which is 0 to 5, and false at the first invalid one, which is 6. Strict less-than against len does exactly that. Less-than-or-equal against len runs one pass too many and raises an IndexError; strict less-than against len minus one stops one short and silently skips the last character. The fourth option is arithmetically the same as the second and is written that way far less often, because the subtraction is an extra opportunity to be wrong.

62. Explain it to someone else

Explain it

Zero-based indexing is the thing beginners most want explained.

Discussion prompt

A classmate is annoyed that s[1] is not the first character. Explain the convention in a way that makes it feel inevitable rather than arbitrary, and give them one thing it makes easier.

Hint: The word is offset.

Answer:

Say: the index is not a count, it is a distance from the beginning — and the first character is zero steps along, so its index is 0.

What it makes easier: the number of characters between two indices is just the difference, with no adjustment. s[2:5] holds 5 minus 2 characters, which is why slicing comes out clean.

Then be honest about the cost: the last character is at len minus one rather than at len, which is where nearly every off-by-one in this chapter comes from. It is a trade, not a free win, and knowing which side of the trade you are on is what stops the errors.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • Zero-based indexing, and reading an index as an offset
  • len, and why the last index is one less than the length
  • Writing a while traversal with the right condition
  • The for loop over a string, and what the loop variable holds

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The offset reading is worth over-learning now, because slicing in the next lesson depends on it entirely and negative indices only make sense from it. The len-minus-one rule stops surprising you after two or three IndexErrors, which are loud and therefore cheap. The traversal condition is the one that produces silent bugs, so the habit of testing both ends is the durable takeaway. And the for loop is the form you will use nearly always — if its behaviour is unclear, it is worth writing both versions side by side once.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw a six-character string as boxes, with three rows of labels beneath: the positive indices, the negative indices, and the character each names. Mark the two indices that raise an IndexError. Then, beside it, write the while traversal and the for traversal of the same string, and draw an arrow from each part of the while version to the thing in the for version that replaced it — marking clearly the three parts that disappeared entirely.

65. What you can do now

Recap

Two pages, and a type you have used since lesson 1a has come apart into pieces.

If you remember one thingIt is this
From indexingThe index is a distance, not a count. The first character is zero steps along.
From lenThe count and the last index always differ by one.
From negative indicess[-1] is the last character, with no arithmetic to get wrong.
From traversalTwo of the three failures are silent, and both are at the ends. Test both ends.
From the for loopIt holds characters, not positions — and it cannot go out of range.

The next lesson takes several characters at once: slices, the reason a slice excludes its end, the discovery that a string cannot be changed at all, and the two patterns — searching and counting — that most string processing is built from.

Think Python, 2nd edition — Allen B. Downey §8.1-8.3, pp. 71-72 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §8.1-8.3, pp. 71-72
  2. Python documentation — An Informal Introduction to Python
  3. Python documentation — Built-in Types

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108