This lesson introduces the gather and scatter operators for functions that take any number of arguments, then covers zip, iterators, tuple assignment in a for loop, and enumerate — the idioms for traversing two sequences at once or a sequence with its indices.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 12 — Tuples
§12.4-12.5, pp. 118-119
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-119 — the pages these objectives are drawn from
Warm-up
You have met one already, several times.
Discussion prompt
print accepts one argument, or three, or none at all. Every function you have written has a fixed number of parameters. How could a function possibly accept a number it does not know in advance?
Hint: Where would the extra arguments go?
Answer:
They would have to be collected into something — one object holding however many arrived, which the function could then look at.
That object needs to be a sequence of arbitrary values with a fixed shape once built, which is exactly the type the previous lesson introduced.
So the answer is a tuple, and this lesson's first section is the syntax for asking for one. The complementary operation — taking a sequence apart into separate arguments — turns out to use the same symbol.
Concept
A parameter name that begins with * gathers arguments into a tuple. The complement of gather is scatter: if you have a sequence of values and you want to pass it to a function as multiple arguments, you can use the same operator.
gather — The operation of assembling a variable-length argument tuple.
One star, two directions. In a definition it collects separate arguments into a tuple; at a call site it spreads a tuple into separate arguments. Which one happens depends entirely on where you write it.
Figure (svg): Two columns contrasting gathering arguments into a tuple with scattering a tuple into arguments
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-118
Section
Section 1
Concept
Functions can take a variable number of arguments. A parameter name that begins with * gathers arguments into a tuple. For example, printall takes any number of arguments and prints them.
def printall(*args):
print(args)
>>> printall(1, 2.0, '3')
(1, 2.0, '3')| Part | What happens | Note |
|---|---|---|
| three arguments passed | gathered into one tuple | (1, 2.0, '3') |
| inside the function | args is an ordinary tuple | indexable, iterable |
| the types | mixed freely | a tuple holds any values |
The gather parameter can have any name you like, but args is conventional — so a reader who sees *args knows immediately what is happening, and a reader who sees *things has to look twice.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-118
Picture it
The call has three values. The function has one parameter.
Figure (svg): A pipeline showing three separate arguments gathered into a single tuple parameter
And if the caller passes nothing at all, args is the empty tuple — which is a legitimate value, so the function still works.
Worked example
args is an ordinary tuple, so everything from the last lesson applies.
def sum_all(*args):
total = 0
for x in args:
total += x
return total
>>> sum_all(1, 2, 3)
6
>>> sum_all()
0| Call | What args holds | Result |
|---|---|---|
| sum_all(1, 2, 3) | args is (1, 2, 3) | loop adds each |
| sum_all() | args is () | the loop body never runs |
| the result | 0 for no arguments | which is correct |
Gather the arguments.
Why: Whatever the caller passes arrives as one tuple bound to args.
Treat it as a sequence.
Why: It is an ordinary tuple, so you can loop over it, take its length, index it, or pass it on.
Check the empty case.
Why: With no arguments args is the empty tuple, the loop body never runs, and the total stays 0 — which is the right answer for an empty sum.
Figure (svg): The state of the program after each line of Worked example using what was gathered, drawn as a ladder with one rung per traced line
6 and 0. The exercise the book sets, and it needs nothing beyond a loop over the gathered tuple.
Verify: Compare with the built-in sum.
Why: sum([1, 2, 3]) takes one sequence, while sum_all(1, 2, 3) takes separate arguments — and the book notes that sum(1, 2, 3) raises TypeError for exactly that reason. The two are not interchangeable, which is what makes the exercise worth doing.
Prediction
Three arguments, one parameter.
def printall(*args):
print(args)
printall(1, 2.0, '3')| Part | What happens | Result |
|---|---|---|
| three arguments | gathered | into one tuple |
| args | (1, 2.0, '3') | mixed types are fine |
| print(args) | prints the tuple | with its parentheses |
Predict first
What does this print?
Correct: (1, 2.0, '3') — the arguments are gathered into a tuple, and printing the tuple shows its own form.
Why: A parameter name beginning with * gathers arguments into a tuple, so args holds all three and print shows the tuple with its parentheses and commas. Option B is what print(*args) would produce, scattering them back into separate arguments — which is the operator in its other direction.
Worked example
The book's own comparison, and it is easy to get wrong.
>>> max(1, 2, 3)
3
>>> max([1, 2, 3])
3
>>> sum(1, 2, 3)
TypeError: sum expected at most 2 arguments, got 3
>>> sum([1, 2, 3])
6| Call | What happens | Result |
|---|---|---|
| max with separate arguments | gathers them | works |
| max with a sequence | also works | it accepts either |
| sum with separate arguments | does not gather | TypeError |
Note that max and min gather.
Why: Many of the built-in functions use variable-length argument tuples, and max and min can take any number of arguments.
Note that sum does not.
Why: sum expects one sequence, so passing three separate numbers raises a TypeError naming the argument count.
Draw the practical conclusion.
Why: You cannot tell from the outside which a function does — the error message is the usual way people find out.
Figure (svg): Two columns separating built-ins that gather arguments from those requiring a sequence
max accepts either form and sum accepts only a sequence. The inconsistency is real, and it is the reason the book sets sum_all as an exercise.
Verify: Convert between the forms with the scatter operator.
Why: sum([1, 2, 3]) fails and max([1, 2, 3]) works — because scatter turns a sequence into separate arguments, which is what max accepts and sum does not. The next idea is exactly this operation, and it is the tool for converting between the two calling conventions.
Trap
A program has a list of numbers and calls max(numbers_list) — then later has three separate numbers and calls sum(a, b, c).
Assume both built-ins take the same shape
Why: They both reduce a collection to one number, so they look interchangeable.
The first works by luck — max accepts both forms — and the second raises TypeError: sum expected at most 2 arguments, got 3. The habit is only sometimes right, which is worse than being always wrong.
Check which convention the function uses, and convert if needed.
A sequence into a gathering function: scatter it
Why: max(*t) spreads the tuple into separate arguments.
Separate values into a sequence-taking function: collect them
Why: sum([a, b, c]) or sum((a, b, c)).
The error message names the count, which is the tell: expected at most 2 arguments, got 3 means the function wanted a sequence and received separate values.
Faded example
One character in the parameter list.
Fill in the blanks
def sum_all(*****args):
total = 0
for x in args:
total += x
return total
Why: The star makes args a gather parameter, so it collects however many arguments arrive into one tuple — including none at all, in which case args is the empty tuple and the function correctly returns 0. Without the star, the function would take exactly one argument and sum_all(1, 2, 3) would raise a TypeError.
Sorting
max and min gather; sum does not.
Sort into buckets
For each call, does it work as written?
Explain it to yourself
It could have been a list. It is not.
Discussion prompt
Gathered arguments arrive as a tuple rather than a list. Why is that the sensible choice?
Hint: What would it mean for a function to modify the thing it was called with?
Answer:
The arguments are a fixed record of what the caller passed — a shape, not a collection the function is meant to grow. A tuple says that.
And if it were a list, a function could append to it or sort it, which would be modifying something the caller never handed over as a modifiable object. Chapter 10's aliasing surprises would apply to argument lists.
So the immutability is a guarantee about the call itself: whatever the function does, the record of what it was called with stays intact. If a function genuinely needs to modify the collection, list(args) makes a copy it owns.
Section
Section 2
Concept
The complement of gather is scatter. If you have a sequence of values and you want to pass it to a function as multiple arguments, you can use the * operator.
>>> t = (7, 3)
>>> divmod(t)
TypeError: divmod expected 2 arguments, got 1
>>> divmod(*t)
(2, 1)| Call | What arrives | Result |
|---|---|---|
| divmod(t) | one argument: a tuple | TypeError |
| divmod(*t) | the tuple is spread | two arguments |
| the result | quotient and remainder | (2, 1) |
divmod takes exactly two arguments, so it doesn't work with a tuple — passing one gives it a single argument that happens to contain two things. Scattering spreads them, and the function sees the two arguments it wanted.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-118
Picture it
The mirror image of the previous idea's picture.
Figure (svg): A pipeline showing a tuple scattered into two separate arguments at a call site
Compare with the gather picture: the same stages, read in the opposite direction. That symmetry is why the book presents them together.
Worked example
Gather then scatter, and you are back where you started.
def show(*args):
print(args)
printall(*args) # scatter them straight back out
def printall(a, b, c):
print(a, b, c)
>>> show(1, 2, 3)
(1, 2, 3)
1 2 3| Step | Direction | Note |
|---|---|---|
| show(1, 2, 3) | gathers into a tuple | (1, 2, 3) |
| printall(*args) | scatters it back | three arguments |
| printall | takes exactly three | and receives three |
Gather at the definition.
Why: show accepts any number of arguments and holds them in one tuple.
Scatter at the call.
Why: printall wants exactly three separate arguments, and *args supplies them from the tuple.
Notice the two stars do opposite things.
Why: One packed and one unpacked, and the code is back to the shape it started in.
Figure (svg): Two columns showing the same star packing in a definition and unpacking at a call
The tuple, then the three values. The two operators are exact complements, which is why forwarding arguments to another function is written this way.
Verify: Remove the star from the inner call.
Why: printall(args) passes one argument — the tuple — and raises TypeError: missing 2 required positional arguments. That failure is the same one as divmod(t), and it confirms that without the star the tuple stays a single value.
Prediction
divmod takes exactly two arguments.
t = (7, 3)
print(divmod(*t))| Part | What happens | Result |
|---|---|---|
| *t | spreads the tuple | two arguments |
| divmod(7, 3) | the call that results | quotient and remainder |
| the result | (2, 1) | a tuple again |
Predict first
What does this print?
Correct: (2, 1) — the scatter operator spreads the tuple into the two arguments divmod expects.
Why: Without the star, divmod(t) passes one argument — a tuple — and raises TypeError: divmod expected 2 arguments, got 1. The star converts the sequence into separate arguments. Note that divmod's own result is a tuple too, which is why the printed answer has parentheses.
Worked example
The conversion in the other direction.
>>> t = (1, 2, 3)
>>> max(*t) # max gathers, so scattering is fine
3
>>> sum(*t)
TypeError: sum expected at most 2 arguments, got 3
>>> sum(t) # sum wants the sequence itself
6| Call | What arrives | Result |
|---|---|---|
| max(*t) | three separate arguments | max gathers them |
| sum(*t) | three separate arguments | sum wanted one sequence |
| sum(t) | one sequence | correct |
Ask what the function wants.
Why: max gathers, so it is happy with separate arguments; sum expects a single sequence.
Scatter only into functions that gather.
Why: The star helps when the function wants separate values and you have a sequence.
Do nothing when the shapes already match.
Why: sum(t) needs no operator at all, because t is already the sequence sum wanted.
Figure (svg): The state of the program after each line of Worked example making sum work with separate values, drawn as a ladder with one rung per traced line
Scattering helps for max and hurts for sum. The operator is a converter, and using it when no conversion is needed creates the mismatch it was meant to fix.
Verify: Read the two error messages side by side.
Why: divmod(t) says expected 2 arguments, got 1 — too few, because a sequence was passed whole. sum(*t) says expected at most 2, got 3 — too many, because a sequence was spread. The direction of the count tells you which correction to make.
Trap
A program has a list and a function that takes a list, and the star is added anyway because it is being passed to a function.
Use the star whenever a sequence is being passed
Why: It appeared in the working example, so it looks like part of the calling syntax.
The sequence is spread into separate arguments the function never asked for, and it raises about the argument count. Nothing was wrong before the star was added.
Use it only to convert between the two shapes.
Ask what the function wants and what you have
Why: The star is needed only when those differ: you have one sequence and it wants separate values.
Read the count in the error message
Why: Too few arguments means you should have scattered; too many means you should not have.
The star is not decoration on a call — it is an operation with a direction, and applying it when the shapes already match breaks a call that was correct.
Discrimination
The direction depends on where the star appears.
Sort into buckets
For each occurrence of the star, which operation is it?
Faded example
divmod wants two arguments and you have one tuple.
Fill in the blanks
t = (7, 3)
quot, rem = divmod(*t)
print(quot, rem) # 2 1
Why: The star scatters the tuple into two separate arguments. Writing divmod(t) instead raises TypeError: divmod expected 2 arguments, got 1, because a tuple passed whole is a single argument however many elements it contains. Note the result is then unpacked by tuple assignment, which is the previous lesson's idiom.
Explain it
The star seems to do opposite things, and there is a rule.
Discussion prompt
A classmate finds it confusing that * packs in one place and unpacks in another. Give them a rule that makes it predictable.
Hint: Where is it written?
Answer:
The rule is position. In a definition it gathers, because that is where arguments arrive; at a call site it scatters, because that is where arguments depart.
Said another way, the star always sits on the side where the many things are, and points at the one tuple. In a def, many arguments become the tuple; in a call, the tuple becomes many arguments.
So it is one idea rather than two: the star marks the boundary between a sequence and a list of arguments, and which way you cross it depends on which side you are standing on.
Section
Section 3
Concept
zip is a built-in function that takes two or more sequences and interleaves them. The name of the function refers to a zipper, which interleaves two rows of teeth.
iterator — An object that iterates through a sequence, but which does not provide list operators and methods.
>>> s = 'abc'
>>> t = [0, 1, 2]
>>> zip(s, t)
<zip object at 0x7f7d0a9e7c48>
>>> list(zip(s, t))
[('a', 0), ('b', 1), ('c', 2)]| Expression | What you get | Note |
|---|---|---|
| zip(s, t) | a zip object | not a list |
| printing it | shows an address | unhelpful |
| list(zip(s, t)) | a list of tuples | each pairs one element from each |
A zip object is a kind of iterator. Iterators are similar to lists in some ways, but unlike lists, you can't use an index to select an element from an iterator.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-119
Picture it
One element from each sequence, paired in order.
Figure (svg): Two sequences aligned with each pair of corresponding elements joined into a tuple
Note the two sequences need not be the same type — a string and a list zip together perfectly well, because zip only needs to walk them.
Worked example
The first surprise everyone meets.
>>> z = zip('abc', [0, 1, 2])
>>> z
<zip object at 0x7f7d0a9e7c48>
>>> z[0]
TypeError: 'zip' object is not subscriptable
>>> list(z)
[('a', 0), ('b', 1), ('c', 2)]| Expression | What happens | Note |
|---|---|---|
| printing z | an object description | not its contents |
| z[0] | indexing an iterator | TypeError |
| list(z) | walks it and collects | a real list |
Recognise what you have.
Why: The result is a zip object that knows how to iterate through the pairs — it is not the pairs themselves.
Note what it cannot do.
Why: Unlike lists, you can't use an index to select an element from an iterator, so z[0] raises.
Convert if you need list operations.
Why: If you want to use list operators and methods, you can use a zip object to make a list.
Figure (svg): A panel showing the three things people try with a zip object and which of them work
An unhelpful description, a TypeError, and then the pairs. The zip object holds the recipe rather than the result.
Verify: Convert the same zip object twice.
Why: The second list(z) comes back empty, because the iterator was used up by the first. That is a genuine trap and it is why the common use is directly in a for loop, where it is walked exactly once and there is no temptation to reuse it.
Prediction
The sequences are different lengths.
print(list(zip('Anne', 'Elk')))| Sequence | Length | Effect |
|---|---|---|
| 'Anne' | 4 characters | one is unused |
| 'Elk' | 3 characters | the shorter |
| the result | 3 pairs | silently truncated |
Predict first
How many pairs does this produce?
Correct: 3 — if the sequences are not the same length, the result has the length of the shorter one.
Why: zip stops as soon as either sequence runs out, so the final 'e' of 'Anne' is dropped without comment. That silence is convenient when the truncation is intended and dangerous when it is not — which is why checking the lengths first is worth doing whenever the two sequences are supposed to correspond exactly.
Worked example
zip stops when the shorter sequence runs out.
>>> list(zip('Anne', 'Elk'))
[('A', 'E'), ('n', 'l'), ('n', 'k')]| Sequence | Length | Note |
|---|---|---|
| 'Anne' | four characters | one left over |
| 'Elk' | three characters | the shorter |
| the result | three pairs | the length of the shorter |
Compare the lengths.
Why: One sequence has four elements and the other three.
Apply the rule.
Why: If the sequences are not the same length, the result has the length of the shorter one.
Note what happens to the extra.
Why: The final 'e' of 'Anne' is simply not included, silently and without error.
Figure (svg): The state of the program after each line of Worked example the unequal-length rule, drawn as a ladder with one rung per traced line
Three pairs. The rule is convenient and quiet, which means a length mismatch you did not intend will not announce itself.
Verify: Check the length before zipping when the counts should match.
Why: If two sequences are supposed to correspond element by element, asserting len(a) == len(b) first turns a silent truncation into an immediate failure — a sanity check of exactly the kind lesson 11c described, and one worth having whenever the pairing is meant to be complete.
Trap
A program stores z = zip(a, b), loops over it to count the pairs, then loops again to print them.
Treat it as a collection
Why: It behaves like one in the first loop, so it looks like a list.
The second loop produces nothing at all, because the iterator was consumed by the first. There is no error — the loop body simply never runs, and the program reports zero of something it just counted.
Use it once, or make a list.
The most common use of zip is directly in a for loop
Why: Where it is walked once and never named.
If you need it twice, convert first
Why: pairs = list(zip(a, b)) gives a real list you can traverse as often as you like.
The symptom is distinctive: a loop that works and a later identical loop that does nothing. Whenever that happens, the thing being looped over is an iterator that has already been spent.
Error analysis
Mark each and say what happens.
Annotate
The consumed-iterator failure in line 3 is the one to remember, because it produces no error and no output.
Comparison
Fill the blanks. They are similar in one way and different in three.
Comparison matrix
| Question | zip object | list |
|---|---|---|
| Can you loop over it? | yes | yes |
| Can you index it? | no — TypeError | yes |
| Can you loop over it twice? | no — it is consumed | yes, as often as you like |
| What does printing it show? | an object description | its contents |
The first row is why zip works in a for loop, and the rest are why it surprises people everywhere else.
Socratic
It would be simpler. There is a reason it is not.
Discussion prompt
zip could have returned a list of pairs directly. What does returning an iterator buy, and when would the difference matter?
Hint: How much has to exist at once?
Answer:
A list of pairs has to be built in full before you see the first one. An iterator produces each pair as it is asked for, so only one exists at a time.
For three elements that is irrelevant. For a million it is the difference between a program that runs and one that exhausts memory building something it was going to walk through once anyway.
So the iterator is the more general design, and list(zip(...)) is there for when you genuinely need a list. The cost is the surprises — no indexing, and consumed after one pass — which is a fair trade for something usually used directly in a for loop.
Section
Section 4
Concept
You can use tuple assignment in a for loop to traverse a list of tuples. Each time through the loop, Python selects the next tuple in the list and assigns the elements to the loop variables.
t = [('a', 0), ('b', 1), ('c', 2)]
for letter, number in t:
print(number, letter)
# 0 a
# 1 b
# 2 c| Pass | The tuple | The names |
|---|---|---|
| pass 1 | ('a', 0) unpacked | letter='a', number=0 |
| pass 2 | ('b', 1) unpacked | letter='b', number=1 |
| pass 3 | ('c', 2) unpacked | letter='c', number=2 |
This is the previous lesson's tuple assignment, happening once per pass. The two names on the left of in must match the number of elements in each tuple, exactly as in an ordinary unpacking.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 119-119
Picture it
The loop supplies a tuple and the assignment splits it.
Figure (svg): A pipeline showing a list of tuples yielding one tuple per pass which is unpacked into two names
Which is why the book put tuple assignment three sections earlier: this idiom is the payoff.
Worked example
zip, for, and tuple assignment together.
def has_match(t1, t2):
for x, y in zip(t1, t2):
if x == y:
return True
return False| Part | What it does | Note |
|---|---|---|
| zip(t1, t2) | pairs the sequences | one pair per pass |
| x, y | unpacked from each pair | two names |
| if x == y | a match at this index | return True |
| after the loop | no match anywhere | return False |
Pair the sequences with zip.
Why: Each pass gives one tuple holding the elements at the same index in both sequences.
Unpack it in the for statement.
Why: x and y are the two elements, named directly rather than fetched with an index.
Apply the search pattern.
Why: Return True on the first match; if the loop finishes, return False — the shape from chapter 8, and from reverse_lookup.
Figure (svg): A flowchart showing the paired traversal and the two exits of has_match
A function that returns True if there is an index i such that t1[i] == t2[i]. If you combine zip, for and tuple assignment, you get a useful idiom for traversing two or more sequences at the same time.
Verify: Compare with the index-based version.
Why: Writing it with for i in range(len(t1)) and comparing t1[i] with t2[i] gives the same answer, needs a length check to avoid an IndexError, and mentions the index three times without ever using it as a number. The zip version says what is meant, which is the argument for the idiom.
Prediction
The names are unpacked in the order they appear.
t = [('a', 0), ('b', 1)]
for letter, number in t:
print(number, letter)| Pass | The names | Output |
|---|---|---|
| pass 1 | letter='a', number=0 | prints 0 a |
| pass 2 | letter='b', number=1 | prints 1 b |
| the print | number first | the reverse of the tuple |
Predict first
What does this print?
Correct: 0 a and then 1 b — the names are bound by position and the print statement reverses them.
Why: Each tuple is unpacked into letter and number by position, and the print statement happens to name them in the other order. Option C is what for item in t would give, binding one name to the whole tuple. The counts match here, so nothing raises.
Worked example
The same idiom, on a source you already have.
h = {'a': 1, 'b': 2}
for key, value in h.items():
print(key, value)
# a 1
# b 2| Part | What it gives | Note |
|---|---|---|
| h.items() | yields pairs | one per item |
| key, value | unpacked | two names per pass |
| the body | uses both | no lookup needed |
Ask for the pairs.
Why: items yields a key-value pair per pass, rather than the keys that a bare dictionary gives.
Unpack them in the for statement.
Why: Exactly the same tuple assignment as for a list of tuples — the source does not matter.
Compare with the lookup version.
Why: for c in h: print(c, h[c]) does the same thing with an extra lookup per pass, which lesson 11b showed.
Figure (svg): The state of the program after each line of Worked example unpacking a dictionary's items, drawn as a ladder with one rung per traced line
Each key with its value, unpacked directly. The idiom applies to anything that yields tuples, which is why it is worth learning as a shape rather than as a zip trick.
Verify: Try it with the wrong number of names.
Why: for k in h.items() binds k to the whole tuple, so printing it shows ('a', 1) rather than the parts — legal and probably not what was wanted. Three names would raise ValueError instead. The count is what determines whether the pairs are taken apart.
Trap
A loop unpacks into two names, and one of the tuples in the list has three elements.
Assume every item has the same shape
Why: It does for all the data seen so far.
The loop raises ValueError: too many values to unpack partway through, after processing some of the items. The failure is loud but the program has already done part of its work.
Make the shape a stated assumption.
Check the shape before the loop if the data are untrusted
Why: One pass to validate, then one to process — which fails before anything is half-done.
Or unpack inside the body, where you can handle the mismatch
Why: for item in t, then check len(item) before taking it apart.
Tuple assignment in a for loop is ordinary tuple assignment, so the same counting rule applies once per pass — and a failure on the tenth item leaves the first nine already processed.
Faded example
The idiom is zip plus tuple assignment.
Fill in the blanks
def has_match(t1, t2):
for x, y in zip(t1, t2):
if x == y:
return True
return False
Why: zip pairs the two sequences element by element, and the tuple assignment in the for statement unpacks each pair into x and y. This is the book's useful idiom for traversing two or more sequences at the same time, and it avoids both the index arithmetic and the IndexError risk of a range-based version.
Ranking
Four steps, in order.
Put in order
Why: The loop supplies a tuple, the unpacking checks that the number of elements matches the number of names, the names are bound, and the body runs. The count check happens once per pass, which is why a mismatched item partway through a list fails partway through the loop.
Real world
Two lists that correspond position by position are everywhere.
Discussion prompt
Think of two lists where the nth item of one belongs with the nth item of the other. What goes wrong if they fall out of step?
Hint: Names and scores, questions and answers, labels and values.
Answer:
Names and their scores, questions and answers, timestamps and readings — in each case the correspondence is by position and nothing in the data records it.
If one list gains or loses an item, every pair after that point is wrong, and nothing detects it: the data are still well-formed and the program still runs.
zip makes the pairing explicit in the code, which helps, but it also truncates silently to the shorter list — so the safest version checks the lengths first. Where the correspondence really matters, a single list of pairs is safer than two parallel lists, because then they cannot fall out of step at all.
Section
Section 5
Concept
If you need to traverse the elements of a sequence and their indices, you can use the built-in function enumerate.
for index, element in enumerate('abc'):
print(index, element)
# 0 a
# 1 b
# 2 c| Part | What it gives | Note |
|---|---|---|
| enumerate('abc') | pairs of index and element | starting from 0 |
| index, element | unpacked per pass | two names |
| the body | uses both | no manual counter |
The result from enumerate is an enumerate object, which iterates a sequence of pairs; each pair contains an index — starting from 0 — and an element from the given sequence. Like a zip object, it is an iterator.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 119-119
Picture it
Only one of them says what is meant.
Figure (svg): Two columns comparing a manual counter and a range loop against enumerate
The range version is the one to watch for: it names an index only to immediately index with it, which is a sign that enumerate was wanted.
Worked example
The counter version works and has two places to go wrong.
# a manual counter
i = 0
for c in s:
print(i, c)
i += 1
# the same, with enumerate
for i, c in enumerate(s):
print(i, c)| Version | What it costs | Risk |
|---|---|---|
| the counter version | five lines | and i must be advanced |
| forgetting i += 1 | every index is 0 | silent |
| enumerate | two lines | nothing to forget |
Look at what the counter version maintains.
Why: An index that has to be initialised before the loop and advanced at the end of every pass.
Consider the failure modes.
Why: Forgetting the increment gives every element the index 0; advancing in the wrong place gives an off-by-one; an early continue skips the increment entirely.
Replace it with enumerate.
Why: The index and the element arrive together, and there is nothing left to maintain.
Figure (svg): A ladder showing the index and element produced on each pass of an enumerate loop
Identical output, and three fewer opportunities for an off-by-one. The counter was doing work the language will do correctly.
Verify: Add a continue to both versions.
Why: In the counter version, a continue before the increment silently stops the index advancing and every subsequent pass is wrong; in the enumerate version nothing changes, because the index comes from the iteration itself. That difference is the strongest argument for the idiom.
Prediction
Two names, one sequence.
for i, c in enumerate('ab'):
print(i, c)| Pass | The pair | Output |
|---|---|---|
| pass 1 | (0, 'a') | prints 0 a |
| pass 2 | (1, 'b') | prints 1 b |
| the index | starts from 0 | as everywhere in Python |
Predict first
What does this print?
Correct: 0 a and then 1 b — each pair contains an index, starting from 0, and an element from the sequence.
Why: The index comes first in the pair, and it starts at 0 like every other index in Python. Option C reverses the pair, which would need the names swapped in the for statement; option D is what for pair in enumerate(...) would bind, keeping each pair whole.
Worked example
enumerate is for when the position itself is part of the answer.
def find_all(t, target):
positions = []
for i, x in enumerate(t):
if x == target:
positions.append(i)
return positions
>>> find_all(['a', 'b', 'a'], 'a')
[0, 2]| Part | Which name it uses | Note |
|---|---|---|
| enumerate(t) | index and element | both needed |
| if x == target | uses the element | the comparison |
| positions.append(i) | uses the index | the answer |
Notice both names are used.
Why: The element for the comparison and the index for the result — which is the case enumerate exists for.
Compare with a plain loop.
Why: for x in t gives no index, so the positions could not be recorded without maintaining a counter.
Note it returns all matches.
Why: Unlike the search pattern, which returns on the first, this collects every one — the accumulator pattern from lesson 11b.
Figure (svg): The state of the program after each line of Worked example when you genuinely need the index, drawn as a ladder with one rung per traced line
[0, 2]. When the index is part of what you are computing rather than only a way of reaching elements, enumerate is exactly right.
Verify: Ask when enumerate is not needed.
Why: Whenever the index is never used: for x in t is clearer and cannot go wrong. Using enumerate and ignoring the index is as much a mismatch as writing range(len(t)) and only indexing with it — in each case the loop is asking for something it does not want.
Trap
A loop is written as for i in range(len(t)) and every use of i is immediately t[i].
Loop over positions and index into the sequence
Why: Which is how loops work in many other languages.
The index is named and never used as a number, so the loop is doing bookkeeping for nothing — and it introduces an IndexError risk that for x in t simply does not have.
Loop over what you actually want.
Only the elements: for x in t
Why: Which is the chapter 8 traversal, and cannot go out of range.
Both: for i, x in enumerate(t)
Why: Which gives the index without maintaining it.
The test is whether the index appears anywhere except inside a bracket. If it does not, it was never wanted, and range(len(t)) is a habit carried in from another language.
Discrimination
Ask what the body actually uses.
Sort into buckets
For each task, which loop form is right?
Faded example
No counter, and no range.
Fill in the blanks
for index, element in enumerate('abc'):
print(index, element)
Why: enumerate yields a pair per pass — an index starting from 0 and the corresponding element — which the tuple assignment unpacks into two names. The alternatives both do more work: a manual counter must be advanced by hand, and range(len(...)) gives only the index, leaving you to fetch the element yourself.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: Both return iterators, not lists. You cannot index either one — z[0] raises TypeError — and each is consumed by a single traversal, so a second loop over the same object produces nothing at all. If you need list behaviour, wrap it: list(zip(a, b)) gives a real list of tuples you can use as often as you like.
Comparison
Fill the blanks. One symbol, two directions.
Comparison matrix
| Question | Gather | Scatter |
|---|---|---|
| Where is the star written? | in the definition | at the call site |
| Which direction? | many arguments into one tuple | one tuple into many arguments |
| An example | def printall(*args) | divmod(*t) |
| What goes wrong without it? | the function takes exactly one argument | the tuple arrives as a single argument |
The star always sits on the side where the many things are. Which side that is depends on whether you are defining or calling.
Pattern
Five steps, and the second is the one people skip.
Step 2 exists because the truncation is silent. A length mismatch you did not intend produces a shorter, plausible result rather than an error.
Python documentation — More Control Flow Tools More Control Flow Tools
Check
Three arguments, one starred parameter.
def f(*args):
return len(args)
print(f(1, 2, 3))| Part | What happens | Result |
|---|---|---|
| three arguments | gathered into a tuple | (1, 2, 3) |
| len(args) | the tuple's length | 3 |
| with no arguments | args is () | would give 0 |
Check your understanding
What does this print?
Answer: A
Why: A parameter name beginning with * gathers the arguments into a tuple, so args is (1, 2, 3) and its length is 3. Without the star the function would take exactly one argument and this call would raise the TypeError in option C.
Check
Unequal lengths.
print(len(list(zip([1, 2, 3, 4], 'ab'))))| Sequence | Length | Effect |
|---|---|---|
| the list | 4 elements | longer |
| the string | 2 characters | shorter |
| zip | stops at the shorter | 2 pairs |
Check your understanding
What does this print?
Answer: A
Why: If the sequences are not the same length, the result has the length of the shorter one — so two pairs, and the last two numbers are dropped without comment. That silence is why checking lengths first is worth doing whenever the pairing is meant to be complete.
Check
divmod takes exactly two arguments.
Check your understanding
You have t = (7, 3). Which call gives (2, 1)?
Answer: B
Why: The scatter operator spreads the tuple into the two separate arguments divmod expects. Option D would also work and says the same thing at more length; the star is the general form, which keeps working when the tuple's length changes.
Real world
Packing and unpacking is a shape that recurs far outside programming.
Discussion prompt
Think of somewhere a variable number of things is bundled into one item and later taken apart again — a parcel, a form, a message. What has to be agreed for the unpacking to work?
Hint: What does the receiver need to know?
Answer:
The count and the order. A parcel of six items unpacked expecting five leaves someone holding an extra, and unpacking in the wrong order puts things in the wrong places.
Which is exactly the two failure modes of tuple assignment: a count mismatch raises ValueError, and a wrong order binds silently and produces confident nonsense.
The systems that handle this well are the ones that label rather than rely on position — which is what keyword arguments do, and why the double star exists for dictionaries. Position is compact and fragile; labels are verbose and robust.
Commit first
Answer, then rate your confidence.
Predict first
You write z = zip(a, b), then list(z) twice in a row. What does the second call return?
Correct: An empty list — the iterator was consumed by the first call.
Why: A zip object is a kind of iterator: it knows how to walk the pairs, and once walked there is nothing left. The second conversion finds an exhausted iterator and returns nothing, with no error at all. This is the quietest bug in the lesson, because the symptom is a loop that produces no output rather than a failure — and it is why the most common use of zip is directly in a for loop, where it is walked once and never stored. If you need the pairs more than once, convert immediately with pairs = list(zip(a, b)) and use that list as often as you like.
Explain it
Two idioms that replace a whole family of index bugs.
Discussion prompt
A classmate writes every loop as for i in range(len(t)). Show them the two idioms from this lesson and say what each one removes.
Hint: What is the index actually used for in their loops?
Answer:
If the index only ever appears inside brackets, they wanted for x in t — which cannot go out of range and does not mention a number at all.
If they need the position too, enumerate gives both names at once, with no counter to initialise or advance and nothing that a continue statement can knock out of step.
And if they are indexing two sequences with the same i, that is zip: for x, y in zip(t1, t2). It states the correspondence in the code rather than leaving it implicit in an index used twice — though it does truncate to the shorter, so the lengths are worth checking.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The two star operators are one idea seen from two sides, and the rule about position — definition gathers, call scatters — makes them predictable rather than something to memorise. The iterator point is where nearly everyone is caught once, and the consumed-iterator failure is silent, so it is worth over-learning. And the loop idioms are the part you will use every day: once for x, y in zip(...) and for i, x in enumerate(...) are automatic, a whole family of index bugs stops being possible.
Connect it up
One page, from memory.
Draw it
Draw two arrows facing each other: one labelled gather going from several arguments into one tuple, one labelled scatter going the other way, and write beside each where the star is written. Underneath, write the three loop forms — for x in t, for i, x in enumerate(t), and for x, y in zip(t1, t2) — and next to each, what the body would have to do without it. Finally write down the three things a zip object cannot do that a list can.
Recap
Two pages, and the idioms you will use for the rest of the book.
| If you remember one thing | It is this |
|---|---|
| From gather | The starred parameter is an ordinary tuple inside the function. |
| From scatter | The star sits on the side where the many things are. |
| From zip | It stops at the shorter sequence, and says nothing about it. |
| From iterators | Walk it once. A second pass finds nothing and raises no error. |
| From enumerate | If the index never leaves a bracket, you did not want it. |
The next lesson combines all three types: dictionaries with tuple keys, tuples as items, and the sorting idiom that inverts a histogram in one line — plus the state diagrams for sequences of sequences.
Think Python, 2nd edition — Allen B. Downey §12.4-12.5, pp. 118-119 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.