This lesson converts between dictionaries and lists of tuples, uses tuples as dictionary keys, gives the criteria for choosing between strings, lists and tuples, and introduces shape errors in compound data structures.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 12 — Tuples
§12.6-12.8, pp. 120-122
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-122 — the pages these objectives are drawn from
Warm-up
You want to look someone up by surname and forename together.
Discussion prompt
A telephone directory maps a person to a number, and a person is a surname and a forename. What would you use as the dictionary key, and what would have gone wrong if you had tried this two chapters ago?
Hint: The key has to be one thing, and it cannot be a list.
Answer:
The key has to be a single value holding both names — a sequence of two — and it cannot be a list, because keys must be hashable.
Two chapters ago there was no such type. You would have had to join the names into a string with a separator and hope no name contained it.
A tuple solves it exactly: two values, one object, immutable and therefore hashable. This is the payoff the last two chapters have been building towards.
Concept
It is common to use tuples as keys in dictionaries, primarily because you can't use lists. A telephone directory might map from last-name, first-name pairs to telephone numbers.
directory[last, first] = number
for last, first in directory:
print(first, last, directory[last, first])| Line | What it means | Note |
|---|---|---|
| directory[last, first] | the brackets hold a TUPLE | not two indices |
| for last, first in directory | the keys are tuples | unpacked per pass |
| the lookup inside | the same tuple key | rebuilt from the two names |
The expression in brackets is a tuple. This is not two-dimensional indexing, however much it looks like it — Python sees the comma, builds a tuple, and uses that one object as the key.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120
Section
Section 1
Concept
Dictionaries have a method called items that returns a sequence of tuples, where each tuple is a key-value pair.
>>> d = {'a': 0, 'b': 1, 'c': 2}
>>> t = d.items()
>>> t
dict_items([('c', 2), ('a', 0), ('b', 1)])
>>> for key, value in d.items():
... print(key, value)
c 2
a 0
b 1| Part | What it gives | Note |
|---|---|---|
| d.items() | a dict_items object | an iterator |
| each element | a key-value pair | as a tuple |
| the for loop | unpacks each pair | two names per pass |
The result is a dict_items object, which is an iterator that iterates the key-value pairs — the same kind of thing as a zip object. And as you should expect from a dictionary, the items are in no particular order.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120
Picture it
The same three items, seen as a sequence of tuples.
Figure (svg): A dictionary shown alongside the sequence of key-value tuples that items produces
Once a dictionary is a sequence of tuples, everything the last two lessons taught about tuples applies to it — including sorting, unpacking and comparison.
Worked example
The same output, with different amounts of work.
# keys, then look each value up
for c in d:
print(c, d[c])
# the pairs, kept whole
for pair in d.items():
print(pair[0], pair[1])
# the pairs, unpacked
for key, value in d.items():
print(key, value)| Form | What each pass gives | Verdict |
|---|---|---|
| the first | a lookup per pass | works, does extra work |
| the second | a tuple per pass, indexed | works, reads badly |
| the third | two names per pass | says what it means |
Start with what lesson 11b gave you.
Why: Looping over the dictionary yields keys, and each value needs a lookup — correct, and one operation more than necessary.
Ask for the pairs instead.
Why: items yields tuples, so both halves arrive together and no lookup is needed.
Unpack them in the for statement.
Why: Tuple assignment names the two halves directly, which removes the indices as well.
Figure (svg): The state of the program after each line of Worked example three ways to traverse a dictionary, drawn as a ladder with one rung per traced line
All three print the same thing. The third does the least work and reads the closest to its intention, which is why it is the standard form.
Verify: Check that the order is the same in all three.
Why: It is unpredictable in all three, and identical between them for a given dictionary — because all three walk the same items. items does not impose an order any more than looping over the keys does, which is worth confirming before relying on either.
Prediction
Two names on the left of in.
d = {'a': 0}
for key, value in d.items():
print(key, value)| Part | What happens | Output |
|---|---|---|
| d.items() | yields ('a', 0) | one pair |
| key, value | unpacked by position | 'a' and 0 |
| both, separated | a 0 |
Predict first
What does this print?
Correct: a 0 — items yields a key-value tuple, which the two names unpack in order.
Why: Each element of items is a tuple whose first component is the key, so key binds to 'a' and value to 0. Writing for pair in d.items() instead would bind one name to the whole tuple and print option B. Looping over the dictionary directly, without items, would give option C — just the keys.
Worked example
Now that the items are tuples, the comparison rule applies.
>>> d = {'b': 1, 'a': 3, 'c': 2}
>>> sorted(d.items())
[('a', 3), ('b', 1), ('c', 2)]
>>> h = {'a': 3, 'b': 1}
>>> sorted([(v, k) for k, v in h.items()])
[(1, 'b'), (3, 'a')]| Expression | What it sorts by | Note |
|---|---|---|
| sorted(d.items()) | sorts the tuples | by key, since it comes first |
| swapping the pairs | value first | sorts by value instead |
| the rule | position 0 dominates | lesson 12a's comparison |
Sort the items directly.
Why: Each item is a tuple, and tuples compare element by element — so the key at position 0 decides the order.
Swap to sort by value.
Why: Building pairs the other way round puts the value at position 0, which makes it the primary ordering.
Note what makes this work.
Why: Nothing about dictionaries: it is the sequence comparison rule from lesson 12a applied to the pairs that items produces.
Figure (svg): A pipeline showing a dictionary becoming pairs, being swapped, and then sorted by value
Sorted by key in the first case and by value in the second. The position of the field inside the tuple is what chooses the ordering.
Verify: Check what sorted returns.
Why: A list, not a dictionary — sorted always returns a list whatever it was given. There is no such thing as a sorted dictionary, for the reason lesson 11a gave, so what you have is a sequence of pairs in an order you chose.
Trap
A program stores pairs = d.items() and later indexes it with pairs[0].
Treat it as the list it prints like
Why: dict_items([('c', 2), ...]) looks very much like a list of tuples.
It is an iterator, so indexing raises a TypeError — the same restriction as a zip object, for the same reason.
Use it in a for loop, or convert it.
for key, value in d.items()
Why: The intended use, and the one that needs no conversion.
list(d.items()) if you need list operations
Why: Which gives a real list of tuples you can index and sort.
The printed form is genuinely misleading here — dict_items wraps what looks like a list — so the type name in the output is the thing to read, not the brackets inside it.
Faded example
Both halves at once, with no lookup.
Fill in the blanks
for key, value in d.items():
print(key, value)
Why: items returns a sequence of tuples, each a key-value pair, which the tuple assignment unpacks into two names. Looping over d alone would give only the keys and need d[key] inside the body; d.values() would give only the values and no way back to the keys.
Discrimination
Three methods, three views of the same dictionary.
Sort into buckets
For each expression, what does one pass of a loop over it give?
Explain it to yourself
A dictionary cannot be sorted. Its items can.
Discussion prompt
Explain why sorted(d.items()) is meaningful when sorting a dictionary is not.
Hint: What does sorted return?
Answer:
Sorting means arranging things in an order, and a dictionary has no arrangement — its storage location comes from hashing the key, so there is nowhere to put an order.
items produces tuples, and a list of tuples is an ordinary sequence with positions. Sorting that is perfectly meaningful.
So sorted(d.items()) does not sort the dictionary; it produces a new list, ordered, from the dictionary's contents. The dictionary is untouched, which is the returning-not-modifying distinction again.
Section
Section 2
Concept
Going in the other direction, you can use a list of tuples to initialize a new dictionary. And combining dict with zip yields a concise way to create one.
>>> t = [('a', 0), ('c', 2), ('b', 1)]
>>> d = dict(t)
>>> d
{'a': 0, 'c': 2, 'b': 1}
>>> d = dict(zip('abc', range(3)))
>>> d
{'a': 0, 'c': 2, 'b': 1}| Expression | What it does | Note |
|---|---|---|
| dict(t) | a list of pairs | becomes a dictionary |
| zip('abc', range(3)) | pairs the two sequences | ('a',0), ('b',1), ('c',2) |
| dict(zip(...)) | the pairs become items | one line |
The dictionary method update also takes a list of tuples and adds them, as key-value pairs, to an existing dictionary — so the same shape works for building one from scratch or extending one you have.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120
Picture it
Dictionary to pairs and back, with a different tool in each direction.
Figure (svg): A pipeline showing a dictionary converted to pairs and pairs converted back to a dictionary
The round trip is worth knowing because so many operations — sorting, filtering, zipping — are natural on a list of pairs and impossible on a dictionary.
Worked example
Two sequences become a mapping in one line.
>>> letters = 'abc'
>>> numbers = range(3)
>>> d = dict(zip(letters, numbers))
>>> d
{'a': 0, 'b': 1, 'c': 2}
>>> d['b']
1| Step | What it produces | Note |
|---|---|---|
| zip | pairs them by position | ('a',0), ('b',1), ('c',2) |
| dict | each pair becomes an item | key from position 0 |
| the lookup | ordinary | 1 |
Pair the sequences.
Why: zip interleaves them into two-element tuples, one per position.
Hand the pairs to dict.
Why: Each tuple becomes one item, with position 0 as the key and position 1 as the value.
Note the direction is fixed.
Why: The first sequence supplies the keys, so swapping the arguments inverts the mapping.
Figure (svg): Two sequences aligned and paired, with each pair becoming a dictionary item
A three-item dictionary, built from two independent sequences without a loop. This is the concise idiom the book names.
Verify: Try it with sequences of different lengths.
Why: dict(zip('abcd', range(2))) gives a two-item dictionary, because zip truncates to the shorter — silently, as lesson 12b warned. So a mismatch produces a smaller dictionary rather than an error, which is worth checking for when the two sequences are meant to correspond exactly.
Prediction
One key appears twice in the list.
t = [('a', 1), ('b', 2), ('a', 3)]
d = dict(t)
print(len(d), d['a'])| Pair | What happens | Size |
|---|---|---|
| three pairs | converted in order | 'a' set to 1 |
| ('b', 2) | a new key | two items |
| ('a', 3) | replaces the earlier 'a' | still two items |
Predict first
What does this print?
Correct: 2 3 — two items, and the later pair for 'a' replaced the earlier one.
Why: Each key is associated with a single value, so the third pair overwrites the first rather than adding an item. Nothing warns you that a pair was absorbed; comparing len(t) with len(d) is the check that finds it, and it is worth running whenever duplicates would be a bug rather than an expectation.
Worked example
A list of pairs can contain the same key twice. A dictionary cannot.
>>> t = [('a', 1), ('b', 2), ('a', 3)]
>>> len(t)
3
>>> d = dict(t)
>>> d
{'a': 3, 'b': 2}
>>> len(d)
2| Stage | What happens | Size |
|---|---|---|
| the list | three pairs | two of them share a key |
| dict(t) | later wins | 'a' maps to 3 |
| the length | 2, not 3 | one pair was absorbed |
Count the pairs.
Why: The list has three, which is a perfectly ordinary list.
Convert and count again.
Why: Each key is associated with a single value, so the second 'a' replaces the first rather than joining it.
Notice nothing is reported.
Why: The conversion succeeds quietly and one pair is gone, which is silent data loss if the duplicate was not intended.
Figure (svg): The state of the program after each line of Worked example what happens to duplicate keys, drawn as a ladder with one rung per traced line
A two-item dictionary. Later pairs win, and nothing tells you a pair was absorbed except the length.
Verify: Compare the lengths as a check.
Why: len(dict(t)) < len(t) means there were duplicate keys, which is a one-line consistency check of exactly the kind lesson 11c described. Where duplicates are expected rather than a bug, invert_dict's list-valued approach keeps them all instead.
Trap
A program builds dict(zip(values, keys)) and then cannot find anything.
Name the arguments after what they hold
Why: Both sequences are being passed, so the order can look arbitrary.
The first sequence supplies the keys, so the mapping is inverted. Lookups fail with KeyError on things that are plainly in the dictionary — as values.
Keys first, always.
dict(zip(keys, values))
Why: Position 0 of each pair becomes the key, and zip preserves the argument order.
Check with one lookup after building
Why: One d[known_key] confirms the direction immediately.
The symptom is distinctive: a KeyError for something you can see in the printed dictionary, on the wrong side of the colon. That is always an inverted mapping.
Faded example
One line, two built-in functions.
Fill in the blanks
d = dict(zip('abc', range(3)))
print(d['b']) # 1
Why: zip pairs the two sequences into tuples and dict turns each pair into an item, with the first sequence supplying the keys. Swapping the arguments would invert the mapping — a mistake whose symptom is a KeyError for something visibly present in the dictionary, on the value side.
Comparison
Fill the blanks. Each direction has its own tool.
Comparison matrix
| Question | Dictionary to pairs | Pairs to dictionary |
|---|---|---|
| Which tool? | d.items() | dict(pairs) |
| What comes out? | an iterator of tuples | a dictionary |
| Can it lose information? | no — every item becomes a pair | yes — duplicate keys silently collapse |
| What about ordering? | unpredictable, as always | the input order is not preserved |
The third row is the asymmetry: pairs allow duplicates and a dictionary does not, so only one direction can lose anything.
Explain it
The round trip looks pointless until you see what happens in the middle.
Discussion prompt
A classmate asks why anyone would turn a dictionary into a list of tuples. Give them two things that are easy on pairs and impossible on a dictionary.
Hint: Both involve order.
Answer:
Sorting. A dictionary has no order and cannot be given one, but a list of tuples sorts naturally — by key, or by value if you swap the pairs round first.
Filtering into a new dictionary. Building the pairs you want and calling dict on them is often clearer than looping and assigning item by item.
The general shape is that dictionaries are for lookup and sequences are for order and position. Converting between them is how you get a structure that is good at the operation you need next.
Section
Section 3
Concept
It is common to use tuples as keys in dictionaries, primarily because you can't use lists. A telephone directory might map from last-name, first-name pairs to telephone numbers.
directory[last, first] = number
for last, first in directory:
print(first, last, directory[last, first])| Line | What happens | Note |
|---|---|---|
| directory[last, first] | the comma builds a tuple | one key |
| for last, first in directory | keys are tuples | unpacked per pass |
| the lookup in the body | the same two names | rebuilds the key |
The expression in brackets is a tuple — the comma is doing the same job it did in lesson 12a. This loop traverses the keys in the directory, which are tuples, assigns the elements of each to last and first, and then prints the name and the corresponding number.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120
Picture it
There are two ways to represent a tuple in a state diagram, and the choice is about detail.
Figure (svg): A state diagram of a directory whose keys are name tuples shown in Python syntax
The detailed version — figure 12.1 — draws each tuple as a box with indices and elements, like a list. In a larger diagram you leave the details out, which is what makes the directory legible.
Worked example
The line looks like a grid access and is not.
>>> directory = {}
>>> directory['Cleese', 'John'] = '08700 100 222'
>>> directory
{('Cleese', 'John'): '08700 100 222'}
>>> directory[('Cleese', 'John')]
'08700 100 222'| Expression | What it means | Note |
|---|---|---|
| the two names in brackets | a comma between them | a tuple |
| the printed dictionary | shows the tuple key | with parentheses |
| the explicit form | identical | parentheses are optional |
Read the comma, not the brackets.
Why: The expression in brackets is a tuple — exactly as it is anywhere else, because the comma is what constructs.
Confirm from the printed form.
Why: The dictionary prints its key with parentheses, showing that one tuple went in rather than two indices.
Note both spellings work.
Why: directory[last, first] and directory[(last, first)] are the same expression; the parentheses only group.
Figure (svg): The state of the program after each line of Worked example it is not two-dimensional indexing, drawn as a ladder with one rung per traced line
One key, which happens to be a pair. There is no two-dimensional indexing here — the brackets take one expression, and the comma made that expression a tuple.
Verify: Try the same thing with a list.
Why: directory[[last, first]] raises TypeError: unhashable type: 'list'. That failure confirms the key is a single object being hashed, and it is the whole reason a tuple is used here — which is the point chapter 11 pointed forward to.
Prediction
Two values inside the brackets.
d = {}
d['Cleese', 'John'] = 1
print(list(d))| Part | What happens | Result |
|---|---|---|
| the comma | builds a tuple | one key |
| the key | ('Cleese', 'John') | a two-element tuple |
| list(d) | the keys | one of them |
Predict first
What does this print?
Correct: [('Cleese', 'John')] — one key, which is a tuple of two strings.
Why: The expression in brackets is a tuple, so the dictionary has exactly one item whose key is that pair. Option B is what two separate keys would look like, and option C is what a grid-style reading would predict. Nothing raises: brackets take one expression, and the comma made it a single tuple.
Worked example
The keys are pairs, so unpack them as you go.
for last, first in directory:
print(first, last, directory[last, first])
# Or, better, with items:
for (last, first), number in directory.items():
print(first, last, number)| Form | What happens | Note |
|---|---|---|
| for last, first in directory | keys are tuples | unpacked into two names |
| directory[last, first] | rebuilds the key | one lookup per pass |
| the items version | nested unpacking | no lookup at all |
Unpack the key in the for statement.
Why: Looping over the dictionary gives keys, and here each key is a two-element tuple, so tuple assignment splits it.
Rebuild the key to look up the value.
Why: The book's version writes directory[last, first], reconstructing the tuple from the two names.
Or take the pairs instead.
Why: items yields (key, value), and the key is itself a tuple — so (last, first), number unpacks both levels at once and needs no lookup.
Figure (svg): A flowchart showing a key-value pair unpacked at two levels
Both print each name with its number. The nested unpacking is the tidier form once you can read it, and it does one less operation per pass.
Verify: Check the parentheses in the nested version.
Why: They are required there: (last, first), number tells Python the first element is itself a pair, whereas last, first, number would ask for three elements and raise ValueError. This is the one place in the chapter where parentheses around a tuple are not optional.
Trap
A student reads it as row-and-column indexing and expects directory[last] to give a sub-dictionary.
Read the comma as separating two indices
Why: Which is how several other languages spell two-dimensional access.
directory[last] raises KeyError, because the only keys are two-element tuples and a bare surname is not one of them. The structure is flat, with compound keys.
Read it as one key that happens to be a pair.
The comma builds a tuple, here as everywhere
Why: So the brackets contain exactly one expression.
If you want lookup by surname alone, that is a different structure
Why: A dictionary of dictionaries, or a second dictionary mapping surnames to lists.
The flat form is usually what you want: it is one lookup rather than two, and it treats the whole name as the identifier — which is what it is.
Faded example
One key, made of two names.
Fill in the blanks
directory[last, first] = number
Why: The comma makes the bracketed expression a tuple, so the two names together form one key. Without it the line would be a syntax error, and using a list — directory[[last, first]] — would raise TypeError: unhashable type, which is exactly why a tuple is the right choice here.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C treats the structure as nested when it is flat. The only keys are two-element tuples, so a bare surname is not a key at all and the lookup raises KeyError. Getting every entry with a surname would need a different structure — a dictionary of dictionaries, or a second mapping from surname to a list of full names.
Real world
Identifying something by several fields at once is completely ordinary.
Discussion prompt
Think of something identified by a combination of fields rather than a single one. What goes wrong if you try to squash them into one string instead?
Hint: What if a field contains your separator?
Answer:
A seat is a row and a number, a cell is a column and a row, an appointment is a date and a time. In each case the identity is the combination and neither half is enough.
Joining them into a string — 'Cleese|John' — works until a value contains the separator, at which point the key splits in the wrong place and the entry is silently lost or misfiled.
A tuple key has no separator to collide with, keeps the fields as their own types, and can be taken apart again by unpacking. That is why it is the right structure rather than merely a convenient one.
Section
Section 4
Concept
In many contexts the different kinds of sequences can be used interchangeably. So how should you choose one over the others? Lists are more common than tuples, mostly because they are mutable — but there are a few cases where you might prefer tuples.
And to start with the obvious: strings are more limited than other sequences because the elements have to be characters, and they are also immutable. If you need the ability to change the characters in a string — as opposed to creating a new string — you might want to use a list of characters instead.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 121-121
Picture it
Each question rules out one option.
Figure (svg): A decision flowchart choosing between a string, a tuple and a list
The book is explicit that lists are more common, mostly because they are mutable. The three tuple cases are exceptions worth recognising rather than a general preference.
Worked example
The third reason is the subtlest and the most valuable.
def process(items):
items.sort() # modifies the caller's list!
return items[0]
>>> data = [3, 1, 2]
>>> process(data)
1
>>> data
[1, 2, 3] # reordered, without being asked| Argument type | What a function can do | Note |
|---|---|---|
| a list argument | the function can modify it | chapter 10's aliasing |
| a tuple argument | sort would raise | the mistake is impossible |
| the difference | a guarantee, not a convention | enforced by the type |
Recall what a list argument allows.
Why: A function receives a reference, so modifying the list is visible to the caller — lesson 10c's central point.
Consider the same call with a tuple.
Why: items.sort() would raise AttributeError immediately, because tuples have no modifying methods.
Read the guarantee.
Why: Using tuples reduces the potential for unexpected behaviour due to aliasing — not by discipline, but because the operation does not exist.
Figure (svg): The state of the program after each line of Worked example the aliasing reason, drawn as a ladder with one rung per traced line
The list version silently reorders the caller's data; the tuple version cannot. The type turns a documentation promise into something the language enforces.
Verify: Ask what the tuple version costs.
Why: The function can no longer sort in place, so it would use sorted(items) and work on a new list — which is what it should have been doing anyway. The restriction pushes you towards the design that does not surprise the caller, which is why it counts as a reason rather than a limitation.
Sorting
Ask what the data are and what will be done with them.
Sort into buckets
For each job, which type fits best?
Worked example
Immutability removes methods; built-in functions replace them.
>>> t = (3, 1, 2)
>>> t.sort()
AttributeError: 'tuple' object has no attribute 'sort'
>>> sorted(t)
[1, 2, 3]
>>> list(reversed(t))
[2, 1, 3]| Expression | What it is | Result |
|---|---|---|
| t.sort() | a modifying method | tuples have none |
| sorted(t) | takes any sequence | returns a new LIST |
| reversed(t) | takes a sequence | returns an iterator |
Note what is missing and why.
Why: Because tuples are immutable, they don't provide methods like sort and reverse, which modify existing lists.
Use sorted instead.
Why: Python provides the built-in function sorted, which takes any sequence and returns a new list with the same elements in sorted order.
Use reversed for the other one.
Why: reversed takes a sequence and returns an iterator that traverses it in reverse order — an iterator, so it needs list() if you want a list.
Figure (svg): Two columns pairing each missing tuple method with the built-in that replaces it
The methods are absent and the built-ins cover both cases. Note that each returns something new rather than modifying, which is the only thing an immutable type could support.
Verify: Check the return types.
Why: sorted returns a list even when given a tuple, and reversed returns an iterator rather than either. Neither gives you a tuple back, so tuple(sorted(t)) is the round trip if you need to stay immutable — which is worth knowing before a type mismatch surprises you two lines later.
Trap
A program uses tuples everywhere, on the grounds that immutability prevents bugs.
Prefer the more restrictive type
Why: It genuinely does prevent a class of aliasing surprises.
Every accumulation now rebuilds the whole sequence — t = t + (x,) inside a loop copies everything each pass — and code that should append is doing quadratic work for no benefit.
Lists are the default; tuples are for the three reasons.
Use a list when the collection grows or changes
Why: Which is most collections, and it is why the book says lists are more common.
Reach for a tuple for a key, a return value, or a guarantee to a caller
Why: The three specific cases, each with a concrete payoff.
The distinction is really about what the data are: a list is a collection of similar things whose length varies, and a tuple is a record of a fixed shape. Choosing by that rather than by safety gets the right answer more often.
Prediction
The argument is a tuple.
t = (3, 1, 2)
print(sorted(t))| Part | What is true | Result |
|---|---|---|
| sorted | takes any sequence | including a tuple |
| what it returns | a new list | always a list |
| the original | unchanged | nothing was modified |
Predict first
What does this print?
Correct: [1, 2, 3] — sorted takes any sequence and returns a new list, whatever it was given.
Why: Tuples have no sort method, because sorting in place would modify them — that is what raises the AttributeError in option C. sorted is the replacement, and it returns a list rather than matching the input type, so tuple(sorted(t)) is the round trip if you need to stay immutable.
Comparison
Fill the blanks. Two properties decide almost everything.
Comparison matrix
| Question | String | List | Tuple |
|---|---|---|---|
| What can it hold? | characters only | any values | any values |
| Mutable? | no | yes | no |
| Usable as a key? | yes | no | yes |
| When is it the default? | for text | for most collections | for records and keys |
The book's advice in one line: lists are more common because they are mutable, and the tuple's three advantages all follow from not being.
Socratic
The third reason is stated briefly and is worth expanding.
Discussion prompt
Chapter 10 said aliasing is error-prone. Why does passing a tuple remove the problem rather than merely making it less likely?
Hint: What has to happen for aliasing to cause a bug?
Answer:
Aliasing itself is harmless — two names for one object is only a problem if one of them changes it.
A tuple cannot be changed, so no holder can affect any other. The aliasing still exists and has no observable consequence, which is exactly the situation strings have always been in.
So it is a removal rather than a reduction: the bug requires a modification, and the modification is impossible. That is stronger than a convention, because it holds however carelessly the receiving function is written.
Section
Section 5
Concept
Lists, dictionaries and tuples are examples of data structures, and in this chapter we are starting to see compound data structures — lists of tuples, or dictionaries that contain tuples as keys and lists as values.
shape error — An error caused because a value has the wrong shape; that is, the wrong type or size.
Compound data structures are useful, but they are prone to shape errors: errors caused when a data structure has the wrong type, size, or structure. For example, if you are expecting a list with one integer and I give you a plain old integer — not in a list — it won't work.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 122-122
Picture it
Each would be described the same way in English and none is interchangeable.
Figure (svg): Four differently shaped values that a casual description would not distinguish
That is what makes shape errors distinctive: the data are right and the packaging is wrong, so error messages talk about types rather than about values.
Worked example
The book provides a module that summarises any structure.
>>> from structshape import structshape
>>> t = [1, 2, 3]
>>> structshape(t)
'list of 3 int'
>>> t2 = [[1, 2], [3, 4], [5, 6]]
>>> structshape(t2)
'list of 3 list of 2 int'| Structure | Its shape | Note |
|---|---|---|
| a flat list | list of 3 int | one level |
| a list of lists | list of 3 list of 2 int | two levels |
| the summary | type and size at each level | not the values |
Ask for the shape rather than the contents.
Why: structshape takes any kind of data structure and returns a string that summarises its shape.
Read it as nested description.
Why: list of 3 list of 2 int says three elements, each itself two integers.
Note what it leaves out.
Why: The values. A shape error is about packaging, so the summary deliberately describes the packaging alone.
Figure (svg): The state of the program after each line of Worked example describing a shape, drawn as a ladder with one rung per traced line
Two one-line descriptions that would have taken several prints to establish by hand. The module is downloadable from the book's own site.
Verify: Compare with printing the structure itself.
Why: Printing [[1,2],[3,4],[5,6]] shows the same information mixed with the values, and for a large structure the shape is invisible in the volume. That is the summary technique from lesson 11c applied specifically to structure — print a description rather than the data.
Prediction
A list of pairs, from zip.
lt = list(zip([1, 2, 3], 'abc'))
# structshape(lt) reports ...| Part | What it is | Note |
|---|---|---|
| zip | pairs the two sequences | three pairs |
| each pair | an int and a str | a tuple of two |
| the whole thing | a list of three tuples | one level of nesting |
Predict first
What shape does structshape report?
Correct: 'list of 3 tuple of (int, str)' — three tuples, each holding an integer and a string.
Why: list(zip(...)) always produces a list of tuples, and here each tuple pairs a number with a character. Option D is what dict(lt) would report, which is the natural next step — converting the pairs into a mapping, exactly as idea 2 did.
Worked example
The descriptions get more useful as the structures get worse.
>>> t3 = [1, 2, 3, 4.0, '5', '6', [7], [8], 9]
>>> structshape(t3)
'list of (3 int, float, 2 str, 2 list of int, int)'
>>> lt = list(zip([1, 2, 3], 'abc'))
>>> structshape(lt)
'list of 3 tuple of (int, str)'
>>> structshape(dict(lt))
'dict of 3 int->str'| Structure | Its shape | Note |
|---|---|---|
| a mixed list | grouped in order by type | the parentheses show a run |
| a list of tuples | list of 3 tuple of (int, str) | the zip result |
| a dictionary | dict of 3 int->str | key type to value type |
Read the grouped form.
Why: If the elements of the list are not the same type, structshape groups them, in order, by type — so 3 int, float, 2 str describes a run of three integers, then a float, then two strings.
Read the list of tuples.
Why: list of 3 tuple of (int, str) is exactly what list(zip(...)) produces, which is a shape worth recognising on sight.
Read the dictionary form.
Why: dict of 3 int->str gives the size and the two types, which is usually the whole question.
Figure (svg): A panel showing a uniform shape description beside a grouped one revealing a stray element
Three descriptions of increasingly compound structures. The mixed-type one is the most telling, because an unexpectedly mixed list is a classic shape error.
Verify: Use it on a structure you believe is uniform.
Why: If structshape reports a grouped description like list of (5 int, str, 3 int) for something meant to be all integers, one element is the wrong type — and the position in the grouping tells you roughly where. That is a diagnosis that would take a loop and several prints to reach by hand.
Trap
A function raises a TypeError and the developer starts checking whether the numbers are right.
Assume wrong output means wrong data
Why: Which is the usual case for arithmetic bugs.
The values may be perfectly correct and wrapped in one list too many. Checking them confirms they are right, which is true and unhelpful, and the actual problem is never examined.
Check the shape before the contents.
Print the type, or the structure's shape
Why: A common cause of runtime errors is a value that is not the right type, and printing the type is often enough.
Compare against what the function expects
Why: list with one integer against a plain old integer is a difference in packaging, not in value.
The tell is a TypeError or an AttributeError rather than a wrong number. Those messages are almost always about shape, and they point at the packaging rather than the contents.
Error analysis
Mark each and say what the shape is versus what was wanted.
Annotate
Shape errors produce TypeErrors and ValueErrors rather than wrong numbers, which is a useful signal about where to look.
Faded example
The shape is already right; it needs one conversion.
Fill in the blanks
lt = list(zip([1, 2, 3], 'abc'))
d = dict(lt)
# structshape(d) -> 'dict of 3 int->str'
Why: dict takes a list of tuples and makes each pair an item, so 'list of 3 tuple of (int, str)' becomes 'dict of 3 int->str'. The shapes line up because zip already produced exactly the two-element tuples dict requires — which is why dict(zip(...)) works as a single idiom.
Explain it
The classic symptom of a shape error.
Discussion prompt
A classmate has checked every value in their structure and they are all correct, and the program still raises a TypeError. What should they look at instead?
Hint: Not the values.
Answer:
The packaging. A list containing one integer and a plain integer hold the same value and are not interchangeable, and no amount of checking the number will reveal that.
Have them print the type at each level, or use structshape to get a one-line description of the whole thing, and compare it with what the receiving code expects.
The signal is in the error kind: TypeError and AttributeError are almost always about shape, while a wrong answer with no exception is usually about values. Reading which kind of failure you have tells you which of the two to investigate.
Comparison
Fill the blanks. Each is good at what the other cannot do.
Comparison matrix
| Question | Dictionary | List of tuples |
|---|---|---|
| How do you find something? | by key, in about constant time | by searching, in proportion to the length |
| Is there an order? | no | yes, and it can be sorted |
| Duplicate keys? | impossible — later wins | allowed, and kept |
| How do you convert? | dict(pairs) | list(d.items()) |
The bottom row is why the conversion matters: you move to whichever structure is good at the operation you need next.
Pattern
Five steps. The third is where the shape errors come from.
Step 3's failure is a KeyError for something you can see in the dictionary, which is the standard symptom of a compound key assembled in the wrong order.
Python documentation — Data Structures Data Structures
Check
Two names, and a dictionary.
d = {'a': 1, 'b': 2}
total = 0
for k, v in d.items():
total += v
print(total)| Pass | The names | Running total |
|---|---|---|
| pass 1 | k='a', v=1 | total 1 |
| pass 2 | k='b', v=2 | total 3 |
| the result | the sum of the values | 3 |
Check your understanding
What does this print?
Answer: A
Why: items yields a key-value pair per pass, which the two names unpack, so v holds the values 1 and 2 and the total is 3. Looping over d directly instead would bind k to a key and there would be nothing to unpack, which is where the ValueError in option C would come from.
Check
Two values inside the brackets.
Check your understanding
What does directory['Cleese', 'John'] = number create?
Answer: A
Why: The expression in brackets is a tuple — the comma builds it, exactly as everywhere else — so a single key goes in. Printing the dictionary shows the key with its parentheses, confirming that one object was used rather than two separate indices.
Check
The book gives three reasons to prefer a tuple.
Check your understanding
Which of these is NOT one of the book's reasons for preferring a tuple over a list?
Answer: A
Why: The three reasons the book gives are syntax, keys, and aliasing. Memory is not among them — and the book is explicit that lists are more common than tuples, mostly because they are mutable, so the tuple cases are specific exceptions rather than a general preference.
Real world
The difference between a record and a collection is not a programming idea.
Discussion prompt
Think of a form with a fixed set of fields and a shopping list. What is different about how you would change each, and how does that map onto tuples and lists?
Hint: One has a shape and the other has a length.
Answer:
You add to a shopping list and its length changes; you fill in a form and its shape does not. Adding a field to a form is redesigning it, not using it.
That is the tuple-and-list distinction: a list is a collection of similar things whose length varies, and a tuple is a record of a fixed number of positions where each means something particular.
And it explains the shape errors. Handing someone a form where a list was expected fails not because the contents are wrong but because the packaging is — which is exactly what a TypeError about a structure is telling you.
Commit first
Answer, then rate your confidence.
Predict first
In directory[last, first] = number, how many keys does the assignment create?
Correct: One — the brackets contain a tuple, because the comma builds one.
Why: The book puts it in exactly those words: the expression in brackets is a tuple. This is not two-dimensional indexing, however much it resembles it — Python sees the comma, builds a two-element tuple, and uses that single object as the key. You can confirm it by printing the dictionary, which shows the key with its parentheses, and by trying the same thing with a list: directory[[last, first]] raises TypeError: unhashable type, which is the whole reason a tuple is used here. The types of the two names are irrelevant, so long as each is itself hashable — a tuple containing a list would fail for the same reason a list does.
Explain it
Three types, and when each one is right.
Discussion prompt
A classmate asks how to decide between a string, a list and a tuple. Give them the two questions that settle it, and the one default.
Hint: The book says one of the three is more common.
Answer:
First question: are all the elements characters? If so a string will do, unless you need to change them, in which case a list of characters is the way.
Second: will it be a dictionary key, or must it not change? Then it has to be immutable, so a tuple.
Otherwise a list, which is the default — the book says lists are more common than tuples, mostly because they are mutable. The three tuple cases are worth recognising, and everything else is a list.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The conversions are worth having automatic, because so many operations are natural on pairs and impossible on a dictionary. The dict-with-zip idiom is compact and has two silent failure modes — truncation and duplicate keys — both worth knowing. The tuple-key syntax is where the grid misreading catches people, and stating the expression in brackets is a tuple out loud usually fixes it for good. And the choosing section is the closest the book comes to design advice, which makes it worth more than its length suggests.
Connect it up
One page, from memory.
Draw it
Draw a dictionary on the left and a list of tuples on the right, with arrows between them labelled items() and dict(). Beside each, write one thing it can do that the other cannot. Underneath, draw the telephone directory as the book's figure 12.2 — tuple keys in Python syntax, mapping to numbers — and write beside it the line that stores an entry, with the comma circled. Finally list the three reasons to prefer a tuple over a list.
Recap
Three pages, and chapter 12 is finished: the three collection types working together.
| If you remember one thing | It is this |
|---|---|
| From items | A dictionary's pairs are tuples, so everything about tuples applies to them. |
| From dict(zip(...)) | The first sequence supplies the keys, and duplicates silently collapse. |
| From tuple keys | The expression in brackets is a tuple, not two indices. |
| From choosing | Lists are more common; the tuple's three advantages all follow from immutability. |
| From shape errors | The values can all be right and the packaging still wrong. |
The next chapter is a case study: reading a text file, counting words with the dictionaries of this chapter and the last, and using random selection to generate text — the first program in the book built from all three collection types at once.
Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-122 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.