12c Dictionaries and Tuples, and Sequences of Sequences

This lesson converts between dictionaries and lists of tuples, uses tuples as dictionary keys, gives the criteria for choosing between strings, lists and tuples, and introduces shape errors in compound data structures.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 12c Dictionaries and Tuples, and Sequences of Sequences

Title

Python · Chapter 12 — Tuples

§12.6-12.8, pp. 120-122

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-122 — the pages these objectives are drawn from

3. Before we start: a directory keyed by two things

Warm-up

You want to look someone up by surname and forename together.

Discussion prompt

A telephone directory maps a person to a number, and a person is a surname and a forename. What would you use as the dictionary key, and what would have gone wrong if you had tried this two chapters ago?

Hint: The key has to be one thing, and it cannot be a list.

Answer:

The key has to be a single value holding both names — a sequence of two — and it cannot be a list, because keys must be hashable.

Two chapters ago there was no such type. You would have had to join the names into a string with a separator and hope no name contained it.

A tuple solves it exactly: two values, one object, immutable and therefore hashable. This is the payoff the last two chapters have been building towards.

4. The one idea behind this lesson: the three types compose

Concept

It is common to use tuples as keys in dictionaries, primarily because you can't use lists. A telephone directory might map from last-name, first-name pairs to telephone numbers.

directory[last, first] = number

for last, first in directory:
    print(first, last, directory[last, first])
LineWhat it meansNote
directory[last, first]the brackets hold a TUPLEnot two indices
for last, first in directorythe keys are tuplesunpacked per pass
the lookup insidethe same tuple keyrebuilt from the two names

The expression in brackets is a tuple. This is not two-dimensional indexing, however much it looks like it — Python sees the comma, builds a tuple, and uses that one object as the key.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120

5. items, and the pairs of a dictionary

Section

Section 1

6. A dictionary as a sequence of pairs

Concept

Dictionaries have a method called items that returns a sequence of tuples, where each tuple is a key-value pair.

>>> d = {'a': 0, 'b': 1, 'c': 2}
>>> t = d.items()
>>> t
dict_items([('c', 2), ('a', 0), ('b', 1)])
>>> for key, value in d.items():
...     print(key, value)
c 2
a 0
b 1
PartWhat it givesNote
d.items()a dict_items objectan iterator
each elementa key-value pairas a tuple
the for loopunpacks each pairtwo names per pass

The result is a dict_items object, which is an iterator that iterates the key-value pairs — the same kind of thing as a zip object. And as you should expect from a dictionary, the items are in no particular order.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120

7. Picture it: a dictionary viewed as pairs

Picture it

The same three items, seen as a sequence of tuples.

Figure (svg): A dictionary shown alongside the sequence of key-value tuples that items produces

Each item becomes a two-element tuple. The order is still unpredictable.

Once a dictionary is a sequence of tuples, everything the last two lessons taught about tuples applies to it — including sorting, unpacking and comparison.

8. Worked example: three ways to traverse a dictionary

Worked example

The same output, with different amounts of work.

# keys, then look each value up
for c in d:
    print(c, d[c])

# the pairs, kept whole
for pair in d.items():
    print(pair[0], pair[1])

# the pairs, unpacked
for key, value in d.items():
    print(key, value)
FormWhat each pass givesVerdict
the firsta lookup per passworks, does extra work
the seconda tuple per pass, indexedworks, reads badly
the thirdtwo names per passsays what it means

Start with what lesson 11b gave you.

Why: Looping over the dictionary yields keys, and each value needs a lookup — correct, and one operation more than necessary.

Ask for the pairs instead.

Why: items yields tuples, so both halves arrive together and no lookup is needed.

Unpack them in the for statement.

Why: Tuple assignment names the two halves directly, which removes the indices as well.

Figure (svg): The state of the program after each line of Worked example three ways to traverse a dictionary, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

All three print the same thing. The third does the least work and reads the closest to its intention, which is why it is the standard form.

Verify: Check that the order is the same in all three.

Why: It is unpredictable in all three, and identical between them for a given dictionary — because all three walk the same items. items does not impose an order any more than looping over the keys does, which is worth confirming before relying on either.

9. Predict: what does each pass give?

Prediction

Two names on the left of in.

d = {'a': 0}
for key, value in d.items():
    print(key, value)
PartWhat happensOutput
d.items()yields ('a', 0)one pair
key, valueunpacked by position'a' and 0
printboth, separateda 0

Predict first

What does this print?

  • a 0
  • ('a', 0)
  • a
  • 0 a

Correct: a 0 — items yields a key-value tuple, which the two names unpack in order.

Why: Each element of items is a tuple whose first component is the key, so key binds to 'a' and value to 0. Writing for pair in d.items() instead would bind one name to the whole tuple and print option B. Looping over the dictionary directly, without items, would give option C — just the keys.

10. Worked example: sorting a dictionary's items

Worked example

Now that the items are tuples, the comparison rule applies.

>>> d = {'b': 1, 'a': 3, 'c': 2}
>>> sorted(d.items())
[('a', 3), ('b', 1), ('c', 2)]
>>> h = {'a': 3, 'b': 1}
>>> sorted([(v, k) for k, v in h.items()])
[(1, 'b'), (3, 'a')]
ExpressionWhat it sorts byNote
sorted(d.items())sorts the tuplesby key, since it comes first
swapping the pairsvalue firstsorts by value instead
the ruleposition 0 dominateslesson 12a's comparison

Sort the items directly.

Why: Each item is a tuple, and tuples compare element by element — so the key at position 0 decides the order.

Swap to sort by value.

Why: Building pairs the other way round puts the value at position 0, which makes it the primary ordering.

Note what makes this work.

Why: Nothing about dictionaries: it is the sequence comparison rule from lesson 12a applied to the pairs that items produces.

Figure (svg): A pipeline showing a dictionary becoming pairs, being swapped, and then sorted by value

Turning a dictionary into something ordered means turning it into a list of tuples.

Sorted by key in the first case and by value in the second. The position of the field inside the tuple is what chooses the ordering.

Verify: Check what sorted returns.

Why: A list, not a dictionary — sorted always returns a list whatever it was given. There is no such thing as a sorted dictionary, for the reason lesson 11a gave, so what you have is a sequence of pairs in an order you chose.

11. Trap: expecting items to be a list

Trap

The trap

A program stores pairs = d.items() and later indexes it with pairs[0].

Treat it as the list it prints like

Why: dict_items([('c', 2), ...]) looks very much like a list of tuples.

It is an iterator, so indexing raises a TypeError — the same restriction as a zip object, for the same reason.

The fix

Use it in a for loop, or convert it.

for key, value in d.items()

Why: The intended use, and the one that needs no conversion.

list(d.items()) if you need list operations

Why: Which gives a real list of tuples you can index and sort.

The printed form is genuinely misleading here — dict_items wraps what looks like a list — so the type name in the output is the thing to read, not the brackets inside it.

12. Complete it: traverse the pairs

Faded example

Both halves at once, with no lookup.

Fill in the blanks

for key, value in d.items():
print(key, value)

Why: items returns a sequence of tuples, each a key-value pair, which the tuple assignment unpacks into two names. Looping over d alone would give only the keys and need d[key] inside the body; d.values() would give only the values and no way back to the keys.

13. Discriminate: which half do you get?

Discrimination

Three methods, three views of the same dictionary.

Sort into buckets

For each expression, what does one pass of a loop over it give?

one value per pass
for x in d; for x in d.values(); for x in d.items(); for x in sorted(d)
two names per pass
for k, v in d.items(); for k, v in sorted(d.items())
one
Each binds a single name per pass — a key, a value, or in one case a whole key-value tuple that has not been taken apart. Sorting changes the order and not the shape.
two
Each pairs items() with tuple assignment, so the two halves of every pair are bound to separate names. The sorted version does the same after ordering the pairs.

14. Explain it yourself: why does items make sorting possible?

Explain it to yourself

A dictionary cannot be sorted. Its items can.

Discussion prompt

Explain why sorted(d.items()) is meaningful when sorting a dictionary is not.

Hint: What does sorted return?

Answer:

Sorting means arranging things in an order, and a dictionary has no arrangement — its storage location comes from hashing the key, so there is nowhere to put an order.

items produces tuples, and a list of tuples is an ordinary sequence with positions. Sorting that is perfectly meaningful.

So sorted(d.items()) does not sort the dictionary; it produces a new list, ordered, from the dictionary's contents. The dictionary is untouched, which is the returning-not-modifying distinction again.

15. Building a dictionary from pairs

Section

Section 2

16. The conversion runs both ways

Concept

Going in the other direction, you can use a list of tuples to initialize a new dictionary. And combining dict with zip yields a concise way to create one.

>>> t = [('a', 0), ('c', 2), ('b', 1)]
>>> d = dict(t)
>>> d
{'a': 0, 'c': 2, 'b': 1}
>>> d = dict(zip('abc', range(3)))
>>> d
{'a': 0, 'c': 2, 'b': 1}
ExpressionWhat it doesNote
dict(t)a list of pairsbecomes a dictionary
zip('abc', range(3))pairs the two sequences('a',0), ('b',1), ('c',2)
dict(zip(...))the pairs become itemsone line

The dictionary method update also takes a list of tuples and adds them, as key-value pairs, to an existing dictionary — so the same shape works for building one from scratch or extending one you have.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120

17. Picture it: the round trip

Picture it

Dictionary to pairs and back, with a different tool in each direction.

Figure (svg): A pipeline showing a dictionary converted to pairs and pairs converted back to a dictionary

items goes one way and dict comes back. Nothing is lost except the order, which there never was.

The round trip is worth knowing because so many operations — sorting, filtering, zipping — are natural on a list of pairs and impossible on a dictionary.

18. Worked example: dict and zip together

Worked example

Two sequences become a mapping in one line.

>>> letters = 'abc'
>>> numbers = range(3)
>>> d = dict(zip(letters, numbers))
>>> d
{'a': 0, 'b': 1, 'c': 2}
>>> d['b']
1
StepWhat it producesNote
zippairs them by position('a',0), ('b',1), ('c',2)
dicteach pair becomes an itemkey from position 0
the lookupordinary1

Pair the sequences.

Why: zip interleaves them into two-element tuples, one per position.

Hand the pairs to dict.

Why: Each tuple becomes one item, with position 0 as the key and position 1 as the value.

Note the direction is fixed.

Why: The first sequence supplies the keys, so swapping the arguments inverts the mapping.

Figure (svg): Two sequences aligned and paired, with each pair becoming a dictionary item

The first sequence supplies the keys. Swapping the arguments inverts the mapping.

A three-item dictionary, built from two independent sequences without a loop. This is the concise idiom the book names.

Verify: Try it with sequences of different lengths.

Why: dict(zip('abcd', range(2))) gives a two-item dictionary, because zip truncates to the shorter — silently, as lesson 12b warned. So a mismatch produces a smaller dictionary rather than an error, which is worth checking for when the two sequences are meant to correspond exactly.

19. Predict: how many items?

Prediction

One key appears twice in the list.

t = [('a', 1), ('b', 2), ('a', 3)]
d = dict(t)
print(len(d), d['a'])
PairWhat happensSize
three pairsconverted in order'a' set to 1
('b', 2)a new keytwo items
('a', 3)replaces the earlier 'a'still two items

Predict first

What does this print?

  • 2 3
  • 3 1
  • 2 1
  • 3 3

Correct: 2 3 — two items, and the later pair for 'a' replaced the earlier one.

Why: Each key is associated with a single value, so the third pair overwrites the first rather than adding an item. Nothing warns you that a pair was absorbed; comparing len(t) with len(d) is the check that finds it, and it is worth running whenever duplicates would be a bug rather than an expectation.

20. Worked example: what happens to duplicate keys

Worked example

A list of pairs can contain the same key twice. A dictionary cannot.

>>> t = [('a', 1), ('b', 2), ('a', 3)]
>>> len(t)
3
>>> d = dict(t)
>>> d
{'a': 3, 'b': 2}
>>> len(d)
2
StageWhat happensSize
the listthree pairstwo of them share a key
dict(t)later wins'a' maps to 3
the length2, not 3one pair was absorbed

Count the pairs.

Why: The list has three, which is a perfectly ordinary list.

Convert and count again.

Why: Each key is associated with a single value, so the second 'a' replaces the first rather than joining it.

Notice nothing is reported.

Why: The conversion succeeds quietly and one pair is gone, which is silent data loss if the duplicate was not intended.

Figure (svg): The state of the program after each line of Worked example what happens to duplicate keys, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A two-item dictionary. Later pairs win, and nothing tells you a pair was absorbed except the length.

Verify: Compare the lengths as a check.

Why: len(dict(t)) < len(t) means there were duplicate keys, which is a one-line consistency check of exactly the kind lesson 11c described. Where duplicates are expected rather than a bug, invert_dict's list-valued approach keeps them all instead.

21. Trap: zipping the sequences the wrong way round

Trap

The trap

A program builds dict(zip(values, keys)) and then cannot find anything.

Name the arguments after what they hold

Why: Both sequences are being passed, so the order can look arbitrary.

The first sequence supplies the keys, so the mapping is inverted. Lookups fail with KeyError on things that are plainly in the dictionary — as values.

The fix

Keys first, always.

dict(zip(keys, values))

Why: Position 0 of each pair becomes the key, and zip preserves the argument order.

Check with one lookup after building

Why: One d[known_key] confirms the direction immediately.

The symptom is distinctive: a KeyError for something you can see in the printed dictionary, on the wrong side of the colon. That is always an inverted mapping.

22. Complete it: build a dictionary from two sequences

Faded example

One line, two built-in functions.

Fill in the blanks

d = dict(zip('abc', range(3)))
print(d['b']) # 1

Why: zip pairs the two sequences into tuples and dict turns each pair into an item, with the first sequence supplying the keys. Swapping the arguments would invert the mapping — a mistake whose symptom is a KeyError for something visibly present in the dictionary, on the value side.

23. Compare: the two directions

Comparison

Fill the blanks. Each direction has its own tool.

Comparison matrix

QuestionDictionary to pairsPairs to dictionary
Which tool?d.items()dict(pairs)
What comes out?an iterator of tuplesa dictionary
Can it lose information?no — every item becomes a pairyes — duplicate keys silently collapse
What about ordering?unpredictable, as alwaysthe input order is not preserved

The third row is the asymmetry: pairs allow duplicates and a dictionary does not, so only one direction can lose anything.

24. Explain it: when would I convert to pairs and back?

Explain it

The round trip looks pointless until you see what happens in the middle.

Discussion prompt

A classmate asks why anyone would turn a dictionary into a list of tuples. Give them two things that are easy on pairs and impossible on a dictionary.

Hint: Both involve order.

Answer:

Sorting. A dictionary has no order and cannot be given one, but a list of tuples sorts naturally — by key, or by value if you swap the pairs round first.

Filtering into a new dictionary. Building the pairs you want and calling dict on them is often clearer than looping and assigning item by item.

The general shape is that dictionaries are for lookup and sequences are for order and position. Converting between them is how you get a structure that is good at the operation you need next.

25. Tuples as dictionary keys

Section

Section 3

26. The expression in brackets is a tuple

Concept

It is common to use tuples as keys in dictionaries, primarily because you can't use lists. A telephone directory might map from last-name, first-name pairs to telephone numbers.

directory[last, first] = number

for last, first in directory:
    print(first, last, directory[last, first])
LineWhat happensNote
directory[last, first]the comma builds a tupleone key
for last, first in directorykeys are tuplesunpacked per pass
the lookup in the bodythe same two namesrebuilds the key

The expression in brackets is a tuple — the comma is doing the same job it did in lesson 12a. This loop traverses the keys in the directory, which are tuples, assigns the elements of each to last and first, and then prints the name and the corresponding number.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-120

27. Picture it: figures 12.1 and 12.2

Picture it

There are two ways to represent a tuple in a state diagram, and the choice is about detail.

Figure (svg): A state diagram of a directory whose keys are name tuples shown in Python syntax

The book's figure 12.2: tuples shown using Python syntax as a graphical shorthand.

The detailed version — figure 12.1 — draws each tuple as a box with indices and elements, like a list. In a larger diagram you leave the details out, which is what makes the directory legible.

28. Worked example: it is not two-dimensional indexing

Worked example

The line looks like a grid access and is not.

>>> directory = {}
>>> directory['Cleese', 'John'] = '08700 100 222'
>>> directory
{('Cleese', 'John'): '08700 100 222'}
>>> directory[('Cleese', 'John')]
'08700 100 222'
ExpressionWhat it meansNote
the two names in bracketsa comma between thema tuple
the printed dictionaryshows the tuple keywith parentheses
the explicit formidenticalparentheses are optional

Read the comma, not the brackets.

Why: The expression in brackets is a tuple — exactly as it is anywhere else, because the comma is what constructs.

Confirm from the printed form.

Why: The dictionary prints its key with parentheses, showing that one tuple went in rather than two indices.

Note both spellings work.

Why: directory[last, first] and directory[(last, first)] are the same expression; the parentheses only group.

Figure (svg): The state of the program after each line of Worked example it is not two-dimensional indexing, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

One key, which happens to be a pair. There is no two-dimensional indexing here — the brackets take one expression, and the comma made that expression a tuple.

Verify: Try the same thing with a list.

Why: directory[[last, first]] raises TypeError: unhashable type: 'list'. That failure confirms the key is a single object being hashed, and it is the whole reason a tuple is used here — which is the point chapter 11 pointed forward to.

29. Predict: what is the key?

Prediction

Two values inside the brackets.

d = {}
d['Cleese', 'John'] = 1
print(list(d))
PartWhat happensResult
the commabuilds a tupleone key
the key('Cleese', 'John')a two-element tuple
list(d)the keysone of them

Predict first

What does this print?

  • [('Cleese', 'John')]
  • ['Cleese', 'John']
  • ['Cleese']
  • A TypeError about too many indices

Correct: [('Cleese', 'John')] — one key, which is a tuple of two strings.

Why: The expression in brackets is a tuple, so the dictionary has exactly one item whose key is that pair. Option B is what two separate keys would look like, and option C is what a grid-style reading would predict. Nothing raises: brackets take one expression, and the comma made it a single tuple.

30. Worked example: traversing a tuple-keyed dictionary

Worked example

The keys are pairs, so unpack them as you go.

for last, first in directory:
    print(first, last, directory[last, first])

# Or, better, with items:
for (last, first), number in directory.items():
    print(first, last, number)
FormWhat happensNote
for last, first in directorykeys are tuplesunpacked into two names
directory[last, first]rebuilds the keyone lookup per pass
the items versionnested unpackingno lookup at all

Unpack the key in the for statement.

Why: Looping over the dictionary gives keys, and here each key is a two-element tuple, so tuple assignment splits it.

Rebuild the key to look up the value.

Why: The book's version writes directory[last, first], reconstructing the tuple from the two names.

Or take the pairs instead.

Why: items yields (key, value), and the key is itself a tuple — so (last, first), number unpacks both levels at once and needs no lookup.

Figure (svg): A flowchart showing a key-value pair unpacked at two levels

Two levels of unpacking, written as one pattern in the for statement.

Both print each name with its number. The nested unpacking is the tidier form once you can read it, and it does one less operation per pass.

Verify: Check the parentheses in the nested version.

Why: They are required there: (last, first), number tells Python the first element is itself a pair, whereas last, first, number would ask for three elements and raise ValueError. This is the one place in the chapter where parentheses around a tuple are not optional.

31. Trap: reading directory[last, first] as a grid

Trap

The trap

A student reads it as row-and-column indexing and expects directory[last] to give a sub-dictionary.

Read the comma as separating two indices

Why: Which is how several other languages spell two-dimensional access.

directory[last] raises KeyError, because the only keys are two-element tuples and a bare surname is not one of them. The structure is flat, with compound keys.

The fix

Read it as one key that happens to be a pair.

The comma builds a tuple, here as everywhere

Why: So the brackets contain exactly one expression.

If you want lookup by surname alone, that is a different structure

Why: A dictionary of dictionaries, or a second dictionary mapping surnames to lists.

The flat form is usually what you want: it is one lookup rather than two, and it treats the whole name as the identifier — which is what it is.

32. Complete it: store a number under a name pair

Faded example

One key, made of two names.

Fill in the blanks

directory[last, first] = number

Why: The comma makes the bracketed expression a tuple, so the two names together form one key. Without it the line would be a syntax error, and using a list — directory[[last, first]] — would raise TypeError: unhashable type, which is exactly why a tuple is the right choice here.

33. Two truths and a lie: tuple keys

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. The expression in brackets is a tuple, not two separate indices
  • B. Tuples are used as keys primarily because you can't use lists
  • C. directory[last] will give you every entry with that surname

Survives elimination: C

Why: C treats the structure as nested when it is flat. The only keys are two-element tuples, so a bare surname is not a key at all and the lookup raises KeyError. Getting every entry with a surname would need a different structure — a dictionary of dictionaries, or a second mapping from surname to a list of full names.

34. Where compound keys are already used

Real world

Identifying something by several fields at once is completely ordinary.

Discussion prompt

Think of something identified by a combination of fields rather than a single one. What goes wrong if you try to squash them into one string instead?

Hint: What if a field contains your separator?

Answer:

A seat is a row and a number, a cell is a column and a row, an appointment is a date and a time. In each case the identity is the combination and neither half is enough.

Joining them into a string — 'Cleese|John' — works until a value contains the separator, at which point the key splits in the wrong place and the entry is silently lost or misfiled.

A tuple key has no separator to collide with, keeps the fields as their own types, and can be taken apart again by unpacking. That is why it is the right structure rather than merely a convenient one.

35. Choosing between the sequence types

Section

Section 4

36. Three reasons to prefer a tuple

Concept

In many contexts the different kinds of sequences can be used interchangeably. So how should you choose one over the others? Lists are more common than tuples, mostly because they are mutable — but there are a few cases where you might prefer tuples.

And to start with the obvious: strings are more limited than other sequences because the elements have to be characters, and they are also immutable. If you need the ability to change the characters in a string — as opposed to creating a new string — you might want to use a list of characters instead.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 121-121

37. Picture it: the decision, in three questions

Picture it

Each question rules out one option.

Figure (svg): A decision flowchart choosing between a string, a tuple and a list

Lists are the default; the other two are chosen for specific reasons.

The book is explicit that lists are more common, mostly because they are mutable. The three tuple cases are exceptions worth recognising rather than a general preference.

38. Worked example: the aliasing reason

Worked example

The third reason is the subtlest and the most valuable.

def process(items):
    items.sort()          # modifies the caller's list!
    return items[0]

>>> data = [3, 1, 2]
>>> process(data)
1
>>> data
[1, 2, 3]              # reordered, without being asked
Argument typeWhat a function can doNote
a list argumentthe function can modify itchapter 10's aliasing
a tuple argumentsort would raisethe mistake is impossible
the differencea guarantee, not a conventionenforced by the type

Recall what a list argument allows.

Why: A function receives a reference, so modifying the list is visible to the caller — lesson 10c's central point.

Consider the same call with a tuple.

Why: items.sort() would raise AttributeError immediately, because tuples have no modifying methods.

Read the guarantee.

Why: Using tuples reduces the potential for unexpected behaviour due to aliasing — not by discipline, but because the operation does not exist.

Figure (svg): The state of the program after each line of Worked example the aliasing reason, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The list version silently reorders the caller's data; the tuple version cannot. The type turns a documentation promise into something the language enforces.

Verify: Ask what the tuple version costs.

Why: The function can no longer sort in place, so it would use sorted(items) and work on a new list — which is what it should have been doing anyway. The restriction pushes you towards the design that does not surprise the caller, which is why it counts as a reason rather than a limitation.

39. Sort: which sequence type?

Sorting

Ask what the data are and what will be done with them.

Sort into buckets

For each job, which type fits best?

a list
a growing collection of scores; the lines read from a file, to be filtered; a list of items being sorted in place
a tuple
a coordinate pair used as a dictionary key; the two values a function returns; a sequence passed to a function that must not change it
list
Each grows, shrinks or is reordered, which needs mutability. Lists are more common precisely because most collections do at least one of those.
tuple
Each is one of the book's three cases: a dictionary key, a return statement, and an argument where immutability removes the potential for aliasing surprises.

40. Worked example: the missing methods, and their replacements

Worked example

Immutability removes methods; built-in functions replace them.

>>> t = (3, 1, 2)
>>> t.sort()
AttributeError: 'tuple' object has no attribute 'sort'
>>> sorted(t)
[1, 2, 3]
>>> list(reversed(t))
[2, 1, 3]
ExpressionWhat it isResult
t.sort()a modifying methodtuples have none
sorted(t)takes any sequencereturns a new LIST
reversed(t)takes a sequencereturns an iterator

Note what is missing and why.

Why: Because tuples are immutable, they don't provide methods like sort and reverse, which modify existing lists.

Use sorted instead.

Why: Python provides the built-in function sorted, which takes any sequence and returns a new list with the same elements in sorted order.

Use reversed for the other one.

Why: reversed takes a sequence and returns an iterator that traverses it in reverse order — an iterator, so it needs list() if you want a list.

Figure (svg): Two columns pairing each missing tuple method with the built-in that replaces it

Every replacement returns something rather than modifying — which is the only option for an immutable type.

The methods are absent and the built-ins cover both cases. Note that each returns something new rather than modifying, which is the only thing an immutable type could support.

Verify: Check the return types.

Why: sorted returns a list even when given a tuple, and reversed returns an iterator rather than either. Neither gives you a tuple back, so tuple(sorted(t)) is the round trip if you need to stay immutable — which is worth knowing before a type mismatch surprises you two lines later.

41. Trap: using a tuple because it seems safer

Trap

The trap

A program uses tuples everywhere, on the grounds that immutability prevents bugs.

Prefer the more restrictive type

Why: It genuinely does prevent a class of aliasing surprises.

Every accumulation now rebuilds the whole sequence — t = t + (x,) inside a loop copies everything each pass — and code that should append is doing quadratic work for no benefit.

The fix

Lists are the default; tuples are for the three reasons.

Use a list when the collection grows or changes

Why: Which is most collections, and it is why the book says lists are more common.

Reach for a tuple for a key, a return value, or a guarantee to a caller

Why: The three specific cases, each with a concrete payoff.

The distinction is really about what the data are: a list is a collection of similar things whose length varies, and a tuple is a record of a fixed shape. Choosing by that rather than by safety gets the right answer more often.

42. Predict: what does sorted return?

Prediction

The argument is a tuple.

t = (3, 1, 2)
print(sorted(t))
PartWhat is trueResult
sortedtakes any sequenceincluding a tuple
what it returnsa new listalways a list
the originalunchangednothing was modified

Predict first

What does this print?

  • [1, 2, 3]
  • (1, 2, 3)
  • An AttributeError, because tuples cannot be sorted
  • None

Correct: [1, 2, 3] — sorted takes any sequence and returns a new list, whatever it was given.

Why: Tuples have no sort method, because sorting in place would modify them — that is what raises the AttributeError in option C. sorted is the replacement, and it returns a list rather than matching the input type, so tuple(sorted(t)) is the round trip if you need to stay immutable.

43. Compare: the three sequence types

Comparison

Fill the blanks. Two properties decide almost everything.

Comparison matrix

QuestionStringListTuple
What can it hold?characters onlyany valuesany values
Mutable?noyesno
Usable as a key?yesnoyes
When is it the default?for textfor most collectionsfor records and keys

The book's advice in one line: lists are more common because they are mutable, and the tuple's three advantages all follow from not being.

44. Think it through: why does immutability reduce aliasing problems?

Socratic

The third reason is stated briefly and is worth expanding.

Discussion prompt

Chapter 10 said aliasing is error-prone. Why does passing a tuple remove the problem rather than merely making it less likely?

Hint: What has to happen for aliasing to cause a bug?

Answer:

Aliasing itself is harmless — two names for one object is only a problem if one of them changes it.

A tuple cannot be changed, so no holder can affect any other. The aliasing still exists and has no observable consequence, which is exactly the situation strings have always been in.

So it is a removal rather than a reduction: the bug requires a modification, and the modification is impossible. That is stronger than a convention, because it holds however carelessly the receiving function is written.

45. Shape errors

Section

Section 5

46. Compound structures go wrong in a new way

Concept

Lists, dictionaries and tuples are examples of data structures, and in this chapter we are starting to see compound data structures — lists of tuples, or dictionaries that contain tuples as keys and lists as values.

shape error — An error caused because a value has the wrong shape; that is, the wrong type or size.

Compound data structures are useful, but they are prone to shape errors: errors caused when a data structure has the wrong type, size, or structure. For example, if you are expecting a list with one integer and I give you a plain old integer — not in a list — it won't work.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 122-122

47. Picture it: four values that all look like *some numbers*

Picture it

Each would be described the same way in English and none is interchangeable.

Figure (svg): Four differently shaped values that a casual description would not distinguish

All four hold the number three, and no two of them work in the same place.

That is what makes shape errors distinctive: the data are right and the packaging is wrong, so error messages talk about types rather than about values.

48. Worked example: describing a shape

Worked example

The book provides a module that summarises any structure.

>>> from structshape import structshape
>>> t = [1, 2, 3]
>>> structshape(t)
'list of 3 int'
>>> t2 = [[1, 2], [3, 4], [5, 6]]
>>> structshape(t2)
'list of 3 list of 2 int'
StructureIts shapeNote
a flat listlist of 3 intone level
a list of listslist of 3 list of 2 inttwo levels
the summarytype and size at each levelnot the values

Ask for the shape rather than the contents.

Why: structshape takes any kind of data structure and returns a string that summarises its shape.

Read it as nested description.

Why: list of 3 list of 2 int says three elements, each itself two integers.

Note what it leaves out.

Why: The values. A shape error is about packaging, so the summary deliberately describes the packaging alone.

Figure (svg): The state of the program after each line of Worked example describing a shape, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Two one-line descriptions that would have taken several prints to establish by hand. The module is downloadable from the book's own site.

Verify: Compare with printing the structure itself.

Why: Printing [[1,2],[3,4],[5,6]] shows the same information mixed with the values, and for a large structure the shape is invisible in the volume. That is the summary technique from lesson 11c applied specifically to structure — print a description rather than the data.

49. Predict: what shape is this?

Prediction

A list of pairs, from zip.

lt = list(zip([1, 2, 3], 'abc'))
# structshape(lt) reports ...
PartWhat it isNote
zippairs the two sequencesthree pairs
each pairan int and a stra tuple of two
the whole thinga list of three tuplesone level of nesting

Predict first

What shape does structshape report?

  • 'list of 3 tuple of (int, str)'
  • 'list of 3 int'
  • 'tuple of 3 list of (int, str)'
  • 'dict of 3 int->str'

Correct: 'list of 3 tuple of (int, str)' — three tuples, each holding an integer and a string.

Why: list(zip(...)) always produces a list of tuples, and here each tuple pairs a number with a character. Option D is what dict(lt) would report, which is the natural next step — converting the pairs into a mapping, exactly as idea 2 did.

50. Worked example: mixed and nested shapes

Worked example

The descriptions get more useful as the structures get worse.

>>> t3 = [1, 2, 3, 4.0, '5', '6', [7], [8], 9]
>>> structshape(t3)
'list of (3 int, float, 2 str, 2 list of int, int)'
>>> lt = list(zip([1, 2, 3], 'abc'))
>>> structshape(lt)
'list of 3 tuple of (int, str)'
>>> structshape(dict(lt))
'dict of 3 int->str'
StructureIts shapeNote
a mixed listgrouped in order by typethe parentheses show a run
a list of tupleslist of 3 tuple of (int, str)the zip result
a dictionarydict of 3 int->strkey type to value type

Read the grouped form.

Why: If the elements of the list are not the same type, structshape groups them, in order, by type — so 3 int, float, 2 str describes a run of three integers, then a float, then two strings.

Read the list of tuples.

Why: list of 3 tuple of (int, str) is exactly what list(zip(...)) produces, which is a shape worth recognising on sight.

Read the dictionary form.

Why: dict of 3 int->str gives the size and the two types, which is usually the whole question.

Figure (svg): A panel showing a uniform shape description beside a grouped one revealing a stray element

Three descriptions of increasingly compound structures. The mixed-type one is the most telling, because an unexpectedly mixed list is a classic shape error.

Verify: Use it on a structure you believe is uniform.

Why: If structshape reports a grouped description like list of (5 int, str, 3 int) for something meant to be all integers, one element is the wrong type — and the position in the grouping tells you roughly where. That is a diagnosis that would take a loop and several prints to reach by hand.

51. Trap: reading a shape error as a value error

Trap

The trap

A function raises a TypeError and the developer starts checking whether the numbers are right.

Assume wrong output means wrong data

Why: Which is the usual case for arithmetic bugs.

The values may be perfectly correct and wrapped in one list too many. Checking them confirms they are right, which is true and unhelpful, and the actual problem is never examined.

The fix

Check the shape before the contents.

Print the type, or the structure's shape

Why: A common cause of runtime errors is a value that is not the right type, and printing the type is often enough.

Compare against what the function expects

Why: list with one integer against a plain old integer is a difference in packaging, not in value.

The tell is a TypeError or an AttributeError rather than a wrong number. Those messages are almost always about shape, and they point at the packaging rather than the contents.

52. Error analysis: four shape mistakes

Error analysis

Mark each and say what the shape is versus what was wanted.

Annotate

  • Line 1 passes one pair where a list of pairs was expected — 'tuple of 2 int' instead of 'list of tuple of (int, int)'. The function will iterate over two numbers instead of two pairs.
  • Line 2 hands dict a list of integers rather than a list of pairs, and it raises TypeError: cannot convert dictionary update sequence element to a sequence.
  • Line 3 loops over a dictionary, which yields keys, and unpacks each into two names. It works only if the keys are two-element tuples — otherwise ValueError. It probably wanted d.items().
  • Line 4 is a list whose tuples are not all the same size, so any loop unpacking two names raises ValueError partway through, after processing the first item.
  • So three raise and one silently misinterprets — and all four are about packaging rather than about the values, which are perfectly reasonable in every case.
  • The common diagnosis is to describe the shape you have and the shape the code expects, and compare the two descriptions rather than the data.

Shape errors produce TypeErrors and ValueErrors rather than wrong numbers, which is a useful signal about where to look.

53. Complete it: turn pairs into a mapping

Faded example

The shape is already right; it needs one conversion.

Fill in the blanks

lt = list(zip([1, 2, 3], 'abc'))
d = dict(lt)
# structshape(d) -> 'dict of 3 int->str'

Why: dict takes a list of tuples and makes each pair an item, so 'list of 3 tuple of (int, str)' becomes 'dict of 3 int->str'. The shapes line up because zip already produced exactly the two-element tuples dict requires — which is why dict(zip(...)) works as a single idiom.

54. Explain it: my data look right and nothing works

Explain it

The classic symptom of a shape error.

Discussion prompt

A classmate has checked every value in their structure and they are all correct, and the program still raises a TypeError. What should they look at instead?

Hint: Not the values.

Answer:

The packaging. A list containing one integer and a plain integer hold the same value and are not interchangeable, and no amount of checking the number will reveal that.

Have them print the type at each level, or use structshape to get a one-line description of the whole thing, and compare it with what the receiving code expects.

The signal is in the error kind: TypeError and AttributeError are almost always about shape, while a wrong answer with no exception is usually about values. Reading which kind of failure you have tells you which of the two to investigate.

55. Compare: dictionary and list of tuples

Comparison

Fill the blanks. Each is good at what the other cannot do.

Comparison matrix

QuestionDictionaryList of tuples
How do you find something?by key, in about constant timeby searching, in proportion to the length
Is there an order?noyes, and it can be sorted
Duplicate keys?impossible — later winsallowed, and kept
How do you convert?dict(pairs)list(d.items())

The bottom row is why the conversion matters: you move to whichever structure is good at the operation you need next.

56. The procedure: keying a dictionary by several fields

Pattern

Five steps. The third is where the shape errors come from.

  1. Decide which fields together identify one entry — that combination is the key.
  2. Store with directory[field1, field2] = value; the comma makes the bracketed expression a tuple.
  3. Keep the field order fixed, since the key is matched by position and a swapped lookup simply misses.
  4. Traverse by unpacking: for f1, f2 in directory, or for (f1, f2), value in directory.items().
  5. If a field is itself mutable, convert it — tuple(the_list) — because anything reachable from the key must be immutable.

Step 3's failure is a KeyError for something you can see in the dictionary, which is the standard symptom of a compound key assembled in the wrong order.

Python documentation — Data Structures Data Structures

57. Check yourself 1 of 3: items

Check

Two names, and a dictionary.

d = {'a': 1, 'b': 2}
total = 0
for k, v in d.items():
    total += v
print(total)
PassThe namesRunning total
pass 1k='a', v=1total 1
pass 2k='b', v=2total 3
the resultthe sum of the values3

Check your understanding

What does this print?

  • A. 3 (correct)
  • B. 0
  • C. A ValueError about unpacking
  • D. A TypeError, because k is a string

Answer: A

Why: items yields a key-value pair per pass, which the two names unpack, so v holds the values 1 and 2 and the total is 3. Looping over d directly instead would bind k to a key and there would be nothing to unpack, which is where the ValueError in option C would come from.

Why B tempts people
The loop runs twice and adds a positive value each time.
Why C tempts people
Each item is a two-element tuple and there are two names, so the counts match.
Why D tempts people
Nothing adds the string. Only v is accumulated, and it holds the integer values.

58. Check yourself 2 of 3: tuple keys

Check

Two values inside the brackets.

Check your understanding

What does directory['Cleese', 'John'] = number create?

  • A. One item whose key is the tuple ('Cleese', 'John') (correct)
  • B. Two items, one for each name
  • C. A nested dictionary, keyed first by surname
  • D. A TypeError, because brackets take only one index

Answer: A

Why: The expression in brackets is a tuple — the comma builds it, exactly as everywhere else — so a single key goes in. Printing the dictionary shows the key with its parentheses, confirming that one object was used rather than two separate indices.

Why B tempts people
Nothing splits the pair. The comma constructs one value rather than separating two.
Why C tempts people
This is what a grid-style reading would predict. The structure is flat, with compound keys.
Why D tempts people
The brackets do take one expression, and the comma made the two names into exactly one.

59. Check yourself 3 of 3: choosing a type

Check

The book gives three reasons to prefer a tuple.

Check your understanding

Which of these is NOT one of the book's reasons for preferring a tuple over a list?

  • A. Tuples use less memory than lists (correct)
  • B. A sequence used as a dictionary key must be immutable
  • C. In a return statement it is syntactically simpler to create a tuple
  • D. Passing a tuple reduces the potential for unexpected behaviour due to aliasing

Answer: A

Why: The three reasons the book gives are syntax, keys, and aliasing. Memory is not among them — and the book is explicit that lists are more common than tuples, mostly because they are mutable, so the tuple cases are specific exceptions rather than a general preference.

Why B tempts people
This is the second reason, and it is the one chapter 11 pointed forward to.
Why C tempts people
This is the first reason: return min(t), max(t) needs no brackets at all.
Why D tempts people
This is the third, and it works because a function cannot modify what it cannot modify.

60. Where this shows up outside this course

Real world

The difference between a record and a collection is not a programming idea.

Discussion prompt

Think of a form with a fixed set of fields and a shopping list. What is different about how you would change each, and how does that map onto tuples and lists?

Hint: One has a shape and the other has a length.

Answer:

You add to a shopping list and its length changes; you fill in a form and its shape does not. Adding a field to a form is redesigning it, not using it.

That is the tuple-and-list distinction: a list is a collection of similar things whose length varies, and a tuple is a record of a fixed number of positions where each means something particular.

And it explains the shape errors. Handing someone a form where a list was expected fails not because the contents are wrong but because the packaging is — which is exactly what a TypeError about a structure is telling you.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

In directory[last, first] = number, how many keys does the assignment create?

  • One — the brackets contain a tuple
  • Two — one for each name
  • One, but only if last and first are strings
  • It raises a TypeError: brackets take one index

Correct: One — the brackets contain a tuple, because the comma builds one.

Why: The book puts it in exactly those words: the expression in brackets is a tuple. This is not two-dimensional indexing, however much it resembles it — Python sees the comma, builds a two-element tuple, and uses that single object as the key. You can confirm it by printing the dictionary, which shows the key with its parentheses, and by trying the same thing with a list: directory[[last, first]] raises TypeError: unhashable type, which is the whole reason a tuple is used here. The types of the two names are irrelevant, so long as each is itself hashable — a tuple containing a list would fail for the same reason a list does.

62. Explain it to someone else

Explain it

Three types, and when each one is right.

Discussion prompt

A classmate asks how to decide between a string, a list and a tuple. Give them the two questions that settle it, and the one default.

Hint: The book says one of the three is more common.

Answer:

First question: are all the elements characters? If so a string will do, unless you need to change them, in which case a list of characters is the way.

Second: will it be a dictionary key, or must it not change? Then it has to be immutable, so a tuple.

Otherwise a list, which is the default — the book says lists are more common than tuples, mostly because they are mutable. The three tuple cases are worth recognising, and everything else is a list.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • items, and converting between dictionaries and lists of tuples
  • dict with zip, and what happens to duplicate keys
  • Tuples as keys, and why the brackets hold one expression
  • Choosing a sequence type, and recognising shape errors

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The conversions are worth having automatic, because so many operations are natural on pairs and impossible on a dictionary. The dict-with-zip idiom is compact and has two silent failure modes — truncation and duplicate keys — both worth knowing. The tuple-key syntax is where the grid misreading catches people, and stating the expression in brackets is a tuple out loud usually fixes it for good. And the choosing section is the closest the book comes to design advice, which makes it worth more than its length suggests.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw a dictionary on the left and a list of tuples on the right, with arrows between them labelled items() and dict(). Beside each, write one thing it can do that the other cannot. Underneath, draw the telephone directory as the book's figure 12.2 — tuple keys in Python syntax, mapping to numbers — and write beside it the line that stores an entry, with the comma circled. Finally list the three reasons to prefer a tuple over a list.

65. What you can do now

Recap

Three pages, and chapter 12 is finished: the three collection types working together.

If you remember one thingIt is this
From itemsA dictionary's pairs are tuples, so everything about tuples applies to them.
From dict(zip(...))The first sequence supplies the keys, and duplicates silently collapse.
From tuple keysThe expression in brackets is a tuple, not two indices.
From choosingLists are more common; the tuple's three advantages all follow from immutability.
From shape errorsThe values can all be right and the packaging still wrong.

The next chapter is a case study: reading a text file, counting words with the dictionaries of this chapter and the last, and using random selection to generate text — the first program in the book built from all three collection types at once.

Think Python, 2nd edition — Allen B. Downey §12.6-12.8, pp. 120-122 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §12.6-12.8, pp. 120-122
  2. Python documentation — Data Structures
  3. Python documentation — Built-in Types

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108