This lesson loops over dictionaries, writes a reverse lookup that has to search, introduces the raise statement and LookupError, inverts a dictionary using lists as values, and explains why keys must be hashable.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 11 — Dictionaries
§11.3-11.5, pp. 106-108
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-108 — the pages these objectives are drawn from
Warm-up
You have a histogram. Try asking it a question it was not built for.
Discussion prompt
You have a dictionary mapping letters to counts, and you want to know which letter appeared twice. There is no syntax for that. How would you find out, and what could go wrong with your method?
Hint: You know the search pattern from chapter 8.
Answer:
You would have to look at every item and check whether its value is 2 — a search, rather than a lookup.
Two things could go wrong. There might be more than one letter with a count of 2, in which case which letter is not a well-posed question. And there might be none, in which case there is no answer to return.
Both problems are named in this lesson, and the second one is what introduces a new statement: a way for your own code to raise an exception rather than returning something misleading.
Concept
Given a dictionary d and a key k, it is easy to find the corresponding value v = d[k]. This operation is called a lookup. But what if you have v and you want to find k?
lookup — A dictionary operation that takes a key and finds the corresponding value.
You have two problems: first, there might be more than one key that maps to the value v — you might be able to pick one, or you might have to make a list of all of them. Second, there is no simple syntax to do a reverse lookup; you have to search.
Figure (svg): Two columns contrasting a forward lookup with a reverse lookup
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-106
Section
Section 1
Concept
If you use a dictionary in a for statement, it traverses the keys of the dictionary. To see the values too, you look each one up inside the loop.
def print_hist(h):
for c in h:
print(c, h[c])
>>> h = histogram('parrot')
>>> print_hist(h)
a 1
p 1
r 2
t 1
o 1| Part | What it does | Note |
|---|---|---|
| for c in h | c takes each KEY in turn | not the value, not the pair |
| h[c] | looks up the corresponding value | the ordinary forward lookup |
| the output order | no particular order | as always |
Again, the keys are in no particular order — this is the same unpredictability from the previous lesson, showing up in the place where it is most visible.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-106
Picture it
Each pass gives one key, and the value comes from a lookup.
Figure (svg): A pipeline showing a dictionary yielding one key per pass and a lookup producing the value
That is a design choice worth noticing: looping gives keys because keys are what everything else in a dictionary is indexed by.
Worked example
The dictionary has no order. sorted makes one.
>>> h = histogram('parrot')
>>> for key in sorted(h):
... print(key, h[key])
a 1
o 1
p 1
r 2
t 1| Part | What it produces | Note |
|---|---|---|
| sorted(h) | a sorted LIST of the keys | the dictionary is unchanged |
| for key in ... | walks that list | in a defined order |
| h[key] | the value, by lookup | as before |
Recognise what sorted is given.
Why: A dictionary in a for statement yields keys, and sorted works the same way — it receives the keys and returns them sorted.
Note that it returns a list.
Why: sorted creates a new list, so the order lives in that list and not in the dictionary, which has no order to change.
Look up inside the loop as usual.
Why: The sorted list carries keys only, so the values still come from h[key].
Figure (svg): The state of the program after each line of Worked example traversing the keys in sorted order, drawn as a ladder with one rung per traced line
The five letters in alphabetical order with their counts. The ordering was created by sorted and does not belong to the dictionary.
Verify: Print the dictionary again afterwards.
Why: It comes back in the same unpredictable order it had before, because sorted returned a new list rather than rearranging anything — the returning-not-modifying distinction from lesson 10c. There is no operation that sorts a dictionary in place, and after the previous lesson it should be clear why the idea does not make sense.
Prediction
Watch what the loop variable holds.
h = {'a': 1, 'b': 2}
for x in h:
print(x)| Part | What happens | Output |
|---|---|---|
| for x in h | x takes each key | 'a', then 'b' |
| print(x) | prints the key | not the value |
| to get values | print(h[x]) | or loop over h.values() |
Predict first
What does this print?
Correct: the keys: a and b — a dictionary in a for statement traverses its keys.
Why: The name x has no influence on what the loop yields; a dictionary always gives its keys. To print the values you would write print(h[x]) inside the loop, or loop over h.values(). Note that the order of the two lines is not guaranteed, which is the previous lesson's point showing up again.
Worked example
One is written with what you know now, and one is the tool for the job.
# loop over keys and look each value up
for c in h:
print(c, h[c])
# or ask for the pairs directly
for c, count in h.items():
print(c, count)| Form | What each pass gives | Note |
|---|---|---|
| for c in h | one key, then a lookup | two steps |
| h.items() | one key-value pair per pass | unpacked into two names |
| both | produce the same output | the second says the intent |
Read the first form.
Why: This is what the book uses at this point: the loop gives the key and h[c] fetches the value.
Read the second.
Why: items yields the pairs, and the two names on the left of in unpack each pair — which is a shape you will meet properly in the next chapter.
Prefer the one that states the intent.
Why: If you need both parts of every item, saying so is clearer than fetching the second one yourself.
Figure (svg): Two columns comparing looping over keys with looping over items
Identical output. The first is the form built from what you already know; the second says for each key and its value, which is usually what you meant.
Verify: Ask what the first form does that the second does not.
Why: It performs a lookup per pass, which is fast but not free, and it works when you only sometimes need the value. Neither is wrong — but a loop that looks up every key it is given is asking for the pairs the long way round.
Trap
A student writes for count in h: print(count) expecting to see the counts.
Name the loop variable after what you want
Why: The name is the only clue in the line, and naming it count makes it look like a count.
The loop gives keys whatever you call them, so this prints the letters. Nothing raises — the name is a label, not a request, and the program quietly reports the wrong column.
Name the loop variable after what it holds.
for c in h names a key
Why: And if you want the value, h[c] fetches it.
Or loop over h.values() if the keys are irrelevant
Why: Which says explicitly that you want the other column.
This is the same silent-wrong-answer failure as using in on values: a dictionary has two halves and the default is always the keys. When a program reports the wrong half, the first thing to check is which half it asked for.
Faded example
The dictionary has no order, so make one.
Fill in the blanks
for key in sorted(h):
print(key, h[key])
Why: sorted receives the keys and returns them as a new sorted list, which the loop then walks in a defined order. The dictionary itself is unchanged — there is no way to sort one, because it has no arrangement to sort. Note that h.sort() would raise an AttributeError: that method belongs to lists.
Discrimination
A dictionary has two columns and each expression picks one.
Sort into buckets
For each expression, which part of the dictionary does it produce?
Socratic
It could have given values, or pairs. It gives keys.
Discussion prompt
A dictionary could plausibly yield values or pairs when looped over. Why is yielding keys the sensible default?
Hint: What can you do with a key that you cannot do with a value?
Answer:
From a key you can reach everything: h[key] gives the value, so a loop over keys can produce anything a loop over pairs could.
From a value you can reach nothing. There is no way back to the key without a search, which is the entire subject of the next section.
So keys are the more useful thing to be given, because the mapping runs that way. Yielding the half you can navigate from is the same reasoning that makes the in operator check keys — both defaults point along the direction the structure is built for.
Section
Section 2
Concept
Here is a function that takes a value and returns the first key that maps to that value. It is yet another example of the search pattern, but it uses a feature we haven't seen before.
def reverse_lookup(d, v):
for k in d:
if d[k] == v:
return k
raise LookupError()| Line | What it does | Note |
|---|---|---|
| for k in d | traverse the keys | the only way in |
| if d[k] == v | look up and compare | the search pattern |
| return k | the FIRST key that matches | there may be others |
| raise LookupError() | reached only if nothing matched | no key to return |
A reverse lookup is much slower than a forward lookup; if you have to do it often, or if the dictionary gets big, the performance of your program will suffer.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-107
Picture it
The forward direction touches one item. The reverse direction touches all of them.
Figure (svg): A growth chart comparing constant-time forward lookup against linear reverse lookup
If a program does reverse lookups often, the fix is usually to build the inverse dictionary once, which is exactly what idea 4 does.
Worked example
Find the letter that appeared twice.
>>> h = histogram('parrot')
>>> h
{'a': 1, 'p': 1, 'r': 2, 't': 1, 'o': 1}
>>> key = reverse_lookup(h, 2)
>>> key
'r'| Pass | The comparison | What happens |
|---|---|---|
| k = 'a' | h['a'] is 1, not 2 | keep going |
| k = 'p' | h['p'] is 1, not 2 | keep going |
| k = 'r' | h['r'] is 2 | return 'r' |
| the rest | never examined | return leaves at once |
Traverse the keys.
Why: The loop is the only way to reach the items, since there is no syntax that goes from a value to a key.
Compare each value.
Why: h[k] == v is a forward lookup used to test one item at a time, which is the search pattern from chapter 8.
Return on the first match.
Why: return leaves the function immediately, so the remaining items are never examined — which is why the answer is the first key found, not the only one.
Figure (svg): The state of the program after each line of Worked example a successful reverse lookup, drawn as a ladder with one rung per traced line
'r'. It is the first key whose value is 2 in the order the dictionary happens to traverse, which for this histogram is the only one.
Verify: Ask what would happen with two matching keys.
Why: For histogram('parrots'), both 'r' and 's' map to 2, and the function returns whichever comes first in an unpredictable order. It would be correct and unrepeatable — which is the first of the two problems the book named, and it is the reason idea 4 collects all the matches rather than one.
Prediction
Two keys map to the same value.
h = {'r': 2, 's': 2, 'a': 1}
print(reverse_lookup(h, 2))| Aspect | What is true | Note |
|---|---|---|
| the search | returns on the first match | and stops |
| which is first | depends on traversal order | unpredictable |
| the consequence | 'r' or 's', not both | and not reliably either |
Predict first
What can you say about the result?
Correct: It is 'r' or 's', and which one is not something you can rely on — the function returns on the first match in an unpredictable traversal order.
Why: reverse_lookup returns the first key it finds with the given value, and the order of items in a dictionary is unpredictable, so the result is one of the two matches with no promise about which. The function raises only when nothing matches, and it never returns a list — collecting all the matches is what invert_dict does.
Worked example
No key maps to 3. The function has nothing to return.
>>> key = reverse_lookup(h, 3)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 5, in reverse_lookup
LookupError| Stage | What happens | Note |
|---|---|---|
| the loop | examines every key | none has the value 3 |
| falls off the end | reaches the raise | no key to return |
| the traceback | two frames deep | the caller, then the function |
Follow the loop to its end.
Why: Every key is tested and none matches, so control reaches the line after the loop.
See what that line does.
Why: If we get to the end of the loop, that means v doesn't appear in the dictionary as a value, so we raise an exception.
Read the traceback.
Why: It shows two frames — the call at the prompt and the function itself — which is chapter 5's stack diagram appearing in an error message.
Figure (svg): A panel showing the traceback from a failed reverse lookup with its two frames
A LookupError, raised deliberately by the function. The effect when you raise an exception is the same as when Python raises one: it prints a traceback and an error message.
Verify: Ask what the alternative would have been.
Why: Returning None. That would be legal and quiet, and the caller would carry a None around until it was used as a key or an index and failed somewhere else entirely — the NoneType failures from chapter 6. Raising here puts the error where the problem is, which is worth more than the convenience.
Trap
A program calls reverse_lookup to find the word with a given frequency and uses the result as though it were unique.
Read the function's name as a lookup
Why: A forward lookup has exactly one answer, so the reverse of one sounds like it should too.
There might be more than one key that maps to the value, and the function returns whichever the traversal reaches first — in an order the language does not promise. The program is correct only by accident.
Decide what multiple matches should mean before you write the function.
If any one will do, say so
Why: Name the function first_key_with or similar, so the caller knows what they are getting.
If you need them all, collect them
Why: Build a list inside the loop and return it, which is what invert_dict does for every value at once.
The book names this as the first of the two problems with reverse lookup, and it is the more dangerous one — the second problem, the lack of syntax, merely costs you a loop.
Invariant
Step through and watch how much work each case costs.
Step through it
Which case does the most work, and what does that say about using reverse_lookup in a loop?
The failing search is the most expensive: it examines every item and returns nothing. So a program that reverse-looks-up many values, most of which are absent, does the maximum work every time — which is exactly the situation the book warns about when it says the performance of your program will suffer.
Ranking
Four steps, in the order they happen.
Put in order
Why: The loop takes a key, performs a forward lookup to get its value, compares, and either returns or continues. The raise is reached only after the loop has ended without returning — which is the point of putting it after the loop rather than inside it.
Explain it
The same dictionary, and two questions that cost very different amounts.
Discussion prompt
A classmate asks why d[k] is instant but finding the key for a value needs a loop. Explain it in terms of how a dictionary stores things.
Hint: Where does the dictionary put an item when you add it?
Answer:
When you add an item, the dictionary computes a location from the KEY and stores the pair there. So given a key it can compute where to look and go straight to it.
Nothing is computed from the value, so there is no location to compute when you start from one. The only option is to examine items until you find a match.
That is why the book says a reverse lookup is much slower than a forward one, and why the fix — when you need it often — is to build a second dictionary keyed the other way. You pay for the traversal once instead of every time.
Section
Section 3
Concept
The raise statement causes an exception. In reverse_lookup it causes a LookupError, which is a built-in exception used to indicate that a lookup operation failed.
>>> raise LookupError()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
LookupError
>>> raise LookupError('value does not appear in the dictionary')
Traceback (most recent call last):
File "<stdin>", line 1, in ?
LookupError: value does not appear in the dictionary| Statement | What it produces | Note |
|---|---|---|
| raise LookupError() | an exception with no message | the name alone |
| raise LookupError('...') | an exception with a message | shown after the colon |
| the effect | a traceback and an error message | as if Python had raised it |
The effect when you raise an exception is the same as when Python raises one. And when you raise an exception, you can provide a detailed error message as an optional argument — which is worth doing, because the person reading it will not have your function in front of them.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 107-107
Picture it
An exception you raise and one Python raises behave identically.
Figure (svg): Two columns showing that a raised exception and a Python exception behave the same way
That is the point of the section: error reporting is not something only the interpreter can do.
Worked example
The message is what the reader will actually see.
def reverse_lookup(d, v):
for k in d:
if d[k] == v:
return k
raise LookupError('value does not appear in the dictionary')| Stage | What happens | Note |
|---|---|---|
| no match found | the loop ends | control reaches the raise |
| the message | passed as an argument | printed after the colon |
| the reader | sees what went wrong | without reading the source |
Put the raise after the loop.
Why: It is reached only when the loop finished without returning, which is exactly the case where there is no key to give back.
Supply a message.
Why: You can provide a detailed error message as an optional argument, and here it says which of the many possible lookup failures this is.
Consider what to include.
Why: Naming the value that was not found — as a KeyError names its key — makes the message diagnostic rather than merely descriptive.
Figure (svg): The state of the program after each line of Worked example raising with a useful message, drawn as a ladder with one rung per traced line
A function that fails clearly. Compare the two forms in the traceback: bare LookupError versus LookupError with a sentence, and only the second one tells a reader anything.
Verify: Read the traceback as someone who did not write the function.
Why: With the bare form they see a name they may not recognise and a line number in code they did not write. With the message they see what the function was trying to do and why it could not — which is the difference between an error they can act on and one they have to investigate.
Prediction
The optional argument is the part a reader sees.
raise LookupError('value does not appear in the dictionary')| Part | What it is | Where it appears |
|---|---|---|
| the exception type | LookupError | before the colon |
| the argument | the detailed message | after the colon |
| the effect | traceback plus message | as if Python raised it |
Predict first
What is printed on the last line of the traceback?
Correct: LookupError: value does not appear in the dictionary — the type, a colon, and the message you supplied.
Why: The message is an optional argument, and when you supply one it appears after the exception name and a colon. That is the same shape as every error message you have read so far — KeyError: 'four', TypeError: ... — because it is the same mechanism, whoever raised it.
Worked example
Two ways to report failure, and they fail in different places.
# raises: fails at the point of the problem
def lookup_or_raise(d, v):
for k in d:
if d[k] == v:
return k
raise LookupError('no such value')
# returns None: fails later, somewhere else
def lookup_or_none(d, v):
for k in d:
if d[k] == v:
return k
return None| Version | What happens on failure | Where the error appears |
|---|---|---|
| the raising version | stops at the failure | the traceback points here |
| the None version | returns quietly | the caller carries None away |
| where the None version fails | wherever the None is used | possibly far away |
Consider what the caller does with a None.
Why: It usually gets used — as a key, an index, or in arithmetic — and fails there with a NoneType message that names nothing about the real problem.
Consider what the caller does with an exception.
Why: It stops, with a traceback whose deepest frame is the function that could not do its job.
Choose by whether absence is expected.
Why: If the caller routinely asks about values that may not be there, returning None is reasonable and the caller must check. If absence means something is wrong, raise.
Figure (svg): A flowchart comparing where a raising function and a None-returning function report a failure
Raising when failure is exceptional, returning None when it is routine and the caller will check. The cost of the None version is the distance between where it fails and where the problem is.
Verify: Trace the None version's failure to its cause.
Why: The traceback names a line that uses the result, several steps from the lookup that returned nothing — the exact gap chapter 10 warned about with NoneType errors. That distance is the price, and it is why the book's version raises.
Trap
A student writes the raise inside the for loop, in an else clause of the if.
Handle both outcomes of the comparison
Why: Every if wants an else, so the failure case goes there.
Now the function raises on the very first key that does not match, which is almost always the first key. It reports failure before it has finished looking, so a value that IS in the dictionary is reported as absent.
The raise belongs after the loop.
Ask when you actually know the search has failed
Why: Only when every key has been checked — which is when the loop ends.
Read the indentation as the answer
Why: At the loop's indentation the raise runs after all passes; inside, it runs during one.
This is the search pattern's standard shape, and it is the same reason a found-flag is initialised before a loop rather than inside it: a conclusion about all the items cannot be drawn from one of them.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C confuses raising with returning. raise does not return anything — it stops the function and unwinds to whoever is prepared to handle the exception, printing a traceback if nobody is. Returning a value the caller checks is the alternative design, and choosing between them is exactly the raise-or-return-None decision.
Faded example
The bare exception name is legal. A message is better.
Fill in the blanks
def reverse_lookup(d, v):
for k in d:
if d[k] == v:
return k
raise LookupError('value does not appear in the dictionary')
Why: The raise goes after the loop, because only then do you know that no key matched. LookupError is the built-in exception used to indicate that a lookup operation failed, and the string argument becomes the message a reader sees after the colon.
Real world
The raise-or-return-None choice is a general one about error reporting.
Discussion prompt
Think of a system that detects a problem and does not report it until much later. What does the delay cost the person diagnosing it?
Hint: Anything with a form, a queue, or a batch process.
Answer:
A form that accepts a bad value and fails at submission; a batch job that reads a malformed file at midnight and reports at nine; a build that compiles a bad configuration and fails at deploy.
What the delay costs is the connection between the symptom and the cause. By the time the failure appears, the information that would explain it — which field, which line, which setting — is gone.
Raising at the point of failure is the programming version of validating at the point of entry. It is more annoying in the moment and much cheaper afterwards, which is the trade in every case.
Section
Section 4
Concept
Lists can appear as values in a dictionary. If you have a dictionary that maps from letters to frequencies, you might want to invert it — but since there might be several letters with the same frequency, each value in the inverted dictionary should be a list of letters.
singleton — A list with a single element.
def invert_dict(d):
inverse = dict()
for key in d:
val = d[key]
if val not in inverse:
inverse[val] = [key]
else:
inverse[val].append(key)
return inverse| Pass | The test | The inverse after |
|---|---|---|
| key = 'a', val = 1 | 1 not in inverse | inverse[1] = ['a'] |
| key = 'p', val = 1 | 1 is in inverse | append: [1] becomes ['a', 'p'] |
| key = 'r', val = 2 | 2 not in inverse | inverse[2] = ['r'] |
| key = 't', val = 1 | 1 is in inverse | ['a', 'p', 't'] |
Each time through the loop, key gets a key from d and val gets the corresponding value. If val is not in inverse, we haven't seen it before, so we create a new item and initialize it with a singleton. Otherwise we append the corresponding key to the list.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 107-108
Picture it
The book draws integers and strings inside the box and lists outside it, to keep the diagram simple.
Figure (svg): A state diagram showing the histogram dictionary and its inverse with list values drawn outside the box
Every list is drawn outside the box because it is a separate object the dictionary refers to — the same reference idea as chapter 10, in a new picture.
Worked example
Letters to counts becomes counts to lists of letters.
>>> hist = histogram('parrot')
>>> hist
{'a': 1, 'p': 1, 'r': 2, 't': 1, 'o': 1}
>>> inverse = invert_dict(hist)
>>> inverse
{1: ['a', 'p', 't', 'o'], 2: ['r']}| Aspect | What happened | Note |
|---|---|---|
| the keys become values | letters end up in lists | several per count |
| the values become keys | counts are now the index | and they are integers |
| the sizes differ | five items become two | because counts repeat |
Notice the shape changed.
Why: The original has one letter per item; the inverse has one count per item and several letters inside each.
Notice why it had to.
Why: Four letters share the count 1, and each key is associated with a single value — so the four letters have to go into one value, which means a list.
Notice the keys are now integers.
Why: A dictionary's keys can be almost any type, and here they are counts rather than characters, which is the generality from the previous lesson being used for real.
Figure (svg): The state of the program after each line of Worked example inverting a histogram, drawn as a ladder with one rung per traced line
{1: ['a', 'p', 't', 'o'], 2: ['r']}. Five items become two, because the original values repeat and the inverse collects the duplicates rather than losing them.
Verify: Check that no information was lost.
Why: Counting the letters across all the lists gives five, which is the number of items in the original. That check is the point of using lists: a naive inversion writing inverse[val] = key would have kept one letter per count and silently discarded three.
Prediction
Two keys share a value.
d = {'a': 1, 'b': 2, 'c': 1}
print(invert_dict(d))| Pass | The test | The inverse after |
|---|---|---|
| key 'a', val 1 | 1 is new | {1: ['a']} |
| key 'b', val 2 | 2 is new | {1: ['a'], 2: ['b']} |
| key 'c', val 1 | 1 is known | {1: ['a', 'c'], 2: ['b']} |
Predict first
What does this print?
Correct: {1: ['a', 'c'], 2: ['b']} — the two keys sharing the value 1 both end up in one list.
Why: Three items become two, because 'a' and 'c' share a value and the inverse collects them rather than losing one. The second option is what a naive inversion would produce, discarding 'a'. The third is impossible — a dictionary cannot show the same key twice, since each key is associated with a single value.
Worked example
The first sighting has to create something an append can add to.
if val not in inverse:
inverse[val] = [key] # a singleton, not key
else:
inverse[val].append(key)| Case | What is stored | Why it matters |
|---|---|---|
| first sighting | inverse[val] = [key] | a list holding one thing |
| later sightings | inverse[val].append(key) | the list grows |
| if it were = key | a string, not a list | append would raise |
Look at the brackets on the first branch.
Why: inverse[val] = [key] stores a list containing the key, not the key itself. That list is a singleton — a list that contains a single element.
Ask what the else branch needs.
Why: It calls append on whatever is there, and append is a list method. So the first branch must leave a list behind.
Predict the failure without the brackets.
Why: Storing the key directly leaves a string, and 'a'.append('p') raises an AttributeError, because strings have no append.
Figure (svg): A flowchart showing the create-a-singleton branch and the append branch of invert_dict
The singleton exists so that every value in the inverse is a list from the start, which lets the else branch treat every later sighting identically.
Verify: Compare with the get trick from the previous lesson.
Why: The same structure appears: a first sighting that has nothing to add to. Here it could be written inverse.setdefault(val, []).append(key) or inverse[val] = inverse.get(val, []) + [key] — both remove the conditional the same way get did for the histogram, by making the first sighting produce an empty starting value.
Trap
A student writes inverse[val] = key in the first branch, since only one key has been seen.
Store what you have
Why: One key has been found, so storing one key looks right.
The value is now a string rather than a list, so the first time a second key shares that value, inverse[val].append(key) raises AttributeError — a failure that appears only on data with a repeat.
Store a list containing the key.
inverse[val] = [key]
Why: A singleton: a list with a single element, ready to grow.
Keep the value type uniform
Why: Every value in the inverse is a list, whether it holds one key or four — so nothing has to check.
The general rule is worth naming: when a value may need to hold several things, make it a collection from the beginning. Special-casing the first one produces a structure whose values have two possible types, and every reader afterwards has to handle both.
Faded example
What has to be stored so that append works later?
Fill in the blanks
if val not in inverse:
inverse[val] = [key]
else:
inverse[val].append(key)
Why: A singleton — a list containing the one key seen so far. Storing key alone would leave a string in the dictionary, and the else branch's append would raise AttributeError as soon as a second key shared that value. Making the value a list from the start keeps every value the same type.
Comparison
Fill the blanks. The same shape, with a different thing accumulated.
Comparison matrix
| Question | histogram | invert_dict |
|---|---|---|
| What are the keys? | characters | the original dictionary's values |
| What is accumulated? | a count | a list of keys |
| First sighting stores | 1 | [key], a singleton |
| Later sightings do what? | d[c] += 1 | inverse[val].append(key) |
Both are the same pattern: for each distinct thing, accumulate a fact about it. Only the fact differs.
Explain it to yourself
The original held single values. The inverse cannot.
Discussion prompt
Explain why inverting a dictionary needs lists as values, using the fact that each key maps to exactly one value.
Hint: What happens when two keys have the same value?
Answer:
In the original, several keys may map to the same value — four letters can all have the count 1, and nothing forbids that.
When you turn it around, that count becomes a key, and a key is associated with a single value. Four letters cannot each be that single value.
So the single value has to be one thing that contains four — a list. The constraint is not about lists at all; it is that keys are unique and values are not, so inverting has to collapse the many into one collection.
Section
Section 5
Concept
Lists can be values in a dictionary, as invert_dict shows, but they cannot be keys. Here's what happens if you try.
hash function — A function used by a hashtable to compute the location for a key.
>>> t = [1, 2, 3]
>>> d = dict()
>>> d[t] = 'oops'
Traceback (most recent call last):
File "<stdin>", line 1, in ?
TypeError: list objects are unhashable| Line | What happens | Note |
|---|---|---|
| d[t] = 'oops' | t is a list | TypeError |
| the word unhashable | names the actual requirement | keys must be hashable |
| why | because of how dictionaries store items | the next few lines |
A hash is a function that takes a value of any kind and returns an integer. Dictionaries use these integers, called hash values, to store and look up key-value pairs.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 108-108
Picture it
Hash the key, store it there. Then change the key.
Figure (svg): A flowchart showing an item stored at a computed location and then becoming unfindable after the key changes
So the dictionary would have an item it cannot find, or two entries for what is now the same key. Either way, it would not work correctly.
Worked example
The book's reasoning, step by step.
# if this were allowed:
t = [1, 2]
d = {}
d[t] = 'x' # hashed and filed under hash([1, 2])
t.append(3) # the key is now [1, 2, 3]
d[t] # looks under hash([1, 2, 3]) - nothing there| Stage | What is computed | Consequence |
|---|---|---|
| store | location computed from [1, 2] | the pair goes there |
| modify | the key becomes [1, 2, 3] | the stored pair does not move |
| look up | location computed from [1, 2, 3] | a different place |
Store the pair.
Why: When you create a key-value pair, Python hashes the key and stores it in the corresponding location.
Modify the key.
Why: This is only possible because the key is mutable — a string or an integer could not be changed like this.
Look it up.
Why: If you modify the key and then hash it again, it would go to a different location. In that case you might have two entries for the same key, or you might not be able to find a key.
Figure (svg): The state of the program after each line of Worked example the argument in full, drawn as a ladder with one rung per traced line
The dictionary would not work correctly. That's why keys have to be hashable, and why mutable types like lists aren't.
Verify: Check that immutable keys cannot cause this.
Why: A string key cannot be modified at all, so its hash value is fixed for as long as it exists and the item can always be found where it was filed. The restriction is not arbitrary: it is exactly what is needed to make the storage scheme work, and it is the payoff of chapter 10's mutability distinction.
Prediction
A list appears twice, in two different positions.
t = [1, 2]
d = {}
d['nums'] = t
d[t] = 'nums'| Line | What it does | Result |
|---|---|---|
| d['nums'] = t | a list as a VALUE | fine |
| d[t] = 'nums' | a list as a KEY | TypeError |
| the reason | keys are hashed | values are not |
Predict first
Which line raises an error?
Correct: Line 4, because a list cannot be a key — keys must be hashable and lists are not.
Why: Line 3 is fine: lists can appear as values in a dictionary, which is exactly what invert_dict relies on. Line 4 raises TypeError: list objects are unhashable, because the key is what gets hashed to decide where the pair is stored. The restriction lands on one side of the colon only.
Worked example
Immutable things, and the way round the restriction.
d = {}
d['abc'] = 1 # a string: fine
d[42] = 2 # an integer: fine
d[3.14] = 3 # a float: fine
d[(1, 2)] = 4 # a tuple: fine, and we meet these next chapter
# d[[1, 2]] = 5 # a list: TypeError, unhashable| Type | Mutable? | Usable as a key? |
|---|---|---|
| strings, numbers | immutable | hashable |
| tuples | immutable sequences | hashable, and the way round the restriction |
| lists, dictionaries | mutable | unhashable |
Notice the pattern.
Why: Everything immutable you have met can be a key, and the mutable types cannot. Hashability tracks mutability almost exactly.
Notice what values may be.
Why: Anything at all. invert_dict put lists in as values, and the restriction has never applied to that side.
Note the escape route.
Why: The simplest way to get around this limitation is to use tuples, which we will see in the next chapter.
Figure (svg): Two columns contrasting what may be a key with what may be a value
Immutable things make keys; anything makes a value. The tuple is the immutable sequence, and it exists partly so that a sequence can be a key.
Verify: Ask why the restriction is one-sided.
Why: Because only the key determines where the pair is stored. A value can change without moving anything, so there is no reason to restrict it — which is why the same list that is illegal on the left of the colon is perfectly legal on the right.
Trap
A student concludes that lists cannot be used with dictionaries at all.
Generalise from the error
Why: TypeError: list objects are unhashable does sound like a blanket ban.
Lists are perfectly good values — invert_dict is built on that — and the restriction applies only to keys. Avoiding lists entirely rules out the chapter's own main example.
Read the error as being about the position, not the type.
Keys are hashed; values are not
Why: So the requirement lands on one side of the colon only.
If you need a sequence as a key, use a tuple
Why: The book points forward to this explicitly.
The precise statement is that keys have to be hashable, and mutable types like lists aren't. Everything else about lists in dictionaries is unaffected.
Sorting
The question is whether it can change.
Sort into buckets
For each value, can it be used as a dictionary key?
Error analysis
Mark each and say what happens.
Annotate
The gap between the mistake in line 3 and the failure in line 4 is the same cause-and-symptom distance the whole course keeps meeting.
Socratic
The restriction sounds technical. It is a direct consequence of how storage works.
Discussion prompt
Explain, without using the word hashable, why a key that can change would break a dictionary.
Hint: Where does the dictionary decide to put the item?
Answer:
The dictionary decides where to put an item by computing something from the key. So the location is a fact about the key's contents at the moment it was stored.
If the contents then change, the computation gives a different answer, and the dictionary looks in a place where the item never was. The item has not moved; the address has.
So the requirement is really that a key's contents be fixed for as long as it is in use — which is precisely what immutability means. The restriction is not an implementation quirk, it is the minimum condition for the storage scheme to work at all.
Comparison
Fill the blanks. One is an operation and the other is a search.
Comparison matrix
| Question | Forward: d[k] | Reverse: value to key |
|---|---|---|
| Is there syntax for it? | yes, bracket lookup | no — you write a loop |
| How many answers? | exactly one | zero, one, or many |
| Cost | about constant | grows with the dictionary |
| What happens on failure? | KeyError, automatically | whatever you write — the book raises LookupError |
Every row is a consequence of the same fact: the dictionary computes its storage location from the key and never from the value.
Pattern
Six steps. It is the histogram pattern with a list where the counter was.
Step 4's brackets are the whole pattern. Storing the item rather than a list containing it leaves values of two different types, and the else branch fails on the first repeat.
Python documentation — Data Structures Data Structures
Check
A dictionary in a for statement.
h = {'a': 1, 'b': 2}
total = 0
for x in h:
total += h[x]
print(total)| Pass | The lookup | Running total |
|---|---|---|
| x = 'a' | h['a'] is 1 | total is 1 |
| x = 'b' | h['b'] is 2 | total is 3 |
| the sum of the values | 3 |
Check your understanding
What does this print?
Answer: A
Why: The loop gives keys, and h[x] converts each key to its value, so the additions are 1 and 2 and the total is 3. Writing total += x instead would try to add a string to an integer and raise the TypeError described in option C — which is the difference the lookup inside the loop is there to make.
Check
The search ends without a match.
Check your understanding
In reverse_lookup, why does the raise statement come after the loop rather than inside it?
Answer: B
Why: A conclusion about all the items cannot be drawn from one of them. Inside the loop, the raise would fire on the first key that does not match — almost always the first key — and report a value as absent when it is present two items later. Placing it after the loop means it is reached only when the loop finished without returning.
Check
One of these positions is restricted.
Check your understanding
Which statement is true about lists and dictionaries?
Answer: B
Why: Lists can appear as values — invert_dict depends on it — and cannot appear as keys, because the key is what gets hashed to decide where the pair is stored and a mutable key could move afterwards. The simplest way round the restriction is a tuple, which the next chapter introduces.
Real world
Building an index is the everyday version of inverting a dictionary.
Discussion prompt
The index at the back of a book maps topics to page numbers, and the book itself maps page numbers to content. What work did someone do to produce the index, and why is each entry a list?
Hint: They read it once and wrote down what they found.
Answer:
Someone traversed the whole book once, and for each topic they encountered, added the page number to that topic's list. That is invert_dict, done by hand.
Each entry is a list because a topic appears on several pages — the same reason the inverse of a histogram holds lists. Several keys mapping to one value become one key mapping to several.
And the payoff is the same one the book names: a reverse lookup is much slower than a forward lookup, so if you need it often, you traverse once and build the inverse. That is what an index is for, and it is why nobody reads a whole book to find a topic.
Commit first
Answer, then rate your confidence.
Predict first
You write d[[1, 2]] = 'x'. What happens, and why?
Correct: TypeError, because keys must be hashable and lists are mutable — the error message is 'list objects are unhashable'.
Why: The reason is worth being able to state. A dictionary computes a location from the key and stores the pair there. If the key could change, hashing it again would give a different location, so you might end up with two entries for the same key or be unable to find one — either way, the dictionary wouldn't work correctly. Python cannot trust a promise not to modify a list, so it rejects mutable keys outright rather than allowing the fourth option. The same list is perfectly legal as a value, because nothing is computed from values.
Explain it
Two directions, and only one of them is cheap.
Discussion prompt
A classmate wants to find which key maps to a given value and asks why there is no d.reverse(). Explain what they would have to write and why the language does not provide it.
Hint: Ask what the answer would even be.
Answer:
They would have to loop over every key, look up its value, and compare — the search pattern, and the whole dictionary examined in the worst case.
The language does not provide it partly because the answer is not well defined: several keys can map to one value, so the key does not exist. Any built-in would have to choose, and no choice is right for everyone.
And partly because it would hide the cost. d[k] looks cheap and is cheap; a built-in reverse lookup would look equally cheap and traverse everything. Making it a visible loop keeps the price visible — and if they need it often, the answer is invert_dict, paid for once.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The looping is mechanical once you have internalised that the loop yields keys, which is the one thing people get wrong. Reverse lookup is worth being uncomfortable about, because its two problems — several answers, and the cost — are the reason the rest of the chapter exists. The raise statement is your first time causing an error rather than receiving one, and the design question behind it recurs everywhere. And the hashability argument is the payoff of chapter 10: it is the moment mutability stops being a curiosity and starts constraining what you can write.
Connect it up
One page, from memory.
Draw it
Draw a dictionary and, beside it, two arrows: one from a key to its value labelled one operation, and one from a value back to a key labelled a search over everything. Under the second, write the two problems the book names. Then write invert_dict from memory, circling the singleton, and beneath it write in one sentence why a list can be on the right of the colon and not on the left.
Recap
Three pages, and the asymmetry that shapes every use of a mapping.
| If you remember one thing | It is this |
|---|---|
| From looping | A dictionary yields keys, whatever you name the loop variable. |
| From reverse lookup | It is a search, not a lookup — and it may have zero or many answers. |
| From raise | Report the failure where it happens, not by returning None for someone else to trip over. |
| From invert_dict | When a value may hold several things, make it a list from the first one. |
| From hashability | Only keys are hashed, so only keys are restricted. |
The next lesson uses a dictionary as machinery rather than as data: memos, which turn an impossibly slow recursive Fibonacci into an instant one, and the global statement — plus the debugging advice for when your data are too big to print.
Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-108 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.