11b Looping, Reverse Lookup, and Dictionaries of Lists

This lesson loops over dictionaries, writes a reverse lookup that has to search, introduces the raise statement and LookupError, inverts a dictionary using lists as values, and explains why keys must be hashable.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 11b Looping, Reverse Lookup, and Dictionaries of Lists

Title

Python · Chapter 11 — Dictionaries

§11.3-11.5, pp. 106-108

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-108 — the pages these objectives are drawn from

3. Before we start: going the wrong way

Warm-up

You have a histogram. Try asking it a question it was not built for.

Discussion prompt

You have a dictionary mapping letters to counts, and you want to know which letter appeared twice. There is no syntax for that. How would you find out, and what could go wrong with your method?

Hint: You know the search pattern from chapter 8.

Answer:

You would have to look at every item and check whether its value is 2 — a search, rather than a lookup.

Two things could go wrong. There might be more than one letter with a count of 2, in which case which letter is not a well-posed question. And there might be none, in which case there is no answer to return.

Both problems are named in this lesson, and the second one is what introduces a new statement: a way for your own code to raise an exception rather than returning something misleading.

4. The one idea behind this lesson: a mapping only runs one way

Concept

Given a dictionary d and a key k, it is easy to find the corresponding value v = d[k]. This operation is called a lookup. But what if you have v and you want to find k?

lookup — A dictionary operation that takes a key and finds the corresponding value.

You have two problems: first, there might be more than one key that maps to the value v — you might be able to pick one, or you might have to make a list of all of them. Second, there is no simple syntax to do a reverse lookup; you have to search.

Figure (svg): Two columns contrasting a forward lookup with a reverse lookup

The two directions are not symmetric, and every difference in the right column causes work.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-106

5. Looping over a dictionary

Section

Section 1

6. The loop variable holds a key

Concept

If you use a dictionary in a for statement, it traverses the keys of the dictionary. To see the values too, you look each one up inside the loop.

def print_hist(h):
    for c in h:
        print(c, h[c])

>>> h = histogram('parrot')
>>> print_hist(h)
a 1
p 1
r 2
t 1
o 1
PartWhat it doesNote
for c in hc takes each KEY in turnnot the value, not the pair
h[c]looks up the corresponding valuethe ordinary forward lookup
the output orderno particular orderas always

Again, the keys are in no particular order — this is the same unpredictability from the previous lesson, showing up in the place where it is most visible.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-106

7. Picture it: the loop walks the key column

Picture it

Each pass gives one key, and the value comes from a lookup.

Figure (svg): A pipeline showing a dictionary yielding one key per pass and a lookup producing the value

The loop supplies the key and the lookup supplies the value. Nothing gives you both at once.

That is a design choice worth noticing: looping gives keys because keys are what everything else in a dictionary is indexed by.

8. Worked example: traversing the keys in sorted order

Worked example

The dictionary has no order. sorted makes one.

>>> h = histogram('parrot')
>>> for key in sorted(h):
...     print(key, h[key])
a 1
o 1
p 1
r 2
t 1
PartWhat it producesNote
sorted(h)a sorted LIST of the keysthe dictionary is unchanged
for key in ...walks that listin a defined order
h[key]the value, by lookupas before

Recognise what sorted is given.

Why: A dictionary in a for statement yields keys, and sorted works the same way — it receives the keys and returns them sorted.

Note that it returns a list.

Why: sorted creates a new list, so the order lives in that list and not in the dictionary, which has no order to change.

Look up inside the loop as usual.

Why: The sorted list carries keys only, so the values still come from h[key].

Figure (svg): The state of the program after each line of Worked example traversing the keys in sorted order, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The five letters in alphabetical order with their counts. The ordering was created by sorted and does not belong to the dictionary.

Verify: Print the dictionary again afterwards.

Why: It comes back in the same unpredictable order it had before, because sorted returned a new list rather than rearranging anything — the returning-not-modifying distinction from lesson 10c. There is no operation that sorts a dictionary in place, and after the previous lesson it should be clear why the idea does not make sense.

9. Predict: what does the loop print?

Prediction

Watch what the loop variable holds.

h = {'a': 1, 'b': 2}
for x in h:
    print(x)
PartWhat happensOutput
for x in hx takes each key'a', then 'b'
print(x)prints the keynot the value
to get valuesprint(h[x])or loop over h.values()

Predict first

What does this print?

  • the keys: a and b
  • the values: 1 and 2
  • the pairs: ('a', 1) and ('b', 2)
  • an error, because x is not defined

Correct: the keys: a and b — a dictionary in a for statement traverses its keys.

Why: The name x has no influence on what the loop yields; a dictionary always gives its keys. To print the values you would write print(h[x]) inside the loop, or loop over h.values(). Note that the order of the two lines is not guaranteed, which is the previous lesson's point showing up again.

10. Worked example: two ways to get both parts

Worked example

One is written with what you know now, and one is the tool for the job.

# loop over keys and look each value up
for c in h:
    print(c, h[c])

# or ask for the pairs directly
for c, count in h.items():
    print(c, count)
FormWhat each pass givesNote
for c in hone key, then a lookuptwo steps
h.items()one key-value pair per passunpacked into two names
bothproduce the same outputthe second says the intent

Read the first form.

Why: This is what the book uses at this point: the loop gives the key and h[c] fetches the value.

Read the second.

Why: items yields the pairs, and the two names on the left of in unpack each pair — which is a shape you will meet properly in the next chapter.

Prefer the one that states the intent.

Why: If you need both parts of every item, saying so is clearer than fetching the second one yourself.

Figure (svg): Two columns comparing looping over keys with looping over items

Same result, and the second one states the intent.

Identical output. The first is the form built from what you already know; the second says for each key and its value, which is usually what you meant.

Verify: Ask what the first form does that the second does not.

Why: It performs a lookup per pass, which is fast but not free, and it works when you only sometimes need the value. Neither is wrong — but a loop that looks up every key it is given is asking for the pairs the long way round.

11. Trap: expecting the loop to give you values

Trap

The trap

A student writes for count in h: print(count) expecting to see the counts.

Name the loop variable after what you want

Why: The name is the only clue in the line, and naming it count makes it look like a count.

The loop gives keys whatever you call them, so this prints the letters. Nothing raises — the name is a label, not a request, and the program quietly reports the wrong column.

The fix

Name the loop variable after what it holds.

for c in h names a key

Why: And if you want the value, h[c] fetches it.

Or loop over h.values() if the keys are irrelevant

Why: Which says explicitly that you want the other column.

This is the same silent-wrong-answer failure as using in on values: a dictionary has two halves and the default is always the keys. When a program reports the wrong half, the first thing to check is which half it asked for.

12. Complete it: print keys in sorted order

Faded example

The dictionary has no order, so make one.

Fill in the blanks

for key in sorted(h):
print(key, h[key])

Why: sorted receives the keys and returns them as a new sorted list, which the loop then walks in a defined order. The dictionary itself is unchanged — there is no way to sort one, because it has no arrangement to sort. Note that h.sort() would raise an AttributeError: that method belongs to lists.

13. Discriminate: which half does this give you?

Discrimination

A dictionary has two columns and each expression picks one.

Sort into buckets

For each expression, which part of the dictionary does it produce?

works with the keys
for c in h; sorted(h); 'a' in h
works with the values
for v in h.values(); h[c]; max(h.values())
keys
Looping, sorting and membership all default to the keys, because keys are what a dictionary is indexed by. Nothing in any of these mentions values.
vals
Each names the values explicitly, either with the values method or by performing a lookup. Getting at the value column always takes an extra word.

14. Think it through: why does looping give keys?

Socratic

It could have given values, or pairs. It gives keys.

Discussion prompt

A dictionary could plausibly yield values or pairs when looped over. Why is yielding keys the sensible default?

Hint: What can you do with a key that you cannot do with a value?

Answer:

From a key you can reach everything: h[key] gives the value, so a loop over keys can produce anything a loop over pairs could.

From a value you can reach nothing. There is no way back to the key without a search, which is the entire subject of the next section.

So keys are the more useful thing to be given, because the mapping runs that way. Yielding the half you can navigate from is the same reasoning that makes the in operator check keys — both defaults point along the direction the structure is built for.

15. Reverse lookup

Section

Section 2

16. There is no syntax, so you search

Concept

Here is a function that takes a value and returns the first key that maps to that value. It is yet another example of the search pattern, but it uses a feature we haven't seen before.

def reverse_lookup(d, v):
    for k in d:
        if d[k] == v:
            return k
    raise LookupError()
LineWhat it doesNote
for k in dtraverse the keysthe only way in
if d[k] == vlook up and comparethe search pattern
return kthe FIRST key that matchesthere may be others
raise LookupError()reached only if nothing matchedno key to return

A reverse lookup is much slower than a forward lookup; if you have to do it often, or if the dictionary gets big, the performance of your program will suffer.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-107

17. Picture it: one operation against a whole traversal

Picture it

The forward direction touches one item. The reverse direction touches all of them.

Figure (svg): A growth chart comparing constant-time forward lookup against linear reverse lookup

The same two shapes as list search against dictionary lookup — because a reverse lookup IS a list-style search.

If a program does reverse lookups often, the fix is usually to build the inverse dictionary once, which is exactly what idea 4 does.

18. Worked example: a successful reverse lookup

Worked example

Find the letter that appeared twice.

>>> h = histogram('parrot')
>>> h
{'a': 1, 'p': 1, 'r': 2, 't': 1, 'o': 1}
>>> key = reverse_lookup(h, 2)
>>> key
'r'
PassThe comparisonWhat happens
k = 'a'h['a'] is 1, not 2keep going
k = 'p'h['p'] is 1, not 2keep going
k = 'r'h['r'] is 2return 'r'
the restnever examinedreturn leaves at once

Traverse the keys.

Why: The loop is the only way to reach the items, since there is no syntax that goes from a value to a key.

Compare each value.

Why: h[k] == v is a forward lookup used to test one item at a time, which is the search pattern from chapter 8.

Return on the first match.

Why: return leaves the function immediately, so the remaining items are never examined — which is why the answer is the first key found, not the only one.

Figure (svg): The state of the program after each line of Worked example a successful reverse lookup, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

'r'. It is the first key whose value is 2 in the order the dictionary happens to traverse, which for this histogram is the only one.

Verify: Ask what would happen with two matching keys.

Why: For histogram('parrots'), both 'r' and 's' map to 2, and the function returns whichever comes first in an unpredictable order. It would be correct and unrepeatable — which is the first of the two problems the book named, and it is the reason idea 4 collects all the matches rather than one.

19. Predict: which key comes back?

Prediction

Two keys map to the same value.

h = {'r': 2, 's': 2, 'a': 1}
print(reverse_lookup(h, 2))
AspectWhat is trueNote
the searchreturns on the first matchand stops
which is firstdepends on traversal orderunpredictable
the consequence'r' or 's', not bothand not reliably either

Predict first

What can you say about the result?

  • It is 'r', because 'r' was written first
  • It is 'r' or 's', and which one is not something you can rely on
  • It is a list containing both 'r' and 's'
  • It raises LookupError, because the answer is ambiguous

Correct: It is 'r' or 's', and which one is not something you can rely on — the function returns on the first match in an unpredictable traversal order.

Why: reverse_lookup returns the first key it finds with the given value, and the order of items in a dictionary is unpredictable, so the result is one of the two matches with no promise about which. The function raises only when nothing matches, and it never returns a list — collecting all the matches is what invert_dict does.

20. Worked example: an unsuccessful one

Worked example

No key maps to 3. The function has nothing to return.

>>> key = reverse_lookup(h, 3)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "<stdin>", line 5, in reverse_lookup
LookupError
StageWhat happensNote
the loopexamines every keynone has the value 3
falls off the endreaches the raiseno key to return
the tracebacktwo frames deepthe caller, then the function

Follow the loop to its end.

Why: Every key is tested and none matches, so control reaches the line after the loop.

See what that line does.

Why: If we get to the end of the loop, that means v doesn't appear in the dictionary as a value, so we raise an exception.

Read the traceback.

Why: It shows two frames — the call at the prompt and the function itself — which is chapter 5's stack diagram appearing in an error message.

Figure (svg): A panel showing the traceback from a failed reverse lookup with its two frames

A LookupError, raised deliberately by the function. The effect when you raise an exception is the same as when Python raises one: it prints a traceback and an error message.

Verify: Ask what the alternative would have been.

Why: Returning None. That would be legal and quiet, and the caller would carry a None around until it was used as a key or an index and failed somewhere else entirely — the NoneType failures from chapter 6. Raising here puts the error where the problem is, which is worth more than the convenience.

21. Trap: assuming the reverse lookup has one answer

Trap

The trap

A program calls reverse_lookup to find the word with a given frequency and uses the result as though it were unique.

Read the function's name as a lookup

Why: A forward lookup has exactly one answer, so the reverse of one sounds like it should too.

There might be more than one key that maps to the value, and the function returns whichever the traversal reaches first — in an order the language does not promise. The program is correct only by accident.

The fix

Decide what multiple matches should mean before you write the function.

If any one will do, say so

Why: Name the function first_key_with or similar, so the caller knows what they are getting.

If you need them all, collect them

Why: Build a list inside the loop and return it, which is what invert_dict does for every value at once.

The book names this as the first of the two problems with reverse lookup, and it is the more dangerous one — the second problem, the lack of syntax, merely costs you a loop.

22. Watch the search: reverse_lookup on a five-item dictionary

Invariant

Step through and watch how much work each case costs.

Step through it

Which case does the most work, and what does that say about using reverse_lookup in a loop?

  1. The first key's value is 1, which is not what we want, so the loop continues.
  2. The second is also 1. Two lookups done, no match yet.
  3. The third has value 2, so the function returns immediately and the last two keys are never examined.
  4. Searching for a value that is not there costs a full traversal every time — the worst case is also the case where you learn the least.

The failing search is the most expensive: it examines every item and returns nothing. So a program that reverse-looks-up many values, most of which are absent, does the maximum work every time — which is exactly the situation the book warns about when it says the performance of your program will suffer.

23. Rank: what reverse_lookup does, in order

Ranking

Four steps, in the order they happen.

Put in order

  1. take the next key from the dictionary
  2. look up the value for that key
  3. compare that value with the one being searched for
  4. raise LookupError, having run out of keys

Why: The loop takes a key, performs a forward lookup to get its value, compares, and either returns or continues. The raise is reached only after the loop has ended without returning — which is the point of putting it after the loop rather than inside it.

24. Explain it: why is one direction so much more expensive?

Explain it

The same dictionary, and two questions that cost very different amounts.

Discussion prompt

A classmate asks why d[k] is instant but finding the key for a value needs a loop. Explain it in terms of how a dictionary stores things.

Hint: Where does the dictionary put an item when you add it?

Answer:

When you add an item, the dictionary computes a location from the KEY and stores the pair there. So given a key it can compute where to look and go straight to it.

Nothing is computed from the value, so there is no location to compute when you start from one. The only option is to examine items until you find a match.

That is why the book says a reverse lookup is much slower than a forward one, and why the fix — when you need it often — is to build a second dictionary keyed the other way. You pay for the traversal once instead of every time.

25. The raise statement

Section

Section 3

26. Causing an exception on purpose

Concept

The raise statement causes an exception. In reverse_lookup it causes a LookupError, which is a built-in exception used to indicate that a lookup operation failed.

>>> raise LookupError()
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
LookupError

>>> raise LookupError('value does not appear in the dictionary')
Traceback (most recent call last):
  File "<stdin>", line 1, in ?
LookupError: value does not appear in the dictionary
StatementWhat it producesNote
raise LookupError()an exception with no messagethe name alone
raise LookupError('...')an exception with a messageshown after the colon
the effecta traceback and an error messageas if Python had raised it

The effect when you raise an exception is the same as when Python raises one. And when you raise an exception, you can provide a detailed error message as an optional argument — which is worth doing, because the person reading it will not have your function in front of them.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 107-107

27. Picture it: two sources, one mechanism

Picture it

An exception you raise and one Python raises behave identically.

Figure (svg): Two columns showing that a raised exception and a Python exception behave the same way

The mechanism is the same. Writing raise puts your function on the same footing as the language's own.

That is the point of the section: error reporting is not something only the interpreter can do.

28. Worked example: raising with a useful message

Worked example

The message is what the reader will actually see.

def reverse_lookup(d, v):
    for k in d:
        if d[k] == v:
            return k
    raise LookupError('value does not appear in the dictionary')
StageWhat happensNote
no match foundthe loop endscontrol reaches the raise
the messagepassed as an argumentprinted after the colon
the readersees what went wrongwithout reading the source

Put the raise after the loop.

Why: It is reached only when the loop finished without returning, which is exactly the case where there is no key to give back.

Supply a message.

Why: You can provide a detailed error message as an optional argument, and here it says which of the many possible lookup failures this is.

Consider what to include.

Why: Naming the value that was not found — as a KeyError names its key — makes the message diagnostic rather than merely descriptive.

Figure (svg): The state of the program after each line of Worked example raising with a useful message, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A function that fails clearly. Compare the two forms in the traceback: bare LookupError versus LookupError with a sentence, and only the second one tells a reader anything.

Verify: Read the traceback as someone who did not write the function.

Why: With the bare form they see a name they may not recognise and a line number in code they did not write. With the message they see what the function was trying to do and why it could not — which is the difference between an error they can act on and one they have to investigate.

29. Predict: what does the message look like?

Prediction

The optional argument is the part a reader sees.

raise LookupError('value does not appear in the dictionary')
PartWhat it isWhere it appears
the exception typeLookupErrorbefore the colon
the argumentthe detailed messageafter the colon
the effecttraceback plus messageas if Python raised it

Predict first

What is printed on the last line of the traceback?

  • LookupError: value does not appear in the dictionary
  • LookupError
  • value does not appear in the dictionary
  • Nothing — a raised exception prints no message

Correct: LookupError: value does not appear in the dictionary — the type, a colon, and the message you supplied.

Why: The message is an optional argument, and when you supply one it appears after the exception name and a colon. That is the same shape as every error message you have read so far — KeyError: 'four', TypeError: ... — because it is the same mechanism, whoever raised it.

30. Worked example: raise or return None?

Worked example

Two ways to report failure, and they fail in different places.

# raises: fails at the point of the problem
def lookup_or_raise(d, v):
    for k in d:
        if d[k] == v:
            return k
    raise LookupError('no such value')

# returns None: fails later, somewhere else
def lookup_or_none(d, v):
    for k in d:
        if d[k] == v:
            return k
    return None
VersionWhat happens on failureWhere the error appears
the raising versionstops at the failurethe traceback points here
the None versionreturns quietlythe caller carries None away
where the None version failswherever the None is usedpossibly far away

Consider what the caller does with a None.

Why: It usually gets used — as a key, an index, or in arithmetic — and fails there with a NoneType message that names nothing about the real problem.

Consider what the caller does with an exception.

Why: It stops, with a traceback whose deepest frame is the function that could not do its job.

Choose by whether absence is expected.

Why: If the caller routinely asks about values that may not be there, returning None is reasonable and the caller must check. If absence means something is wrong, raise.

Figure (svg): A flowchart comparing where a raising function and a None-returning function report a failure

Both report the failure. They differ in how far the report is from the cause.

Raising when failure is exceptional, returning None when it is routine and the caller will check. The cost of the None version is the distance between where it fails and where the problem is.

Verify: Trace the None version's failure to its cause.

Why: The traceback names a line that uses the result, several steps from the lookup that returned nothing — the exact gap chapter 10 warned about with NoneType errors. That distance is the price, and it is why the book's version raises.

31. Trap: putting the raise inside the loop

Trap

The trap

A student writes the raise inside the for loop, in an else clause of the if.

Handle both outcomes of the comparison

Why: Every if wants an else, so the failure case goes there.

Now the function raises on the very first key that does not match, which is almost always the first key. It reports failure before it has finished looking, so a value that IS in the dictionary is reported as absent.

The fix

The raise belongs after the loop.

Ask when you actually know the search has failed

Why: Only when every key has been checked — which is when the loop ends.

Read the indentation as the answer

Why: At the loop's indentation the raise runs after all passes; inside, it runs during one.

This is the search pattern's standard shape, and it is the same reason a found-flag is initialised before a loop rather than inside it: a conclusion about all the items cannot be drawn from one of them.

32. Two truths and a lie: raise

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. A raised exception prints a traceback, just like one Python raises
  • B. You can provide a detailed error message as an optional argument
  • C. raise returns a value to the caller, so the caller can check it

Survives elimination: C

Why: C confuses raising with returning. raise does not return anything — it stops the function and unwinds to whoever is prepared to handle the exception, printing a traceback if nobody is. Returning a value the caller checks is the alternative design, and choosing between them is exactly the raise-or-return-None decision.

33. Complete it: fail with a message

Faded example

The bare exception name is legal. A message is better.

Fill in the blanks

def reverse_lookup(d, v):
for k in d:
if d[k] == v:
return k
raise LookupError('value does not appear in the dictionary')

Why: The raise goes after the loop, because only then do you know that no key matched. LookupError is the built-in exception used to indicate that a lookup operation failed, and the string argument becomes the message a reader sees after the colon.

34. Where *report the failure where it happens* applies

Real world

The raise-or-return-None choice is a general one about error reporting.

Discussion prompt

Think of a system that detects a problem and does not report it until much later. What does the delay cost the person diagnosing it?

Hint: Anything with a form, a queue, or a batch process.

Answer:

A form that accepts a bad value and fails at submission; a batch job that reads a malformed file at midnight and reports at nine; a build that compiles a bad configuration and fails at deploy.

What the delay costs is the connection between the symptom and the cause. By the time the failure appears, the information that would explain it — which field, which line, which setting — is gone.

Raising at the point of failure is the programming version of validating at the point of entry. It is more annoying in the moment and much cheaper afterwards, which is the trade in every case.

35. Dictionaries and lists: inverting a mapping

Section

Section 4

36. Lists can be values

Concept

Lists can appear as values in a dictionary. If you have a dictionary that maps from letters to frequencies, you might want to invert it — but since there might be several letters with the same frequency, each value in the inverted dictionary should be a list of letters.

singleton — A list with a single element.

def invert_dict(d):
    inverse = dict()
    for key in d:
        val = d[key]
        if val not in inverse:
            inverse[val] = [key]
        else:
            inverse[val].append(key)
    return inverse
PassThe testThe inverse after
key = 'a', val = 11 not in inverseinverse[1] = ['a']
key = 'p', val = 11 is in inverseappend: [1] becomes ['a', 'p']
key = 'r', val = 22 not in inverseinverse[2] = ['r']
key = 't', val = 11 is in inverse['a', 'p', 't']

Each time through the loop, key gets a key from d and val gets the corresponding value. If val is not in inverse, we haven't seen it before, so we create a new item and initialize it with a singleton. Otherwise we append the corresponding key to the list.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 107-108

37. Picture it: figure 11.1, the state diagram

Picture it

The book draws integers and strings inside the box and lists outside it, to keep the diagram simple.

Figure (svg): A state diagram showing the histogram dictionary and its inverse with list values drawn outside the box

The book's figure 11.1. A dictionary is a box with the type dict above it and the pairs inside.

Every list is drawn outside the box because it is a separate object the dictionary refers to — the same reference idea as chapter 10, in a new picture.

38. Worked example: inverting a histogram

Worked example

Letters to counts becomes counts to lists of letters.

>>> hist = histogram('parrot')
>>> hist
{'a': 1, 'p': 1, 'r': 2, 't': 1, 'o': 1}
>>> inverse = invert_dict(hist)
>>> inverse
{1: ['a', 'p', 't', 'o'], 2: ['r']}
AspectWhat happenedNote
the keys become valuesletters end up in listsseveral per count
the values become keyscounts are now the indexand they are integers
the sizes differfive items become twobecause counts repeat

Notice the shape changed.

Why: The original has one letter per item; the inverse has one count per item and several letters inside each.

Notice why it had to.

Why: Four letters share the count 1, and each key is associated with a single value — so the four letters have to go into one value, which means a list.

Notice the keys are now integers.

Why: A dictionary's keys can be almost any type, and here they are counts rather than characters, which is the generality from the previous lesson being used for real.

Figure (svg): The state of the program after each line of Worked example inverting a histogram, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

{1: ['a', 'p', 't', 'o'], 2: ['r']}. Five items become two, because the original values repeat and the inverse collects the duplicates rather than losing them.

Verify: Check that no information was lost.

Why: Counting the letters across all the lists gives five, which is the number of items in the original. That check is the point of using lists: a naive inversion writing inverse[val] = key would have kept one letter per count and silently discarded three.

39. Predict: what does the inverse look like?

Prediction

Two keys share a value.

d = {'a': 1, 'b': 2, 'c': 1}
print(invert_dict(d))
PassThe testThe inverse after
key 'a', val 11 is new{1: ['a']}
key 'b', val 22 is new{1: ['a'], 2: ['b']}
key 'c', val 11 is known{1: ['a', 'c'], 2: ['b']}

Predict first

What does this print?

  • {1: ['a', 'c'], 2: ['b']}
  • {1: 'c', 2: 'b'}
  • {1: ['a'], 2: ['b'], 1: ['c']}
  • {'a': 1, 'b': 2, 'c': 1}

Correct: {1: ['a', 'c'], 2: ['b']} — the two keys sharing the value 1 both end up in one list.

Why: Three items become two, because 'a' and 'c' share a value and the inverse collects them rather than losing one. The second option is what a naive inversion would produce, discarding 'a'. The third is impossible — a dictionary cannot show the same key twice, since each key is associated with a single value.

40. Worked example: what the singleton is for

Worked example

The first sighting has to create something an append can add to.

if val not in inverse:
    inverse[val] = [key]      # a singleton, not key
else:
    inverse[val].append(key)
CaseWhat is storedWhy it matters
first sightinginverse[val] = [key]a list holding one thing
later sightingsinverse[val].append(key)the list grows
if it were = keya string, not a listappend would raise

Look at the brackets on the first branch.

Why: inverse[val] = [key] stores a list containing the key, not the key itself. That list is a singleton — a list that contains a single element.

Ask what the else branch needs.

Why: It calls append on whatever is there, and append is a list method. So the first branch must leave a list behind.

Predict the failure without the brackets.

Why: Storing the key directly leaves a string, and 'a'.append('p') raises an AttributeError, because strings have no append.

Figure (svg): A flowchart showing the create-a-singleton branch and the append branch of invert_dict

The same two-branch shape as histogram, with a list where the counter was.

The singleton exists so that every value in the inverse is a list from the start, which lets the else branch treat every later sighting identically.

Verify: Compare with the get trick from the previous lesson.

Why: The same structure appears: a first sighting that has nothing to add to. Here it could be written inverse.setdefault(val, []).append(key) or inverse[val] = inverse.get(val, []) + [key] — both remove the conditional the same way get did for the histogram, by making the first sighting produce an empty starting value.

41. Trap: storing the key instead of a singleton

Trap

The trap

A student writes inverse[val] = key in the first branch, since only one key has been seen.

Store what you have

Why: One key has been found, so storing one key looks right.

The value is now a string rather than a list, so the first time a second key shares that value, inverse[val].append(key) raises AttributeError — a failure that appears only on data with a repeat.

The fix

Store a list containing the key.

inverse[val] = [key]

Why: A singleton: a list with a single element, ready to grow.

Keep the value type uniform

Why: Every value in the inverse is a list, whether it holds one key or four — so nothing has to check.

The general rule is worth naming: when a value may need to hold several things, make it a collection from the beginning. Special-casing the first one produces a structure whose values have two possible types, and every reader afterwards has to handle both.

42. Complete it: the first-sighting branch

Faded example

What has to be stored so that append works later?

Fill in the blanks

if val not in inverse:
inverse[val] = [key]
else:
inverse[val].append(key)

Why: A singleton — a list containing the one key seen so far. Storing key alone would leave a string in the dictionary, and the else branch's append would raise AttributeError as soon as a second key shared that value. Making the value a list from the start keeps every value the same type.

43. Compare: histogram and invert_dict

Comparison

Fill the blanks. The same shape, with a different thing accumulated.

Comparison matrix

Questionhistograminvert_dict
What are the keys?charactersthe original dictionary's values
What is accumulated?a counta list of keys
First sighting stores1[key], a singleton
Later sightings do what?d[c] += 1inverse[val].append(key)

Both are the same pattern: for each distinct thing, accumulate a fact about it. Only the fact differs.

44. Explain it yourself: why must the inverse hold lists?

Explain it to yourself

The original held single values. The inverse cannot.

Discussion prompt

Explain why inverting a dictionary needs lists as values, using the fact that each key maps to exactly one value.

Hint: What happens when two keys have the same value?

Answer:

In the original, several keys may map to the same value — four letters can all have the count 1, and nothing forbids that.

When you turn it around, that count becomes a key, and a key is associated with a single value. Four letters cannot each be that single value.

So the single value has to be one thing that contains four — a list. The constraint is not about lists at all; it is that keys are unique and values are not, so inverting has to collapse the many into one collection.

45. Why keys must be hashable

Section

Section 5

46. Lists can be values, but not keys

Concept

Lists can be values in a dictionary, as invert_dict shows, but they cannot be keys. Here's what happens if you try.

hash function — A function used by a hashtable to compute the location for a key.

>>> t = [1, 2, 3]
>>> d = dict()
>>> d[t] = 'oops'
Traceback (most recent call last):
  File "<stdin>", line 1, in ?
TypeError: list objects are unhashable
LineWhat happensNote
d[t] = 'oops't is a listTypeError
the word unhashablenames the actual requirementkeys must be hashable
whybecause of how dictionaries store itemsthe next few lines

A hash is a function that takes a value of any kind and returns an integer. Dictionaries use these integers, called hash values, to store and look up key-value pairs.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 108-108

47. Picture it: what a mutable key would do

Picture it

Hash the key, store it there. Then change the key.

Figure (svg): A flowchart showing an item stored at a computed location and then becoming unfindable after the key changes

The item did not move. The place the dictionary looks for it did.

So the dictionary would have an item it cannot find, or two entries for what is now the same key. Either way, it would not work correctly.

48. Worked example: the argument in full

Worked example

The book's reasoning, step by step.

# if this were allowed:
t = [1, 2]
d = {}
d[t] = 'x'      # hashed and filed under hash([1, 2])
t.append(3)     # the key is now [1, 2, 3]
d[t]            # looks under hash([1, 2, 3]) - nothing there
StageWhat is computedConsequence
storelocation computed from [1, 2]the pair goes there
modifythe key becomes [1, 2, 3]the stored pair does not move
look uplocation computed from [1, 2, 3]a different place

Store the pair.

Why: When you create a key-value pair, Python hashes the key and stores it in the corresponding location.

Modify the key.

Why: This is only possible because the key is mutable — a string or an integer could not be changed like this.

Look it up.

Why: If you modify the key and then hash it again, it would go to a different location. In that case you might have two entries for the same key, or you might not be able to find a key.

Figure (svg): The state of the program after each line of Worked example the argument in full, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The dictionary would not work correctly. That's why keys have to be hashable, and why mutable types like lists aren't.

Verify: Check that immutable keys cannot cause this.

Why: A string key cannot be modified at all, so its hash value is fixed for as long as it exists and the item can always be found where it was filed. The restriction is not arbitrary: it is exactly what is needed to make the storage scheme work, and it is the payoff of chapter 10's mutability distinction.

49. Predict: which line raises?

Prediction

A list appears twice, in two different positions.

t = [1, 2]
d = {}
d['nums'] = t
d[t] = 'nums'
LineWhat it doesResult
d['nums'] = ta list as a VALUEfine
d[t] = 'nums'a list as a KEYTypeError
the reasonkeys are hashedvalues are not

Predict first

Which line raises an error?

  • Line 3, because a list cannot be stored in a dictionary
  • Line 4, because a list cannot be a key
  • Both lines 3 and 4
  • Neither — both are legal

Correct: Line 4, because a list cannot be a key — keys must be hashable and lists are not.

Why: Line 3 is fine: lists can appear as values in a dictionary, which is exactly what invert_dict relies on. Line 4 raises TypeError: list objects are unhashable, because the key is what gets hashed to decide where the pair is stored. The restriction lands on one side of the colon only.

50. Worked example: what you can use as a key

Worked example

Immutable things, and the way round the restriction.

d = {}
d['abc'] = 1        # a string: fine
d[42] = 2           # an integer: fine
d[3.14] = 3         # a float: fine
d[(1, 2)] = 4       # a tuple: fine, and we meet these next chapter
# d[[1, 2]] = 5     # a list: TypeError, unhashable
TypeMutable?Usable as a key?
strings, numbersimmutablehashable
tuplesimmutable sequenceshashable, and the way round the restriction
lists, dictionariesmutableunhashable

Notice the pattern.

Why: Everything immutable you have met can be a key, and the mutable types cannot. Hashability tracks mutability almost exactly.

Notice what values may be.

Why: Anything at all. invert_dict put lists in as values, and the restriction has never applied to that side.

Note the escape route.

Why: The simplest way to get around this limitation is to use tuples, which we will see in the next chapter.

Figure (svg): Two columns contrasting what may be a key with what may be a value

Only the key decides where the pair is filed, so only the key is restricted.

Immutable things make keys; anything makes a value. The tuple is the immutable sequence, and it exists partly so that a sequence can be a key.

Verify: Ask why the restriction is one-sided.

Why: Because only the key determines where the pair is stored. A value can change without moving anything, so there is no reason to restrict it — which is why the same list that is illegal on the left of the colon is perfectly legal on the right.

51. Trap: reading *unhashable* as a problem with lists

Trap

The trap

A student concludes that lists cannot be used with dictionaries at all.

Generalise from the error

Why: TypeError: list objects are unhashable does sound like a blanket ban.

Lists are perfectly good values — invert_dict is built on that — and the restriction applies only to keys. Avoiding lists entirely rules out the chapter's own main example.

The fix

Read the error as being about the position, not the type.

Keys are hashed; values are not

Why: So the requirement lands on one side of the colon only.

If you need a sequence as a key, use a tuple

Why: The book points forward to this explicitly.

The precise statement is that keys have to be hashable, and mutable types like lists aren't. Everything else about lists in dictionaries is unaffected.

52. Sort: can this be a dictionary key?

Sorting

The question is whether it can change.

Sort into buckets

For each value, can it be used as a dictionary key?

hashable: can be a key
'banana'; 42; 3.14
unhashable: cannot be a key
[1, 2, 3]; {'a': 1}; a list of two strings
yes
Each is immutable, so its hash value cannot change while it is in use and the pair can always be found where it was filed. Strings, integers and floats are all safe keys.
no
Each is mutable, so modifying it after it was filed would send a later lookup to a different location. A list is the book's example and a dictionary has the same problem.

53. Error analysis: four dictionary lines

Error analysis

Mark each and say what happens.

Annotate

  • Line 1 uses a list as a key and raises TypeError: list objects are unhashable. The key is what gets hashed, and a mutable key could move after it was filed.
  • Line 2 uses a list as a value, which is legal and is the basis of invert_dict. The restriction applies to keys only.
  • Line 3 stores the key itself rather than a singleton. Legal, and silently wrong: the value is now a string rather than a list.
  • Line 4 is correct when the value is already a list, and raises AttributeError when line 3 put a string there — so the mistake in line 3 surfaces here, on the first repeated value.
  • So the pairing matters: lines 3 and 4 are individually plausible and jointly broken, and the failure appears only on data where two keys share a value.
  • The correct first branch is inverse[val] = [key], which leaves a list for line 4 to append to.

The gap between the mistake in line 3 and the failure in line 4 is the same cause-and-symptom distance the whole course keeps meeting.

54. Think it through: why is hashability about mutability?

Socratic

The restriction sounds technical. It is a direct consequence of how storage works.

Discussion prompt

Explain, without using the word hashable, why a key that can change would break a dictionary.

Hint: Where does the dictionary decide to put the item?

Answer:

The dictionary decides where to put an item by computing something from the key. So the location is a fact about the key's contents at the moment it was stored.

If the contents then change, the computation gives a different answer, and the dictionary looks in a place where the item never was. The item has not moved; the address has.

So the requirement is really that a key's contents be fixed for as long as it is in use — which is precisely what immutability means. The restriction is not an implementation quirk, it is the minimum condition for the storage scheme to work at all.

55. Compare: forward and reverse lookup

Comparison

Fill the blanks. One is an operation and the other is a search.

Comparison matrix

QuestionForward: d[k]Reverse: value to key
Is there syntax for it?yes, bracket lookupno — you write a loop
How many answers?exactly onezero, one, or many
Costabout constantgrows with the dictionary
What happens on failure?KeyError, automaticallywhatever you write — the book raises LookupError

Every row is a consequence of the same fact: the dictionary computes its storage location from the key and never from the value.

56. The procedure: accumulating a list per key

Pattern

Six steps. It is the histogram pattern with a list where the counter was.

  1. Decide what the key is, and what should be collected under it.
  2. Create an empty dictionary to accumulate into.
  3. Loop over the source data, computing the key for each item.
  4. If the key is new, store a singleton — a list containing this one item.
  5. Otherwise, append to the list that is already there.
  6. Return the accumulator, whose every value is a list — including the ones holding a single element.

Step 4's brackets are the whole pattern. Storing the item rather than a list containing it leaves values of two different types, and the else branch fails on the first repeat.

Python documentation — Data Structures Data Structures

57. Check yourself 1 of 3: what the loop yields

Check

A dictionary in a for statement.

h = {'a': 1, 'b': 2}
total = 0
for x in h:
    total += h[x]
print(total)
PassThe lookupRunning total
x = 'a'h['a'] is 1total is 1
x = 'b'h['b'] is 2total is 3
printthe sum of the values3

Check your understanding

What does this print?

  • A. 3 (correct)
  • B. 0
  • C. An error, because you cannot add a string to a number
  • D. 2

Answer: A

Why: The loop gives keys, and h[x] converts each key to its value, so the additions are 1 and 2 and the total is 3. Writing total += x instead would try to add a string to an integer and raise the TypeError described in option C — which is the difference the lookup inside the loop is there to make.

Why B tempts people
The loop runs twice and adds a positive value each time, so the total cannot stay at zero.
Why C tempts people
This would be right if the loop variable were added directly. The lookup converts each key to a number first.
Why D tempts people
This would be the last value rather than the sum. += accumulates across passes.

58. Check yourself 2 of 3: raise

Check

The search ends without a match.

Check your understanding

In reverse_lookup, why does the raise statement come after the loop rather than inside it?

  • A. Because raise is not allowed inside a loop
  • B. Because you only know the search has failed once every key has been checked (correct)
  • C. Because the loop would catch the exception
  • D. Because indentation inside a loop is not allowed after an if

Answer: B

Why: A conclusion about all the items cannot be drawn from one of them. Inside the loop, the raise would fire on the first key that does not match — almost always the first key — and report a value as absent when it is present two items later. Placing it after the loop means it is reached only when the loop finished without returning.

Why A tempts people
raise is an ordinary statement and is legal anywhere. Where it is placed is a logic decision, not a syntax rule.
Why C tempts people
Loops do not catch exceptions. Nothing here catches anything — the exception propagates to the caller and prints a traceback.
Why D tempts people
Indentation after an if inside a loop is perfectly normal; the else branch of invert_dict does exactly that.

59. Check yourself 3 of 3: keys and values

Check

One of these positions is restricted.

Check your understanding

Which statement is true about lists and dictionaries?

  • A. Lists cannot be used with dictionaries at all
  • B. Lists can be values but not keys, because keys have to be hashable (correct)
  • C. Lists can be keys but not values
  • D. Lists can be keys as long as you do not modify them

Answer: B

Why: Lists can appear as values — invert_dict depends on it — and cannot appear as keys, because the key is what gets hashed to decide where the pair is stored and a mutable key could move afterwards. The simplest way round the restriction is a tuple, which the next chapter introduces.

Why A tempts people
Too broad. The chapter's own main example puts lists into a dictionary as values.
Why C tempts people
Exactly backwards. The restriction is on the key, because only the key determines storage location.
Why D tempts people
Python cannot enforce a promise not to modify something, so the type is rejected outright rather than trusted.

60. Where this shows up outside this course

Real world

Building an index is the everyday version of inverting a dictionary.

Discussion prompt

The index at the back of a book maps topics to page numbers, and the book itself maps page numbers to content. What work did someone do to produce the index, and why is each entry a list?

Hint: They read it once and wrote down what they found.

Answer:

Someone traversed the whole book once, and for each topic they encountered, added the page number to that topic's list. That is invert_dict, done by hand.

Each entry is a list because a topic appears on several pages — the same reason the inverse of a histogram holds lists. Several keys mapping to one value become one key mapping to several.

And the payoff is the same one the book names: a reverse lookup is much slower than a forward lookup, so if you need it often, you traverse once and build the inverse. That is what an index is for, and it is why nobody reads a whole book to find a topic.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

You write d[[1, 2]] = 'x'. What happens, and why?

  • It works — lists can be keys
  • TypeError, because keys must be hashable and lists are mutable
  • KeyError, because the list is not already a key
  • It works, but only until the list is modified

Correct: TypeError, because keys must be hashable and lists are mutable — the error message is 'list objects are unhashable'.

Why: The reason is worth being able to state. A dictionary computes a location from the key and stores the pair there. If the key could change, hashing it again would give a different location, so you might end up with two entries for the same key or be unable to find one — either way, the dictionary wouldn't work correctly. Python cannot trust a promise not to modify a list, so it rejects mutable keys outright rather than allowing the fourth option. The same list is perfectly legal as a value, because nothing is computed from values.

62. Explain it to someone else

Explain it

Two directions, and only one of them is cheap.

Discussion prompt

A classmate wants to find which key maps to a given value and asks why there is no d.reverse(). Explain what they would have to write and why the language does not provide it.

Hint: Ask what the answer would even be.

Answer:

They would have to loop over every key, look up its value, and compare — the search pattern, and the whole dictionary examined in the worst case.

The language does not provide it partly because the answer is not well defined: several keys can map to one value, so the key does not exist. Any built-in would have to choose, and no choice is right for everyone.

And partly because it would hide the cost. d[k] looks cheap and is cheap; a built-in reverse lookup would look equally cheap and traverse everything. Making it a visible loop keeps the price visible — and if they need it often, the answer is invert_dict, paid for once.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • Looping over a dictionary, and getting the keys in sorted order
  • Reverse lookup: the search, and its two problems
  • The raise statement, and choosing between raising and returning None
  • Inverting a dictionary, and why keys must be hashable

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The looping is mechanical once you have internalised that the loop yields keys, which is the one thing people get wrong. Reverse lookup is worth being uncomfortable about, because its two problems — several answers, and the cost — are the reason the rest of the chapter exists. The raise statement is your first time causing an error rather than receiving one, and the design question behind it recurs everywhere. And the hashability argument is the payoff of chapter 10: it is the moment mutability stops being a curiosity and starts constraining what you can write.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw a dictionary and, beside it, two arrows: one from a key to its value labelled one operation, and one from a value back to a key labelled a search over everything. Under the second, write the two problems the book names. Then write invert_dict from memory, circling the singleton, and beneath it write in one sentence why a list can be on the right of the colon and not on the left.

65. What you can do now

Recap

Three pages, and the asymmetry that shapes every use of a mapping.

If you remember one thingIt is this
From loopingA dictionary yields keys, whatever you name the loop variable.
From reverse lookupIt is a search, not a lookup — and it may have zero or many answers.
From raiseReport the failure where it happens, not by returning None for someone else to trip over.
From invert_dictWhen a value may hold several things, make it a list from the first one.
From hashabilityOnly keys are hashed, so only keys are restricted.

The next lesson uses a dictionary as machinery rather than as data: memos, which turn an impossibly slow recursive Fibonacci into an instant one, and the global statement — plus the debugging advice for when your data are too big to print.

Think Python, 2nd edition — Allen B. Downey §11.3-11.5, pp. 106-108 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §11.3-11.5, pp. 106-108
  2. Python documentation — Data Structures
  3. Python documentation — More Control Flow Tools

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108