This lesson introduces the dictionary as a mapping from keys to values, covers creation, lookup, KeyError, len and the in operator, explains why in is fast for dictionaries and slow for lists, and builds the histogram function two ways.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 11 — Dictionaries
§11.1-11.2, pp. 103-105
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 103-105 — the pages these objectives are drawn from
Warm-up
You know lists. Try the problem the next section solves.
Discussion prompt
You are given a string and want to know how many times each letter appears. Sketch how you would do it with the tools from chapter 10 — a list, a loop, and indices. What is awkward about it?
Hint: How do you get from the letter 'q' to a position in a list?
Answer:
You could make a list of 26 counters and convert each character to a number with ord, using that number as an index. It works.
What is awkward is the conversion: the data are letters and the structure is indexed by integers, so every access needs a translation step.
And you have to make room for all 26 whether or not they appear. This lesson's structure is indexed by whatever you like — including letters — which removes both problems at once.
Concept
A dictionary is like a list, but more general. In a list, the indices have to be integers; in a dictionary they can be (almost) any type.
dictionary — A mapping from keys to their corresponding values.
A dictionary contains a collection of indices, which are called keys, and a collection of values. Each key is associated with a single value, and the association of a key and a value is called a key-value pair or an item.
Figure (svg): Two columns comparing how a list and a dictionary are indexed
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 103-103
Section
Section 1
Concept
The function dict creates a new dictionary with no items. Because dict is the name of a built-in function, you should avoid using it as a variable name.
>>> eng2sp = dict()
>>> eng2sp
{}
>>> eng2sp['one'] = 'uno'
>>> eng2sp
{'one': 'uno'}
>>> eng2sp = {'one': 'uno', 'two': 'dos', 'three': 'tres'}| Line | What happens | Note |
|---|---|---|
| dict() | creates an empty dictionary | printed as {} |
| eng2sp['one'] = 'uno' | creates an item | key 'one' maps to 'uno' |
| the brace literal | three items at once | the output format is an input format |
The squiggly-brackets represent an empty dictionary, and a printed dictionary shows each pair with a colon between key and value. That output format is also an input format, so you can copy what Python prints straight back into your code.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 103-104
Picture it
Three items, each a key on the left and a value on the right.
Figure (svg): A dictionary diagram showing three string keys each mapping to a Spanish word
Nothing in the picture says which pair is first, and nothing in the language does either. That is the next section.
Worked example
Start empty and add. The bracket syntax does both jobs.
>>> eng2sp = dict()
>>> eng2sp['one'] = 'uno'
>>> eng2sp['two'] = 'dos'
>>> eng2sp
{'one': 'uno', 'two': 'dos'}
>>> len(eng2sp)
2| Line | What happens | Size after |
|---|---|---|
| dict() | an empty dictionary | 0 items |
| eng2sp['one'] = 'uno' | a new item | 1 item |
| eng2sp['two'] = 'dos' | another new item | 2 items |
Create the empty dictionary.
Why: dict() gives you something with no items, printed as {} — the counterpart of [] for a list.
Add with bracket assignment.
Why: The same syntax that modified a list element creates a new item here, because there was no position to overwrite.
Check the size.
Why: len works on dictionaries and returns the number of key-value pairs.
Figure (svg): The state of the program after each line of Worked example building a dictionary one item at a time, drawn as a ladder with one rung per traced line
A dictionary with two items. Bracket assignment adds when the key is new and replaces when it is not, which is a genuine difference from lists.
Verify: Assign to a key that is already there.
Why: eng2sp['one'] = 'ONE' leaves the length at 2 and changes the value — so the same syntax adds or replaces depending on whether the key exists. A list cannot add this way at all: t[5] = x on a three-element list raises an IndexError.
Prediction
One of the assignments is not what it looks like.
d = dict()
d['a'] = 1
d['b'] = 2
d['a'] = 3
print(len(d))| Line | What happens | Size after |
|---|---|---|
| d['a'] = 1 | a new item | 1 item |
| d['b'] = 2 | a new item | 2 items |
| d['a'] = 3 | the key exists, so it replaces | still 2 items |
Predict first
What does this print?
Correct: 2 — the last line replaces the value for an existing key rather than adding a new item.
Why: Each key is associated with a single value, so assigning to 'a' a second time overwrites the first association. The dictionary holds two items, 'a' mapping to 3 and 'b' mapping to 2. This is what makes bracket assignment do double duty: it adds when the key is new and replaces when it is not.
Worked example
Looking up a missing key is not silent.
>>> eng2sp = {'one': 'uno', 'two': 'dos', 'three': 'tres'}
>>> eng2sp['two']
'dos'
>>> eng2sp['four']
KeyError: 'four'| Expression | Situation | Result |
|---|---|---|
| eng2sp['two'] | the key is present | 'dos' |
| eng2sp['four'] | the key is absent | KeyError |
| the exception name | names the missing key | 'four' |
Look up a key that is present.
Why: You use the keys to look up the corresponding values, and 'two' always maps to 'dos' however the items are ordered.
Look up one that is not.
Why: If the key isn't in the dictionary, you get an exception — a KeyError, which names the key it could not find.
Note how much the message tells you.
Why: The exception carries the actual key, which is usually enough to diagnose the bug without adding a print.
Figure (svg): A panel showing a KeyError raised by looking up an absent key
A present key gives its value; an absent one raises KeyError naming the key. Compare with a list, where an out-of-range index gives IndexError.
Verify: Check the parallel with lists.
Why: The two errors are the same kind of failure — an index that does not exist — reported under different names because the index is a different sort of thing. Both are loud, which is a mercy after chapter 10's silent list bugs.
Trap
A student writes dict = {'a': 1} because the name describes exactly what the variable holds.
Name the variable after what it is
Why: Which is normally good advice, and here the obvious name is already taken.
The name dict now refers to a dictionary rather than to the built-in function, so any later dict() call fails with a TypeError saying a dict object is not callable — possibly far from the assignment.
Avoid using built-in names as variable names.
Name it after the mapping
Why: eng2sp, counts, histogram — each says what maps to what, which is more informative than dict anyway.
Recognise the symptom
Why: A TypeError saying an object is not callable, on a line that looks correct, usually means a built-in name was shadowed earlier.
The book gives this warning at the moment it introduces dict, and the same applies to list, str, len, sum and type — all of which are ordinary names that can be reassigned.
Faded example
The syntax is the same one that modified a list.
Fill in the blanks
eng2sp = dict()
eng2sp['four'] = 'cuatro'
print(len(eng2sp)) # 1
Why: Bracket assignment with the key inside creates the item mapping 'four' to 'cuatro'. Note the key must be in quotes — writing eng2sp[four] would look for a variable called four and raise a NameError, which is the commonest slip when the keys are strings.
Error analysis
Mark each and name the error it raises.
Annotate
The colon is what distinguishes a dictionary literal from a set literal, and it is the only thing that does.
Explain it to yourself
Python prints a dictionary in a form you can paste back.
Discussion prompt
Python prints {'one': 'uno', 'two': 'dos'} and you can type exactly that to create the dictionary. Why is that convenient, and where else have you seen it?
Hint: What do you do when a program produces something you want to test with?
Answer:
It means you can copy a value out of the interpreter and into your code without translating it, which makes interactive experimentation into a way of writing programs.
Lists do the same — Python prints [1, 2, 3] and you can type [1, 2, 3] — and so do strings, floats and booleans. Almost everything you have met prints in its own source form.
The exception is worth noticing: a function prints as something like <function f at 0x...>, which you cannot type back. That is the signature of a value whose printed form is a description rather than a literal.
Section
Section 2
Concept
If you print a dictionary you might be surprised: the order of the key-value pairs might not be the order you typed them, and if you type the same example on your computer you might get a different result.
>>> eng2sp = {'one': 'uno', 'two': 'dos', 'three': 'tres'}
>>> eng2sp
{'one': 'uno', 'three': 'tres', 'two': 'dos'}
>>> eng2sp['two']
'dos'| Observation | What is true | Note |
|---|---|---|
| the order printed | not the order typed | and unpredictable in general |
| eng2sp['two'] | still finds the value | the key is what matters |
| why it is not a problem | no integer indices are used | so order is never consulted |
In general, the order of items in a dictionary is unpredictable. But that's not a problem, because the elements of a dictionary are never indexed with integer indices — you use the keys, and the key 'two' always maps to 'dos' however the pairs are arranged.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 104-104
Picture it
The same three items, arranged two ways. Nothing you can write distinguishes them.
Figure (svg): Two columns showing the same dictionary printed in two different orders
Both columns are the same colour deliberately: there is no meaningful difference between them, which is exactly what unpredictable order means.
Worked example
Two programs, one of which is asking for trouble.
# fine: asks by key
for word in ['one', 'two', 'three']:
print(word, eng2sp[word])
# risky: relies on the order the dictionary happens to have
first = list(eng2sp)[0]
print('the first word is', first)| Approach | Depends on order? | Verdict |
|---|---|---|
| asking by key | the answer never depends on order | safe |
| taking the first item | depends on storage order | unpredictable |
| the fix | sort, or keep a separate list | make the order explicit |
Look at what the first loop asks.
Why: It supplies the keys itself, in an order it controls, and only uses the dictionary for lookup. Nothing about storage order can affect it.
Look at what the second assumes.
Why: It takes whatever the dictionary happens to yield first, which is not something the language promises.
State the rule.
Why: Rely on lookups, never on position. If you need an order, impose one — from a list you control, or with sorted.
Figure (svg): The state of the program after each line of Worked example what you may and may not rely on, drawn as a ladder with one rung per traced line
The first is safe and the second is not. The distinction is whether the program supplies the order or takes whatever it is given.
Verify: Ask what would make the second program's bug visible.
Why: Very little: it would produce a plausible-looking answer every time, and possibly a different one on another machine or another Python version. That is the worst kind of bug, and it is why the book flags the unpredictability so early.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C describes something that does not exist. sorted(d) returns a sorted list of the keys and leaves the dictionary alone — there is no operation that rearranges a dictionary, because a dictionary has no arrangement to rearrange. If you need a persistent order, that belongs in a list.
Worked example
When you do want an order, ask for one.
>>> h = {'p': 1, 'a': 1, 'r': 2, 't': 1, 'o': 1}
>>> for key in sorted(h):
... print(key, h[key])
a 1
o 1
p 1
r 2
t 1| Part | What it produces | Note |
|---|---|---|
| sorted(h) | a sorted list of the KEYS | not of the values |
| the loop | traverses that list | in a known order |
| h[key] | looks up each value | by key, as always |
Recognise what sorted receives.
Why: Looping over a dictionary gives its keys, so sorted receives the keys and returns them sorted into a list.
Note what it does not touch.
Why: The dictionary is unchanged — sorted returns a new list, exactly as it did for lists in lesson 10c.
Look up inside the loop.
Why: The values still come from h[key], because the sorted list holds keys only.
Figure (svg): A pipeline showing a dictionary yielding keys, sorted producing a list, and lookup producing values
The keys come out in alphabetical order and the values are looked up as usual. The order lives in the list sorted produced, not in the dictionary.
Verify: Check that the dictionary is unaffected.
Why: Printing h afterwards shows the same unpredictable order it had before, because sorted created a list rather than rearranging anything. There is no way to sort a dictionary itself, and after this section it should be clear why the idea does not even make sense.
Trap
A program stores settings in a dictionary and reads the first one to use as a default.
Assume insertion order is preserved
Why: It often appears to be, especially with small dictionaries and recent Pythons.
The book states plainly that the order of items in a dictionary is unpredictable. A program that reads correctly a thousand times can read differently after an unrelated change, and the failure looks like a data problem rather than a code problem.
If order matters, store it somewhere that has one.
Keep a list of keys alongside the dictionary
Why: The list carries the order and the dictionary carries the lookup — each doing what it is good at.
Or sort the keys when you need them
Why: sorted(d) gives a defined order that does not depend on storage.
The general rule is that a dictionary answers what maps to what and a list answers what comes first. Asking one of them the other's question is where this bug comes from.
Prediction
It looks up by key and never touches position.
eng2sp = {'one': 'uno', 'two': 'dos'}
for word in ['two', 'one']:
print(eng2sp[word])| Part | What it does | Result |
|---|---|---|
| the loop list | supplies its own order | 'two' then 'one' |
| eng2sp[word] | looks up by key | order-independent |
| output | dos then uno | every time |
Predict first
Does this program always produce the same output?
Correct: Yes — it supplies its own order and only looks up by key, so nothing about the dictionary's storage can affect it.
Why: The order comes from the list in the for statement, which the program controls. The dictionary is used only for lookup, and a lookup by key gives the same answer whatever order the pairs are stored in. This is the shape to aim for: order from a list, values from a dictionary.
Discrimination
Ask whether the program supplies the order or takes it.
Sort into buckets
For each expression, does the result depend on the dictionary's storage order?
Socratic
A list keeps its order. A dictionary throws it away. Something must be bought with it.
Discussion prompt
Lists preserve order and dictionaries do not. What does a dictionary get in exchange, and why can a list not have both?
Hint: Think about how each one finds something.
Answer:
A list finds an element by counting from the start, so its positions have to mean something and its order is fundamental to how it works.
A dictionary finds a value by computing where the key belongs, which is why the next section can say the lookup takes about the same time however large the dictionary is. Position is not consulted at all.
So the ordering is not thrown away out of carelessness — it is a consequence of storing things by computed location rather than by position. You get fast lookup by any key, and the price is that there is no natural first item.
Section
Section 3
Concept
The in operator works on dictionaries; it tells you whether something appears as a key in the dictionary. Appearing as a value is not good enough.
>>> eng2sp = {'one': 'uno', 'two': 'dos', 'three': 'tres'}
>>> 'one' in eng2sp
True
>>> 'uno' in eng2sp
False
>>> vals = eng2sp.values()
>>> 'uno' in vals
True| Expression | What is checked | Result |
|---|---|---|
| 'one' in eng2sp | 'one' is a key | True |
| 'uno' in eng2sp | 'uno' is a value, not a key | False |
| 'uno' in eng2sp.values() | searches the values | True |
To see whether something appears as a value, you can use the method values, which returns a collection of values, and then use the in operator on that. The two questions are different and the syntax makes the difference visible.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 104-104
Picture it
The keys are on the left of the colons and that is where in looks.
Figure (svg): A dictionary diagram with the key column highlighted as the side the in operator searches
This is a genuine asymmetry, and it is the same asymmetry that makes the next lesson's reverse lookup awkward: going from key to value is built in, and going back is not.
Worked example
The same operator, two very different amounts of work.
# a list: searches the elements in order
'q' in ['a', 'b', 'c', 'd', 'e']
# a dictionary: computes where 'q' would be, and looks
'q' in {'a': 1, 'b': 2, 'c': 3, 'd': 4, 'e': 5}| Expression | How it searches | Cost |
|---|---|---|
| in on a list | checks each element in turn | longer list, longer search |
| in on a dictionary | uses a hashtable | about the same time regardless of size |
| the consequence | the choice of structure changes the speed | not just the syntax |
Look at the list case.
Why: For lists, in searches the elements of the list in order, as in the search pattern from chapter 8. As the list gets longer, the search time gets longer in direct proportion.
Look at the dictionary case.
Why: Python dictionaries use a data structure called a hashtable that has a remarkable property: the in operator takes about the same amount of time no matter how many items are in the dictionary.
Draw the practical conclusion.
Why: For a hundred items the difference hardly shows. For a hundred thousand, a program built on list searches can be unusably slow while the same program built on dictionary lookups is instant.
Figure (svg): A growth chart comparing linear list search against flat dictionary lookup as size increases
The list search grows with the list and the dictionary search does not. The operator is spelled the same way, and the two do quite different amounts of work.
Verify: Check what this costs.
Why: Nothing, in the common case — the price was paid in the previous section, as unpredictable ordering. Fast lookup by key and no natural order are two consequences of the same design, which is why they arrive together.
Prediction
The string appears in the dictionary — on one side of the colon.
eng2sp = {'one': 'uno', 'two': 'dos'}
print('dos' in eng2sp)| Part | What is true | Result |
|---|---|---|
| 'dos' | appears as a value | not as a key |
| in on a dictionary | checks the keys only | False |
| the alternative | 'dos' in eng2sp.values() | would be True |
Predict first
What does this print?
Correct: False — 'dos' appears as a value, and in checks only whether something appears as a key.
Why: Appearing as a value is not good enough for the in operator. To ask about values you have to say so, with 'dos' in eng2sp.values(), which returns True. This is one of the few places in the chapter where a wrong test is legal and silent, so it is worth fixing the habit early.
Worked example
Same data, two structures, and one of them answers the question well.
# a list of words seen so far
seen_list = ['the', 'cat', 'sat']
'cat' in seen_list # searches in order
# a dictionary with the words as keys
seen_dict = {'the': 1, 'cat': 1, 'sat': 1}
'cat' in seen_dict # one hashtable lookup| Aspect | What is true | Note |
|---|---|---|
| the question | have I seen this word? | a membership question |
| the list | answers it by scanning | correct but slow at scale |
| the dictionary | answers it directly | and the values are unused |
Identify the question being asked.
Why: Have I seen this? is a membership question, and membership is exactly what a dictionary key is good for.
Notice the values are doing nothing.
Why: Every value is 1 and no one reads them. The dictionary is being used purely for its keys, which is a common and legitimate pattern.
Weigh the cost.
Why: The dictionary uses more memory per item and answers in about the same time however many items there are. For a membership test repeated many times, that trade is usually worth taking.
Figure (svg): The state of the program after each line of Worked example choosing the structure for the question, drawn as a ladder with one rung per traced line
Both are correct; the dictionary scales and the list does not. Choosing a structure is choosing which questions will be cheap to ask.
Verify: Ask what the list is still better at.
Why: Keeping the order the words arrived in, and allowing duplicates — neither of which a dictionary keyed on the word can do. The structures are not ranked; they answer different questions, and this one happens to suit the dictionary.
Trap
A program checks if 'uno' in eng2sp: to find out whether that Spanish word is in the dictionary.
Read in as appears anywhere in
Why: Which is how it works for a list, where there is only one kind of element.
It reports False, because 'uno' is a value and in checks keys. The test is legal, silent, and wrong — the program simply concludes the word is absent.
Ask about values explicitly.
'uno' in eng2sp.values()
Why: The method name makes it plain which side you are searching.
Or reconsider which way the mapping should go
Why: If you mostly ask about Spanish words, the dictionary may be the wrong way round.
Note that searching the values is a linear search, so it is slow in the way a list search is slow. The asymmetry in speed matches the asymmetry in syntax, and both point the same way: dictionaries are built for asking about keys.
Invariant
Two structures, four sizes. Watch which line moves.
Step through it
Which quantity stays the same across all four frames, and which grows?
The dictionary lookup time stays about constant while the list search grows in direct proportion to the size. That constancy is the property the hashtable buys, and it is the reason the book calls dictionaries the building blocks of efficient algorithms.
Comparison
Fill the blanks. Same operator, different machinery.
Comparison matrix
| Question | List | Dictionary |
|---|---|---|
| What does it search? | the elements | the keys |
| How does it search? | in order, one at a time | by computing a location in a hashtable |
| Cost as the size grows | grows in direct proportion | stays about the same |
| How do you search the other part? | there is no other part | d.values(), which is a linear search |
The last row is the asymmetry: a dictionary has two sides and only one of them is fast to search.
Explain it
The code did not change. The data got bigger.
Discussion prompt
A classmate's program checks each of 50,000 words against a list of 50,000 known words and it now takes minutes. Explain what is happening and what to change.
Hint: How many comparisons is that, in total?
Answer:
Every in on the list scans it, so 50,000 checks against a 50,000-element list is on the order of two and a half billion comparisons. Nothing is broken; the work is simply enormous.
Building a dictionary with the known words as keys turns each check into one lookup that takes about the same time regardless of size, so the total becomes 50,000 lookups rather than billions of comparisons.
The tell for this whole family of problem is a program that was fine on test data and is unusable on real data. It is almost always a linear search inside a loop, and swapping the list for a dictionary is the standard fix.
Section
Section 4
Concept
Suppose you are given a string and you want to count how many times each letter appears. There are several ways you could do it, and each of them implements the same computation in a different way.
implementation — A way of performing a computation.
Each of these options performs the same computation, but each implements it differently — and some implementations are better than others. An advantage of the dictionary is that we don't have to know ahead of time which letters appear, and we only have to make room for the ones that do.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 104-105
Picture it
The result is identical. What differs is what has to be arranged in advance.
Figure (svg): Two columns comparing the fixed 26-slot implementations with the dictionary implementation
Notice the third bullet. The list implementation needs a translation from character to integer, and the dictionary does not — because its indices need not be integers.
Worked example
The book's version, with the conditional made explicit.
def histogram(s):
d = dict()
for c in s:
if c not in d:
d[c] = 1
else:
d[c] += 1
return d| Step | Situation | The dictionary after |
|---|---|---|
| d = dict() | an empty dictionary | {} |
| c = 'b', not in d | first sighting | {'b': 1} |
| c = 'r', not in d | first sighting | {'b': 1, 'r': 1} |
| c = 'o', not in d | first sighting | {'b': 1, 'r': 1, 'o': 1} |
| a repeat, say 'r' | already in d | d['r'] becomes 2 |
Start empty.
Why: The first line creates an empty dictionary — nothing is assumed about which characters will appear.
Traverse the string.
Why: The for loop gives one character at a time, which is the string traversal from chapter 8.
Branch on whether it is new.
Why: If the character c is not in the dictionary, create a new item with key c and value 1, since we have seen this letter once. Otherwise increment d[c].
Figure (svg): The state of the program after each line of Worked example the histogram function, drawn as a ladder with one rung per traced line
histogram('brontosaurus') gives {'a': 1, 'b': 1, 'o': 2, 'n': 1, 's': 2, 'r': 2, 'u': 2, 't': 1} — a statistical histogram, which is a collection of counters or frequencies.
Verify: Check one count by hand.
Why: 'brontosaurus' contains two r's, at positions 1 and 7, and the dictionary reports 'r': 2. Checking a repeated letter rather than a single one tests the else branch, which is the half of the conditional the easy cases never reach.
Prediction
A short string with one repeat.
def histogram(s):
d = dict()
for c in s:
if c not in d:
d[c] = 1
else:
d[c] += 1
return d
print(histogram('aba'))| Step | Situation | The dictionary after |
|---|---|---|
| c = 'a' | not in d | {'a': 1} |
| c = 'b' | not in d | {'a': 1, 'b': 1} |
| c = 'a' | already in d | {'a': 2, 'b': 1} |
Predict first
What does this print?
Correct: {'a': 2, 'b': 1} — 'a' appears twice and 'b' once.
Why: The first 'a' takes the if branch and creates the counter at 1; the 'b' does the same; the second 'a' takes the else branch and increments to 2. The last option is impossible: each key is associated with a single value, so a dictionary can never show the same key twice.
Worked example
Remove the if and the function fails on its very first character.
def broken_histogram(s):
d = dict()
for c in s:
d[c] += 1 # WRONG
return d| Step | What happens | Result |
|---|---|---|
| c = 'b' | d['b'] does not exist | KeyError |
| why | += reads before it writes | and there is nothing to read |
| the fix | the if, or get with a default | give it something to read |
Expand the augmented assignment.
Why: d[c] += 1 means d[c] = d[c] + 1, so it reads the current value before storing a new one.
See what the read finds.
Why: On the first sighting there is no item for c, so the read raises KeyError before any store happens.
Identify what the conditional supplies.
Why: The if branch handles exactly the case where there is nothing to read, by writing 1 directly.
Figure (svg): A flowchart showing the first-sighting branch and the increment branch of the histogram loop
It raises KeyError on the first character. The conditional exists to handle the first sighting, which is the only case where there is no counter to increment.
Verify: Check that the error is immediate.
Why: It fails on the first character of any non-empty string, so the bug cannot hide — which makes it much friendlier than the silent list bugs of chapter 10. The next idea removes the conditional entirely, which is the exercise the book sets.
Trap
A student writes d[c] = 0 in the first branch, reasoning that a new counter should start at zero.
Start counters at zero
Why: Which is right when the counting happens afterwards, and it usually does.
Here the first sighting IS a sighting, so the character ends up counted one short. Every count in the histogram is wrong by one, and nothing raises — the numbers are simply too small.
Set it to 1, because you are looking at the first occurrence.
Ask what has happened by the time this line runs
Why: You have seen c once. The counter should say one.
Or write d[c] = 0 followed by d[c] += 1
Why: Which is longer and makes the reasoning explicit — and is what get with a default does in the next idea.
This is an off-by-one error of the kind lesson 7b warned about, and the check that catches it is counting a single-character string by hand: histogram('a') must be {'a': 1}.
Faded example
What should a brand new counter hold?
Fill in the blanks
for c in s:
if c not in d:
d[c] = 1
else:
d[c] += 1
Why: One, not zero — by the time this line runs you have already seen the character once, so the counter should record that sighting. Initialising to 0 makes every count in the histogram one too small, and nothing raises to tell you.
Ranking
Four steps, in the order they happen.
Put in order
Why: The loop takes a character, tests membership, acts on the answer, and repeats. The membership test has to come before the increment, because incrementing a counter that does not exist raises KeyError — which is precisely what the broken version demonstrated.
Real world
Counting occurrences by category is one of the most common things programs do.
Discussion prompt
Name a situation where you would want to count how often each of many possible things occurred, without knowing in advance what the things would be.
Hint: Anything with a long tail of categories.
Answer:
Words in a document, error codes in a log, products in an order history, characters in a text — all cases where the set of categories is discovered from the data rather than fixed beforehand.
That not-knowing-in-advance is exactly the advantage the book names for the dictionary implementation: you only make room for what actually appears.
The pattern generalises past counting. Any time you want for each distinct thing, some accumulated fact, a dictionary keyed on the thing is the structure — totals, first-seen times, lists of examples. The counting case is just the simplest instance.
Section
Section 5
Concept
Dictionaries have a method called get that takes a key and a default value. If the key appears in the dictionary, get returns the corresponding value; otherwise it returns the default value.
>>> h = histogram('a')
>>> h
{'a': 1}
>>> h.get('a', 0)
1
>>> h.get('c', 0)
0| Expression | Situation | Result |
|---|---|---|
| h.get('a', 0) | the key is present | returns its value, 1 |
| h.get('c', 0) | the key is absent | returns the default, 0 |
| h['c'] | the same question with brackets | would raise KeyError |
The whole point is the second case. Bracket lookup raises when the key is missing, and get returns something you chose — which turns ask and handle the failure into ask and get an answer.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 105-106
Picture it
The same question, and two different things to deal with afterwards.
Figure (svg): Two columns contrasting bracket lookup raising KeyError with get returning a default
The choice is about intent: if a missing key means something has gone wrong, you want the exception. If it means none yet, you want the default.
Worked example
The book sets this as an exercise. get is the whole trick.
def histogram(s):
d = dict()
for c in s:
d[c] = d.get(c, 0) + 1
return d| Step | What get returns | The assignment |
|---|---|---|
| c = 'a', first time | d.get('a', 0) gives 0 | d['a'] = 1 |
| c = 'b', first time | d.get('b', 0) gives 0 | d['b'] = 1 |
| c = 'a', again | d.get('a', 0) gives 1 | d['a'] = 2 |
Replace the read with get.
Why: The only reason the conditional existed was that the read failed on a first sighting. get makes the read always succeed.
Choose the default.
Why: Zero, because not seen yet means a count of zero — and adding one to it gives the right answer for a first sighting.
Notice the branches have merged.
Why: First sighting and repeat are now the same line: get gives 0 or the current count, and both cases add one.
Figure (svg): A ladder showing the value of d['a'] through three sightings of the character a
Four lines instead of seven, with no conditional. The first sighting is no longer a special case because get supplied the value it was missing.
Verify: Check it against the original on a string with a repeat.
Why: Both give {'a': 2, 'b': 1} for 'aba'. Checking a repeat matters because the two versions differ precisely in how they handle first sightings, and a string with no repeats would exercise only half of the original's logic.
Prediction
The key is not in the dictionary.
h = {'a': 1, 'b': 2}
print(h.get('z', 0))| Expression | Situation | Result |
|---|---|---|
| h.get('z', 0) | 'z' is not a key | the default is used |
| the default | the second argument | 0 |
| h['z'] | the same question with brackets | would raise KeyError |
Predict first
What does this print?
Correct: 0 — the key is absent, so get returns the default value it was given.
Why: get takes a key and a default, and returns the default when the key is not present. Nothing is raised and nothing is added to the dictionary — get is a pure lookup, so h is unchanged and still has two items afterwards. That last point matters: get never creates the key it was asked about.
Worked example
get is not always the improvement.
settings = {'width': 80, 'height': 24}
# absence means a bug: let it raise
w = settings['width']
# absence means 'use the default': ask with get
colour = settings.get('colour', 'white')| Line | What absence would mean | Which to use |
|---|---|---|
| settings['width'] | must be present | KeyError would be correct |
| settings.get('colour', ...) | may be absent | a default is correct |
| the choice | what does absence mean? | that decides the syntax |
Ask what a missing key would mean.
Why: For a required setting, it means the configuration is broken and the program should not continue quietly.
Use brackets when absence is an error.
Why: The KeyError names the key and stops the program at the point of the problem, which is the most useful thing that can happen.
Use get when absence is a legitimate state.
Why: An optional setting has a sensible default, and supplying it is not papering over anything.
Figure (svg): The state of the program after each line of Worked example when the exception is what you want, drawn as a ladder with one rung per traced line
Brackets for keys that must be there, get for keys that may not be. Reaching for get everywhere replaces useful failures with plausible wrong answers.
Verify: Imagine the bug each wrong choice produces.
Why: get on a required key means a missing configuration silently becomes a default and the program runs with values nobody chose. Brackets on an optional key means an ordinary situation crashes the program. Both mistakes are about intent, not syntax, which is why the question to ask is what absence means.
Trap
A student writes d.get(c, 0) + 1 on a line by itself, expecting the counter to advance.
Read get as the counter
Why: It does produce the counter's value, so the expression looks like it is operating on the dictionary.
get is a returning method, not a modifying one — lesson 10c's distinction exactly. The expression computes a number and discards it, and the histogram comes back with every count at zero or absent.
Store what you computed.
d[c] = d.get(c, 0) + 1
Why: The assignment is what changes the dictionary; get only supplies the old value.
Check the shape of the line
Why: A line with no assignment and no side effect does nothing, whatever it computes.
This is the returning-function misuse from lesson 10c in a new setting: calling a function that returns a value and not assigning the result. It never raises, and the symptom is a result that is simply empty.
Sorting
Ask what a missing key would mean in each case.
Sort into buckets
For each lookup, which form is the better choice?
Faded example
One line replaces the whole conditional.
Fill in the blanks
def histogram(s):
d = dict()
for c in s:
d[c] = d.get(c, 0) + 1
return d
Why: Zero is the count of a character not yet seen, so adding one gives 1 on a first sighting and the correct increment thereafter. A default of 1 would double-count every first sighting; a default of None would raise a TypeError on the addition, which at least fails loudly.
Explain it to yourself
The two versions of histogram compute the same thing with different amounts of code.
Discussion prompt
Explain why supplying a default value makes the if statement unnecessary, in terms of what the first sighting used to need.
Hint: What was the if actually testing for?
Answer:
The conditional existed for exactly one reason: on a first sighting there was no value to read, so the read had to be skipped and a value written directly.
get removes that difference by making the read always succeed. On a first sighting it produces the default, and on a repeat it produces the stored count — so the code after it is the same either way.
The general move is worth naming: a special case disappears when you can make the general case cover it. That is the same reasoning behind an empty list being a legitimate list and an empty string being a legitimate string, both of which let loops start cleanly rather than needing a first-time branch.
Comparison
Fill the blanks. Everything is the same except what the index is.
Comparison matrix
| Question | List | Dictionary |
|---|---|---|
| What can the index be? | an integer | almost any type |
| Is there an order? | yes, and it is fundamental | no, and it is unpredictable |
| Error for a missing index | IndexError | KeyError |
| How fast is in? | grows with the length | about constant, whatever the size |
The last two rows are the practical difference. The first two explain why: generalising the index costs you the ordering and buys you the speed.
Pattern
Five steps, and the fourth is where the two versions differ.
Step 5's choice matters as much as step 4's. Reading with brackets raises for a thing that never occurred; reading with get reports zero, which is usually what you meant.
Python documentation — Data Structures Data Structures
Check
The value appears. The key does not.
d = {'a': 'apple', 'b': 'banana'}
print('apple' in d)| Part | What is true | Result |
|---|---|---|
| 'apple' | appears as a value | not as a key |
| in on a dictionary | checks keys only | False |
| 'apple' in d.values() | checks values | True |
Check your understanding
What does this print?
Answer: B
Why: The in operator tells you whether something appears as a key; appearing as a value is not good enough. To ask the other question you write 'apple' in d.values(), which is True — and which is a linear search rather than a hashtable lookup, so it is slow in the way a list search is slow.
Check
One key is present and one is not.
h = {'a': 3}
print(h.get('a', 0) + h.get('b', 0))| Expression | Situation | Value |
|---|---|---|
| h.get('a', 0) | 'a' is present | 3 |
| h.get('b', 0) | 'b' is absent | 0, the default |
| the sum | 3 + 0 | 3 |
Check your understanding
What does this print?
Answer: A
Why: get returns the stored value for a present key and the default for an absent one, so the two calls produce 3 and 0. Writing h['a'] + h['b'] instead would raise KeyError on the second lookup — which is the whole reason get exists.
Check
Three implementations compute the same histogram.
Check your understanding
What advantage does the book name for the dictionary implementation over a 26-element list?
Answer: B
Why: The advantage named is exactly this: no advance knowledge and no wasted slots. The list implementation must reserve 26 counters whether or not they are used, and it works only for characters that convert to a known range of integers — so it cannot cope with punctuation, accents or arbitrary keys.
Real world
A mapping is one of the oldest ways of organising information.
Discussion prompt
Think of something outside programming that maps one thing to another and is looked up by the first thing — an index, a directory, a glossary. What does it have in common with a dictionary, and where does the analogy break?
Hint: Try looking one up backwards.
Answer:
A phone directory maps names to numbers, an index maps topics to pages, a glossary maps terms to definitions. In each case you arrive knowing the key and leave with the value, and looking one up is fast.
And each is hard to use backwards. Finding whose number this is, from a printed directory, means reading the whole thing — which is exactly why the next lesson's reverse lookup has to search.
Where the analogy breaks is the ordering. A printed directory is alphabetical because a human has to find the entry; a Python dictionary is unordered because the machine computes where to look. The alphabetisation is a human interface, not part of what a mapping is.
Commit first
Answer, then rate your confidence. This is the chapter's most-missed fact.
Predict first
A dictionary maps 'one' to 'uno'. What does the expression 'uno' in d report?
Correct: False — 'uno' is a value, and the in operator tells you only whether something appears as a key.
Why: The book puts it flatly: the in operator tells you whether something appears as a key in the dictionary, and appearing as a value is not good enough. To ask the other question you need d.values(), which returns a collection of values that in can search — and that search is linear, so it is slow in the way a list search is slow. The asymmetry runs all the way through: dictionaries are built to be asked about keys, which is why the next lesson has to write a whole function to go the other way.
Explain it
One sentence about what changed, and one about what it cost.
Discussion prompt
A classmate who is comfortable with lists asks what a dictionary is for. Explain it as one change to lists, and name what that change gives up and what it buys.
Hint: Start with the index.
Answer:
Say: it is a list where the index can be anything, not just an integer. Everything else — bracket lookup, len, in, looping — is the same shape.
What it gives up is order. There is no first item, and the order things print in is unpredictable, so anything that needs an order has to keep it elsewhere.
What it buys is that in and lookup take about the same time however big it gets, where a list search gets slower in direct proportion. That is the trade, and it is why the book calls dictionaries the building blocks of efficient algorithms.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The creation syntax is mechanical and settles with practice, though the dict-as-a-variable-name trap catches people once. The ordering point matters most when you write a program that depends on it without noticing. The in operator hides a real asymmetry — keys are fast, values are not — and that asymmetry is the whole reason the next lesson exists. And the histogram is the pattern you will reuse most: for each distinct thing, some accumulated fact.
Connect it up
One page, from memory.
Draw it
Draw a dictionary as a box with three key-value pairs inside and the type dict written above it. Beside it write the four operations — bracket lookup, bracket assignment, len, and in — and next to each, what it does and what happens when the key is absent. Underneath, write the histogram loop in both forms, the one with the conditional and the one with get, and mark on the get version which sighting uses the default.
Recap
Three pages, and the type Downey calls one of Python's best features.
| If you remember one thing | It is this |
|---|---|
| From the mapping idea | A dictionary is a list whose indices can be anything. |
| From ordering | There is no first item, and nothing should depend on there being one. |
| From in | It checks keys. Values need .values(), and that search is slow. |
| From the hashtable remark | Dictionary lookup does not get slower as the dictionary grows. |
| From histogram | d[k] = d.get(k, 0) + 1 counts anything, with no special case for the first. |
The next lesson loops over dictionaries, writes a reverse lookup — which has to search, and raises an exception when it fails — and puts lists inside dictionaries to invert a mapping, which is where it becomes clear why keys must be immutable.
Think Python, 2nd edition — Allen B. Downey §11.1-11.2, pp. 103-105 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.