This lesson introduces three containers from the standard library: sets for collections of unique keys, Counters for counting occurrences, and defaultdict for generating a value when a key is missing.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 19 — The Goodies
§19.5-19.7, pp. 186-188
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-188 — the pages these objectives are drawn from
Warm-up
You wrote one in chapter 13.
Discussion prompt
The subtract function built a dictionary and set every value to None, because we never use them. What does that suggest about the structure being used?
Hint: What is actually wanted from it?
Answer:
That only the keys matter — the dictionary is being used for fast membership testing and automatic removal of duplicates, and the values are there because a dictionary requires them.
Which means the program is emulating a container it does not have: a collection of keys with no values.
Python provides exactly that, and this lesson introduces it — along with two others that each replace something the course has already done the long way.
Concept
Each of these three types replaces several lines of dictionary manipulation with one, because each is a dictionary specialised for a job the course has already done by hand.
All three are specialisations of the dictionary, and each one exists because the plain dictionary version of the same job needs boilerplate that the specialised version does not.
Figure (svg): Two columns pairing each specialised container with the dictionary pattern it replaces
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-188
Section
Section 1
Concept
Python provides another built-in type, called a set, that behaves like a collection of dictionary keys with no values. Adding elements to a set is fast; so is checking membership. And sets provide methods and operators to compute common set operations.
# the dictionary version, from chapter 13
def subtract(d1, d2):
res = dict()
for key in d1:
if key not in d2:
res[key] = None
return res
# with sets
def subtract(d1, d2):
return set(d1) - set(d2)| Version | How long | Note |
|---|---|---|
| the dictionary version | six lines | with None values |
| the set version | one line | no values at all |
| the minus operator | set difference | built in |
In all of those dictionaries, the values are None because we never use them, and as a result we waste some storage space. Set subtraction is available as a method called difference or as an operator, and the result is a set rather than a dictionary — but for operations like iteration, the behaviour is the same.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-186
Picture it
Keys, and values nobody reads.
Figure (svg): Two columns contrasting a dictionary used as a set with an actual set
Which is why lesson 11b's subtract felt like an approximation: it was, and the type it was approximating is this one.
Worked example
The set's uniqueness does the work.
# with a dictionary
def has_duplicates(t):
d = {}
for x in t:
if x in d:
return True
d[x] = True
return False
# with a set
def has_duplicates(t):
return len(set(t)) < len(t)| Version | How it works | Note |
|---|---|---|
| the dictionary version | tracks what it has seen | seven lines |
| the set version | compares two lengths | one line |
| why it works | an element can only appear in a set once |
Note what a set enforces.
Why: An element can only appear in a set once, so building one from a sequence discards the duplicates.
Compare the lengths.
Why: If an element in t appears more than once, the set will be smaller than t; if there are no duplicates, it will be the same size.
Note what disappeared.
Why: The loop, the accumulator and the early return — the uniqueness the dictionary was tracking by hand is a property of the container.
Figure (svg): The state of the program after each line of Worked example has duplicates in one line, drawn as a ladder with one rung per traced line
One comparison in place of seven lines. The set's defining property is exactly the question being asked.
Verify: Ask what the two versions do differently.
Why: The dictionary version stops at the first duplicate and the set version builds the whole set first — so on a long sequence with an early duplicate, the loop is faster. Both are correct, and the shorter one is not universally better, which is worth noticing before adopting it everywhere.
Prediction
A set discards duplicates.
t = ['a', 'b', 'a', 'c', 'b']
print(len(set(t)), len(t))| Container | How many | Note |
|---|---|---|
| the list | five elements | with repeats |
| the set | three distinct | a, b, c |
| the comparison | 3 < 5 | duplicates present |
Predict first
What does this print?
Correct: 3 5 — an element can only appear in a set once, so the three distinct values survive.
Why: That difference is exactly what has_duplicates tests: if an element appears more than once, the set will be smaller than the sequence. Equal lengths would mean no duplicates, which is why the one-line version works.
Worked example
uses_only, rewritten.
# with a loop
def uses_only(word, available):
for letter in word:
if letter not in available:
return False
return True
# with sets
def uses_only(word, available):
return set(word) <= set(available)| Part | What it means | Note |
|---|---|---|
| <= | subset, including equal | a set operator |
| the question | are all word's letters available? | |
| the loop version | the same question, spelled out |
Note what the loop asks.
Why: uses_only checks whether all letters in word are in available — which is the subset relation.
Note the operator.
Why: The <= operator checks whether one set is a subset of another, including the possibility that they are equal.
Note the reading.
Why: The letters of word are among the letters available, which is what the mathematical relation means.
Figure (svg): Set operators paired with the questions they answer
One operator in place of a search loop, because the question was a set question all along. Sets provide methods and operators to compute common set operations.
Verify: Check the equal case.
Why: The <= operator includes equality, so a word using exactly the available letters reports True — which is right for uses only. The strict version < would exclude it, which would be a subtly different and less useful question.
Trap
A program builds a set from a list and expects to iterate over it in the original order.
Treat it as a list without duplicates
Why: Which is what it looks like.
A set behaves like a collection of dictionary keys, and dictionary keys have no order — so the traversal order is unpredictable, exactly as chapter 11 warned.
Use a set for membership and uniqueness, not for order.
sorted(set(t)) if you need a defined order
Why: Which gives a list, ordered.
Or keep the original list alongside
Why: The list carries the order and the set carries the membership.
This is chapter 11's rule inherited: a set is a dictionary's keys without the values, so everything true of key ordering is true here. Fast membership and no order arrive together.
Discrimination
Ask whether the values matter.
Sort into buckets
For each job, which container fits?
Faded example
Words in the book that are not in the word list.
Fill in the blanks
def subtract(d1, d2):
return set(d1) - set(d2)
Why: Set subtraction is available as a method called difference or as the minus operator, and it gives the elements of the first that are not in the second. That is the whole of chapter 13's six-line subtract, which was building a dictionary whose values were all None because we never use them.
Explain it to yourself
The book says so directly.
Discussion prompt
chapter 13's subtract set every value to None. Explain what that cost and what it revealed.
Hint: What is a dictionary storing?
Answer:
It cost storage: in all of those dictionaries, the values are None because we never use them, and as a result we waste some space.
What it revealed is that the program wanted a different container. It was using a dictionary for fast membership and automatic uniqueness, and paying for a values column it never read.
So the None was a signal rather than a detail — a dictionary whose values are all the same placeholder is almost always a set that has not been written as one. Recognising that signal is worth more than the specific rewrite.
Section
Section 2
Concept
A Counter is like a set, except that if an element appears more than once, the Counter keeps track of how many times it appears. If you are familiar with the mathematical idea of a multiset, a Counter is a natural way to represent one.
>>> from collections import Counter
>>> count = Counter('parrot')
>>> count
Counter({'r': 2, 't': 1, 'o': 1, 'p': 1, 'a': 1})
>>> count['d']
0| Expression | What happens | Note |
|---|---|---|
| Counter('parrot') | counts the letters | in one call |
| the result | like a dictionary | keys to counts |
| a missing key | returns 0 | rather than raising |
Counter is defined in a standard module called collections, so you have to import it. You can initialise one with a string, list, or anything else that supports iteration — and as in dictionaries, the keys have to be hashable.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 187-187
Picture it
The function you wrote is now a constructor.
Figure (svg): Two columns comparing the hand-written histogram with a Counter
The zero-for-missing behaviour is the substantive difference, and it removes the get-with-a-default that lesson 11a introduced for exactly this problem.
Worked example
Two words are anagrams if their Counters match.
def is_anagram(word1, word2):
return Counter(word1) == Counter(word2)| Part | What it computes | Note |
|---|---|---|
| Counter(word1) | letters to counts | for the first word |
| Counter(word2) | likewise | for the second |
| == | same letters, same counts | which is what anagram means |
State the definition.
Why: If two words are anagrams, they contain the same letters with the same counts.
Note what that is.
Why: Exactly what a Counter records — so two anagrams have equivalent Counters.
Note that == works.
Why: Counters behave like dictionaries in many ways, including comparing by contents rather than by identity.
Figure (svg): The state of the program after each line of Worked example is anagram in one line, drawn as a ladder with one rung per traced line
One line, and it is the definition of an anagram written directly. The container matches the question, which is what makes it so short.
Verify: Check that == compares contents.
Why: Two Counters built from anagrams are equal even though they are separate objects, because they compare like dictionaries rather than like the programmer-defined classes of chapter 15. That is worth confirming, since the default for a class you write is identity — Counter is a library type that defines equality by value.
Prediction
The letter is not in the word.
count = Counter('parrot')
print(count['d'])| Aspect | What happens | Note |
|---|---|---|
| 'd' | not in 'parrot' | absent |
| a plain dictionary | would raise KeyError | |
| a Counter | returns 0 |
Predict first
What does this print?
Correct: 0 — unlike dictionaries, Counters don't raise an exception if you access an element that doesn't appear.
Why: The zero is the natural answer for a count, and it removes the need for the get-with-a-default idiom that a plain dictionary histogram required. It is also the one substantive way a Counter differs from a dictionary, which is why code that expects a KeyError from one silently never triggers.
Worked example
Lesson 13b's ranking, built in.
>>> count = Counter('parrot')
>>> for val, freq in count.most_common(3):
... print(val, freq)
r 2
p 1
a 1| Part | What it gives | Note |
|---|---|---|
| most_common(3) | the three commonest | as pairs |
| each pair | value and frequency | unpacked |
| the order | most common to least | sorted for you |
Note what it returns.
Why: A list of value-frequency pairs, sorted from most common to least.
Note the argument.
Why: How many pairs to return — three here, and omitting it gives all of them.
Compare with chapter 13.
Why: That chapter built the same thing by hand: a list of (frequency, word) tuples with a reverse sort, chosen so the sort field came first.
Figure (svg): Two columns comparing the hand-built ranking with the Counter method
The commonest elements, already ordered. Note the pairs are (value, frequency) rather than the (frequency, value) that lesson 13b needed for sorting.
Verify: Compare the tuple order with lesson 13b's.
Why: most_common gives (value, frequency), where most_common's own sorting removed the need to put the frequency first. That reversal is a small trap when converting code: the hand-written version's unpacking order does not transfer, and the swap is silent.
Trap
A program tests whether a letter appeared by catching a KeyError from a Counter.
Expect dictionary behaviour
Why: Counters behave like dictionaries in many ways.
Unlike dictionaries, Counters don't raise an exception if you access an element that doesn't appear — they return 0. The except clause never runs, and every absent letter reports as present with a count of zero.
Test the count, or use in.
if count[letter] > 0
Why: Which reads naturally, since absent means zero.
Or if letter in count
Why: Which distinguishes absent from present-with-zero.
The zero-for-missing behaviour is a feature: it removes the get-with-a-default that a plain dictionary histogram needed. It is also the one place a Counter does not behave like a dictionary, which is exactly why it catches people.
Faded example
Same letters, same counts.
Fill in the blanks
from collections import Counter
def is_anagram(word1, word2):
return Counter(word1) == Counter(word2)
Why: If two words are anagrams, they contain the same letters with the same counts, so their Counters are equivalent. Counters compare by contents rather than by identity — unlike the programmer-defined classes of chapter 15, where == falls back on is.
Comparison
Fill the blanks. They differ in one substantive way.
Comparison matrix
| Question | A dictionary | A Counter |
|---|---|---|
| How do you build one from a string? | a loop with get | Counter(the string) |
| A missing key? | raises KeyError | returns 0 |
| The commonest elements? | build tuples and sort | most_common(n) |
| Do the keys have to be hashable? | yes | yes — as in dictionaries |
The second row is the one that catches people, and it is also what makes the counting idiom unnecessary.
Explain it
Chapter 11's four lines are now one call.
Discussion prompt
A classmate has just discovered Counter and asks whether chapter 11's histogram was a waste of time. Answer them.
Hint: What did writing it teach?
Answer:
Not a waste: writing it is how you learn the accumulator pattern, the get-with-a-default idiom, and why the first sighting is a special case — all of which recur constantly and none of which is specific to counting.
And Counter is a specialised dictionary, so understanding it means understanding the dictionary underneath. The zero-for-missing behaviour only makes sense once you know what a plain dictionary does instead.
Which is the book's whole approach: teach as little Python as possible first, and come back for the good bits. The shortcut is more useful once you could have written it yourself.
Section
Section 3
Concept
The collections module also provides defaultdict, which is like a dictionary except that if you access a key that doesn't exist, it can generate a new value on the fly.
factory — A function used to create objects, often passed as a parameter to a function like defaultdict.
>>> from collections import defaultdict
>>> d = defaultdict(list)
>>> t = d['new key']
>>> t
[]
>>> t.append('new value')
>>> d
defaultdict(<class 'list'>, {'new key': ['new value']})| Step | What happens | Note |
|---|---|---|
| defaultdict(list) | list is the factory | a class object |
| d['new key'] | the key is absent | so a new list is made |
| the new list | also added to the dictionary | not just returned |
When you create a defaultdict, you provide a function that's used to create new values. Notice that the argument is list, which is a class object, not list(), which is a new list.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 188-188
Picture it
A missing key triggers it; a present one does not.
Figure (svg): A flowchart showing a defaultdict calling its factory only for a missing key
And the new value is added to the dictionary as well as returned — which is what makes the append in the book's example stick.
Worked example
list, not list().
d = defaultdict(list) # correct: the class object
d = defaultdict(list()) # wrong: one list, not a factory| Argument | What it is | Note |
|---|---|---|
| list | a class object | callable — a factory |
| list() | a new empty list | not callable |
| the difference | one makes lists on demand | the other is a list |
Note what is wanted.
Why: A function that's used to create new values — the built-in functions that create lists, sets and other types can be used as factories.
Note the distinction.
Why: The argument is list, which is a class object, not list(), which is a new list.
Note why it matters.
Why: The defaultdict needs to call it once per missing key, so it needs something callable rather than one already-made value.
Figure (svg): Two columns contrasting passing a class object with passing an instance
The class object, uncalled. This is the same distinction as chapter 15's Point against Point(): one is the factory and the other is a product.
Verify: Ask what a single shared list would do.
Why: Every missing key would get the same list, so appending under one key would appear under all of them — the mutable-default trap of lesson 13b, in a new place. Passing a factory rather than a value is precisely what avoids it.
Prediction
The factory is list.
d = defaultdict(list)
t = d['new key']
print(t)| Step | What happens | Result |
|---|---|---|
| the key | absent | so the factory runs |
| list() | a new empty list | |
| the result | returned and stored | [] |
Predict first
What does this print?
Correct: an empty list — the factory generates a new value on the fly, which is both returned and added to the dictionary.
Why: The function you provide doesn't get called unless you access a key that doesn't exist, and here it produces an empty list. The storing is what matters: the new list is also added to the dictionary, so appending to it modifies what is kept.
Worked example
The pattern defaultdict is for.
# with a plain dictionary
if key not in d:
d[key] = [value]
else:
d[key].append(value)
# with a defaultdict
d[key].append(value)| Version | How the first sighting is handled | Note |
|---|---|---|
| the conditional | handles the first sighting | four lines |
| the defaultdict | makes the list on demand | one line |
| the append | works either way | there is always a list |
Recall the plain version.
Why: This is lesson 11b's invert_dict: a singleton on the first sighting and an append thereafter.
Note what the factory removes.
Why: The first sighting stops being a special case, because a missing key produces an empty list to append to.
Note the book's own use.
Why: If you are making a dictionary of lists, you can often write simpler code using defaultdict — as in its solution mapping sorted letters to the words that can be spelled with them.
Figure (svg): The state of the program after each line of Worked example a dictionary of lists, drawn as a ladder with one rung per traced line
One line in place of four, with the first-sighting conditional gone. This is the same simplification get gave the histogram, applied to lists instead of counters.
Verify: Check that the new list persists.
Why: The new list is also added to the dictionary, so appending to what d[key] returned modifies what is stored — which is why the one-line version works. If the factory's result were merely returned, the append would go into a list nobody kept.
Trap
A program checks whether a key is present by reading d[key] on a defaultdict.
Read the value and see what is there
Why: Which is how you check a plain dictionary, via KeyError.
Reading a missing key calls the factory and stores the result, so the check creates the key it was asking about — and the dictionary grows every time something is looked up.
Use in to ask without creating.
if key in d
Why: Which does not invoke the factory.
Or use get, which also does not
Why: d.get(key) returns None for an absent key and adds nothing.
The difference from a plain dictionary is that reading has a side effect here. That is exactly the behaviour that makes the append idiom work, and it means a defaultdict cannot be inspected the way a dictionary can.
Faded example
The factory is a class object.
Fill in the blanks
from collections import defaultdict
d = defaultdict(list)
d['opst'].append('opts')
Why: The argument is list, which is a class object, not list(), which is a new list — the defaultdict calls it once per missing key, so it needs something callable. Passing a single list would share one list between every key, which is the mutable-default trap in another setting.
Sorting
Only a missing key does.
Sort into buckets
For each operation on a defaultdict, is the factory invoked?
Socratic
list rather than list().
Discussion prompt
Explain what would go wrong if defaultdict took a value rather than a function that produces one.
Hint: How many keys might be missing?
Answer:
Every missing key would get the same object. For an immutable default that would be harmless, and for a list it means appending under one key appears under all of them.
Passing a factory means the defaultdict can call it once per missing key, producing a fresh value each time — which is what makes a dictionary of independent lists possible.
It is the same reasoning as lesson 13b's mutable-default warning: a default value created once and shared is almost never what was meant. The fix in both cases is to create the value at the moment it is needed.
Section
Section 4
Concept
All three are dictionaries adapted to a particular job, and the question each time is what you need to store against a key.
And a plain dictionary when the values are arbitrary and every key is assigned deliberately — which is still the common case.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-188
Picture it
The answer picks the container.
Figure (svg): A decision flowchart choosing among a set, a Counter, a defaultdict and a dictionary
Every one of them is a dictionary underneath, which is why everything you know about keys — hashability, unpredictable order — applies to all four.
Worked example
Counting words, with each container.
# a set: which distinct words appear?
set(words)
# a Counter: how often does each appear?
Counter(words)
# a defaultdict: which lines did each appear on?
d = defaultdict(list)
d[word].append(line_number)| Container | What it answers | Note |
|---|---|---|
| a set | distinct words | no values |
| a Counter | words to counts | values are numbers |
| a defaultdict of lists | words to line numbers | values are collections |
Ask what is stored.
Why: Nothing, a number, or a growing collection — which is the whole of the decision.
Note that all three are keyed the same way.
Why: By word, with the same hashability requirement and the same unpredictable order.
Note the plain dictionary's place.
Why: For values that are neither counts nor accumulated — a translation, a definition, a configuration value.
Figure (svg): Two columns pairing each container with the chapter that built it by hand
Three containers for three questions about one collection of words. The choice is decided by what a key needs to carry.
Verify: Ask which of these chapter 13 built by hand.
Why: All three: a dictionary with None values for the set, the histogram for the Counter, and invert_dict's singleton-then-append for the defaultdict. Recognising them retrospectively is the point of this chapter — each container is an answer to something already done the long way.
Discrimination
What does each key need to carry?
Sort into buckets
For each job, which container fits best?
Worked example
Three specialisations, and the same constraints.
# all three:
# keys must be hashable
# order is unpredictable
# membership is fast
set([1, 2]) - set([2]) # works
set([[1], [2]]) # TypeError: unhashable| Property | Inherited? | Note |
|---|---|---|
| hashability | required | as for dictionary keys |
| order | unpredictable | likewise |
| speed | fast membership | the hashtable |
Note the hashability requirement.
Why: A set behaves like a collection of dictionary keys, and as in dictionaries, a Counter's keys have to be hashable.
Note the ordering.
Why: All three are unordered, for the reason chapter 11 gave: the storage location is computed from the key.
Note the speed.
Why: Adding elements to a set is fast; so is checking membership — which is the reason to use one.
Figure (svg): The state of the program after each line of Worked example what they inherit from dictionaries, drawn as a ladder with one rung per traced line
Three types with one set of constraints, all inherited from the dictionary. Nothing new has to be learned about keys.
Verify: Test the hashability rule on a set of lists.
Why: set([[1], [2]]) raises TypeError: unhashable type: 'list' — the same restriction and the same message as a list dictionary key. So chapter 11's argument about mutable keys applies unchanged, and a set of tuples works where a set of lists does not.
Trap
Every dictionary in a program is replaced by a defaultdict, in case a key is missing.
Avoid KeyErrors everywhere
Why: Which the default value does.
A KeyError is often the right outcome — it says a key that should exist does not. Replacing every dictionary means a typo in a key name silently creates an entry rather than raising, which loses a genuinely useful error.
Choose by what the values are.
A defaultdict when the values are accumulated
Why: Lists, counters, sets being built up.
A plain dictionary when a missing key is a mistake
Why: So the KeyError still tells you.
This is lesson 11a's get-against-brackets decision at the level of the container: absence is sometimes a legitimate state and sometimes a bug, and the choice should say which.
Prediction
A set of lists.
s = set([[1], [2]])| Aspect | What is true | Note |
|---|---|---|
| a set | behaves like dictionary keys | hashable required |
| a list | unhashable | as chapter 11 |
| the result | TypeError |
Predict first
What happens?
Correct: A TypeError — a set behaves like a collection of dictionary keys, so its elements must be hashable.
Why: The same restriction as a dictionary key, for the same reason: the storage location is computed from the element, so a mutable one could move after it was filed. A set of tuples works, which is chapter 12's escape route applying here too.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C contradicts what all three inherit from the dictionary. The storage location is computed from the key, so there is no order to preserve — which is exactly chapter 11's point, and it applies to sets, Counters and defaultdicts alike. If an order is needed, sorted gives one.
Real world
A structure that already does what you were about to write.
Discussion prompt
Think of a tool or a form that already handled something you were about to do by hand. What did using it remove?
Hint: Not effort — a class of mistake.
Answer:
Usually not just effort but a class of mistake: the hand-rolled version had edge cases, and the ready-made one had them handled by someone who thought about them for longer.
Which is what these three containers do. has_duplicates with a set cannot get the tracking wrong, because uniqueness is the container's property rather than the program's job.
And the reason to have written the long version first is that you can now recognise which container you need — the None values, the get-with-a-default, and the first-sighting conditional are each a signal that a specialised type is waiting.
Section
Section 5
Concept
Each of these types replaces a specific piece of dictionary boilerplate, and each piece of boilerplate is recognisable on sight.
Recognising the pattern is more useful than remembering the type, because the pattern is what appears in code you are reading or writing.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-188
Picture it
The boilerplate names the container.
Figure (svg): Three dictionary patterns paired with the container each one indicates
Which is deliberate on the book's part: the long version teaches the pattern, and the short version is only obvious once you have written the long one.
Worked example
Two functions, and what each is really asking for.
# signal: values are all True
def has_duplicates(t):
d = {}
for x in t:
if x in d: return True
d[x] = True
return False
# so: len(set(t)) < len(t)| Observation | What it indicates | Note |
|---|---|---|
| the values | all True | never read |
| what is wanted | uniqueness | a set's property |
| the rewrite | one comparison |
Look at the values.
Why: Every one is True, and nothing reads them — the same signal as chapter 13's Nones.
Ask what is actually wanted.
Why: Whether an element has been seen before, which is membership, and whether any repeats, which is uniqueness.
Name the container.
Why: Both are what a set provides, so the tracking code is unnecessary.
Figure (svg): The state of the program after each line of Worked example reading code for the signal, drawn as a ladder with one rung per traced line
A seven-line function whose values are decoration. Reading the values first is the fastest way to spot a set written as a dictionary.
Verify: Check whether the rewrite preserves the behaviour.
Why: It gives the same answer and does different work: the loop stops at the first duplicate and the set version builds everything first. For a long sequence with an early repeat the loop wins, so the rewrite is a simplification rather than an optimisation — which is worth being honest about.
Sorting
Read the values.
Sort into buckets
For each pattern, which container is being emulated?
Worked example
The plain dictionary is still the common case.
# a real mapping: keys to meaningful values
eng2sp = {'one': 'uno', 'two': 'dos'}
# no boilerplate to remove:
# the values are not counts
# they are not accumulated
# they are read| Aspect | What is true | Note |
|---|---|---|
| the values | meaningful | read by the program |
| the assignment | deliberate | one per key |
| no pattern | nothing to specialise |
Note what is absent.
Why: No repeated boilerplate, no first-sighting conditional, and values the program actually uses.
Note what a specialised type would add.
Why: A defaultdict would silently create entries on a typo; a Counter would answer a question nobody asked.
Conclude.
Why: The plain dictionary is right here, and it remains the common case.
Figure (svg): Two columns separating dictionaries worth specialising from those that should stay plain
A dictionary that should stay one. The specialised types exist for specific patterns, and most mappings do not exhibit any of them.
Verify: Ask what a defaultdict would cost here.
Why: A misspelled key would create an entry rather than raising KeyError — losing the error that says a key which should exist does not. That is lesson 11a's brackets-against-get decision: absence is a bug here, so the loud version is right.
Trap
A program's dictionaries are all replaced with defaultdicts and Counters after reading this section.
Use the better container
Why: The specialised types handle their cases well.
A container is only better for the pattern it was designed for. Applied elsewhere it changes behaviour — silently creating keys, or answering a question the program was not asking.
Convert where the boilerplate is.
Look for the three signals
Why: Values nobody reads, a counting idiom, an accumulate-per-key conditional.
And leave the rest
Why: Most dictionaries have no boilerplate to remove.
This is the same discipline as the previous lesson's: the point is to be able to choose. A specialised container adopted without its pattern is a change with no benefit and a behavioural cost.
Prediction
One detail says the container is wrong.
def subtract(d1, d2):
res = dict()
for key in d1:
if key not in d2:
res[key] = None
return res| Part | Its role | Note |
|---|---|---|
| the keys | used | the answer |
| the values | all None | never read |
| the signal | a set in disguise |
Predict first
What indicates that a set is wanted here?
Correct: Every value is None and nothing reads them — the dictionary is being used as a collection of keys.
Why: The book says so directly: in all of those dictionaries, the values are None because we never use them, and as a result we waste some storage space. A dictionary whose values are all one placeholder is almost always a set that has not been written as one.
Faded example
Uniqueness is the container's property.
Fill in the blanks
def has_duplicates(t):
return len(set(t)) < len(t)
Why: An element can only appear in a set once, so if an element in t appears more than once, the set will be smaller than t. The tracking dictionary is unnecessary because uniqueness is what a set is, rather than something the program has to maintain.
Explain it
Three containers and a plain dictionary.
Discussion prompt
A classmate asks how to choose between a set, a Counter, a defaultdict and an ordinary dictionary. Give them the question that decides it.
Hint: What goes against each key?
Answer:
Ask what each key needs to carry. Nothing means a set; a count means a Counter; something built up per key means a defaultdict; anything else means a plain dictionary.
And when reading existing code, look for the signals: values that are all None or all True, a get-with-a-default counting line, or a first-sighting conditional around an append. Each names a container.
But most dictionaries have none of those and should stay as they are — a defaultdict in place of a plain one turns a useful KeyError into a silently created entry, which is a real loss.
Comparison
Fill the blanks. All four are keyed the same way.
Comparison matrix
| Question | A set | A Counter | A defaultdict |
|---|---|---|---|
| What is stored per key? | nothing | a count | whatever the factory makes |
| A missing key? | not present | returns 0 | the factory runs and stores a value |
| Which pattern it replaces | a dictionary with None values | the get-and-increment histogram | the first-sighting conditional |
| What is inherited from dictionaries? | hashable keys, no order | hashable keys, no order | hashable keys, no order |
The bottom row is why nothing new has to be learned about keys: all three are dictionaries underneath.
Pattern
Five steps, and the last is the one that keeps a KeyError useful.
Step 4's class object, not a call is the detail that decides correctness: passing list() would share one list between every key, which is lesson 13b's mutable-default trap.
Python documentation — Data Structures Data Structures
Check
The set discards duplicates.
print(len(set('parrot')), len('parrot'))| Value | How many | Note |
|---|---|---|
| 'parrot' | six characters | with a repeated r |
| the set | p, a, r, o, t | five distinct |
| the comparison | 5 < 6 | a duplicate present |
Check your understanding
What does this print?
Answer: A
Why: An element can only appear in a set once, so the two r's become one and five distinct characters remain. That difference is exactly what has_duplicates tests: if the set is smaller than the sequence, something repeats.
Check
The element is absent.
Check your understanding
What does count['d'] give for a Counter built from 'parrot'?
Answer: A
Why: Unlike dictionaries, Counters don't raise an exception if you access an element that doesn't appear — instead, they return 0. That is the natural answer for a count, and it removes the get-with-a-default idiom that a plain dictionary histogram needed.
Check
The argument matters.
Check your understanding
Why is the argument list rather than list()?
Answer: A
Why: When you create a defaultdict, you provide a function that's used to create new values — a factory. Passing list() would supply a single list shared by every missing key, so appending under one key would appear under all of them: lesson 13b's mutable-default trap in a new setting.
Real world
Using a container that already enforces what you need.
Discussion prompt
Think of a form or a container that prevents a mistake by its own structure rather than by a rule someone has to follow. What does that buy?
Hint: A slot that only fits one way.
Answer:
A connector that only fits one way, a form field that accepts only a date, a box with one compartment per item — none of them needs a rule, because the shape does the enforcing.
What it buys is that the mistake becomes impossible rather than discouraged, so nobody has to remember and nobody has to check.
Which is what a set does for uniqueness. has_duplicates with a tracking dictionary can get the tracking wrong; with a set it cannot, because uniqueness is what the container is rather than what the program maintains.
Commit first
Answer, then rate your confidence.
Predict first
You write defaultdict(list()) instead of defaultdict(list). What goes wrong?
Correct: Every missing key shares one list, so appending under one key appears under all of them.
Why: The book flags this precisely: notice that the argument is list, which is a class object, not list(), which is a new list. A defaultdict needs a function it can call once per missing key — a factory — so that each key gets a fresh value. Passing an already-made list supplies one object to be shared by every key, which is lesson 13b's mutable-default trap in a new setting, and it produces the same bewildering symptom: data appearing under keys nothing has touched. The distinction is the same one chapter 15 drew between Point and Point(): one is the factory and the other is a product. And it matters here because the values are mutable — a shared immutable default would be harmless, which is why the trap is specific to lists, sets and dictionaries.
Explain it
A dictionary whose values are all None.
Discussion prompt
A classmate is reading chapter 13's subtract and asks why every value is None. Explain what that signals.
Hint: What is the dictionary actually being used for?
Answer:
That only the keys matter. The dictionary is providing fast membership testing and automatic uniqueness, and the values are there because a dictionary requires them — the book says so: the values are None because we never use them, and as a result we waste some storage space.
Which means the program is emulating a set: a type that behaves like a collection of dictionary keys with no values. Rewritten, the whole function is set(d1) - set(d2).
And the signal generalises. A dictionary whose values are all one placeholder is almost always a set that has not been written as one, which is worth recognising in code you read as well as code you write.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: Sets are the simplest of the three and the operators are worth knowing, since a subset test replaces a search loop. The Counter's zero for a missing element is its one real departure from a dictionary, and it is what makes the counting idiom unnecessary. The defaultdict's factory is a correctness detail rather than a stylistic one — list against list() decides whether the keys share a value. And the recognition skill is what makes any of it useful, because the patterns appear in code long before anyone names the container.
Connect it up
One page, from memory.
Draw it
Draw three boxes labelled set, Counter and defaultdict, and beside each write what is stored per key and what happens when a key is missing. Underneath, write the dictionary pattern each one replaces, naming the chapter where you wrote it by hand. Finally note what all three inherit from dictionaries, and write the one-line versions of subtract, has_duplicates and is_anagram.
Recap
Three pages, and three containers that were waiting behind the code you wrote.
| If you remember one thing | It is this |
|---|---|
| From sets | A dictionary whose values are all None is a set in disguise. |
| From Counters | A missing element gives 0, not a KeyError. |
| From most_common | It returns (value, frequency) — the reverse of the hand-built version. |
| From defaultdict | The argument is a factory, so each key gets its own value. |
| From all three | They are dictionaries underneath: hashable keys, and no order. |
The next lesson finishes the chapter with named tuples, which give a tuple's elements names and can replace a whole class, and the double star that gathers keyword arguments into a dictionary — plus the operator that scatters one back out.
Think Python, 2nd edition — Allen B. Downey §19.5-19.7, pp. 186-188 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.