This lesson analyses linear and bisection search, then builds a hashtable from a list of tuples in three steps to explain why dictionary operations are constant time.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Appendix B — Analysis of Algorithms
§B.3-B.4, pp. 205-208
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-208 — the pages these objectives are drawn from
Warm-up
Two lessons have now repeated it without proving it.
Discussion prompt
Lesson 11a said most dictionary operations, including in, are constant time, and the previous lesson called that 'one of the minor miracles of computer science'. What would have to be true of the implementation for that to work?
Hint: A linear search looks at items one at a time. What would let it look at fewer?
Answer:
Something would have to take you straight to the right place rather than making you look through everything — a way of computing where an item should be, from the item itself.
And that computation would have to cost the same whatever the size, or the saving would evaporate as the dictionary grew.
This lesson builds exactly that, in three steps, from a list of tuples. The book's word for the result is a hashtable, and Python dictionaries are implemented using hashtables.
Concept
A search is an algorithm that takes a collection and a target item and determines whether the target is in the collection, often returning the index of the target.
Three strategies and three orders of growth, and the third one is what a dictionary does.
Figure (svg): Three search strategies with their orders of growth
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205
Section
Section 1
Concept
The simplest search algorithm is a linear search, which traverses the items of the collection in order, stopping if it finds the target.
linear search — A search that traverses the items of a collection in order, stopping if it finds the target.
def linear_search(t, target):
for i, x in enumerate(t):
if x == target:
return i
return None| Situation | Work done | Note |
|---|---|---|
| target is first | 1 comparison | the best case |
| target is in the middle | about n/2 comparisons | the average |
| target is absent | n comparisons | the worst case |
In the worst case it has to traverse the entire collection, so the run time is linear. The in operator for sequences uses a linear search; so do string methods like find and count.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205
Picture it
The worst case is the whole collection.
Figure (svg): A sequence being examined element by element from the start
Which is why the in operator gets slower as a list grows, and why the same operator on a dictionary does not.
Worked example
In operators you have used all course.
'z' in ['a', 'b', 'c'] # linear
'ell' in 'hello world' # linear
'hello'.find('l') # linear
'banana'.count('a') # linear| Operation | What it does | Note |
|---|---|---|
| in, on a sequence | linear search | list, tuple or string |
| str.find | linear search | |
| str.count | linear | it cannot stop early |
Note the operator.
Why: The in operator for sequences uses a linear search, so its cost grows with the length of the sequence.
Note the string methods.
Why: So do string methods like find and count — the same traversal under a different name.
Note which can stop early.
Why: in and find stop at the first hit; count cannot, because it has to see them all.
Figure (svg): The state of the program after each line of Worked example where linear search is hiding, drawn as a ladder with one rung per traced line
Four familiar operations, all the same algorithm. Recognising it is what lets you predict which parts of a program will slow down as its data grows.
Verify: Check it against the word-list exercises.
Why: Chapter 13's has_duplicates and the word-search exercises used in on lists, and the ones that felt slow were the ones running that linear search once per word — a linear operation inside a linear loop, which the previous lesson's rule makes quadratic.
Prediction
For a linear search.
'z' in ['a', 'b', 'c', 'd', 'e']| Target | Comparisons | Note |
|---|---|---|
| 'a' | 1 comparison | found immediately |
| 'e' | 5 comparisons | found at the end |
| 'z' | 5 comparisons | not found at all |
Predict first
Which target makes this linear search do the most work?
Correct: 'z' or 'e' — in the worst case it has to traverse the entire collection.
Why: A linear search stops if it finds the target, so a hit near the front is cheap and a hit at the end or a miss is not. Analysis uses the worst case, which is why the in operator on a sequence is classified as linear rather than by its average.
Worked example
The pattern that makes programs quadratic.
# for each word, is it in the list?
found = []
for word in words:
if word in wordlist:
found.append(word)| Part | Its class | Note |
|---|---|---|
| the loop | n passes | linear |
| the body | a linear search | linear |
| together | quadratic | O(n squared) |
Count the loop.
Why: One pass per word in words, which is linear in the number of words.
Count the body.
Why: The in operator on a list is a linear search, so each pass is linear in the length of the wordlist.
Compose them.
Why: The previous lesson's rule applies: a linear body inside a linear loop is quadratic.
Figure (svg): Two columns comparing a list and a dictionary as the container searched inside a loop
Quadratic, from two innocent-looking lines. The fix is to make the inner search constant, which is what a dictionary or a set does.
Verify: Check what changing the container does.
Why: Making wordlist a dictionary or a set changes the inner search from linear to constant, which drops the whole loop to linear. Nothing about the loop changed — the improvement came entirely from the data structure, which is chapter 13's point stated in the previous lesson's vocabulary.
Trap
A membership test is treated as one step because it is one operator.
Count operators
Why: One line, one operation, one unit of cost.
The in operator for sequences uses a linear search, so it is n steps rather than one. Written inside a loop it silently makes the program quadratic.
Ask what the operator has to look at.
On a sequence it looks at every item
Why: Until it finds one, or runs out.
On a dictionary or a set it does not
Why: Because those are hashtables.
Which is why the same operator has two different costs, and why the previous lesson's operation table is worth knowing rather than deriving each time.
Sorting
Some of these are hashtable operations.
Sort into buckets
For each operation, is it a linear search?
Faded example
What makes a linear search linear.
Fill in the blanks
# in the worst case a linear search has to
# traverse the entire collection
Why: That is the case analysis uses, because the point of the classification is to give a guarantee rather than a typical figure. A miss is always the worst case, since absence can only be established by looking at everything.
Socratic
A hit can be cheap.
Discussion prompt
A linear search can find an item in one step. Why can it never report an absence in one step?
Hint: What would have to be checked?
Answer:
Because 'not here' is a claim about every item. One comparison tells you about one item, so establishing that none of them matches takes as many comparisons as there are items.
A hit can stop early because one matching item is enough to settle the question. There is no equivalent single piece of evidence for a miss.
Which is why the two structures that beat it change the question. A bisection search rules out half the collection with one comparison, and a hashtable computes where the item would be, so a miss is answered by looking in one place.
Section
Section 2
Concept
If the elements of the sequence are in order, you can use a bisection search, which is O(log n).
Bisection search is similar to the algorithm you might use to look a word up in a dictionary — a paper dictionary, not the data structure. If the sequence has 1,000,000 items, it will take about 20 steps to find the word or conclude that it is not there. So that is about 50,000 times faster than a linear search.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205
Picture it
Each step throws away half of what is left.
Figure (svg): The remaining search space halving on each step from a million down to one
And doubling the collection adds one step rather than doubling the work — which is what O(log n) means in practice.
Worked example
The analogy is exact.
# looking up 'python' in a paper dictionary
# open in the middle -> 'mango'
# 'python' comes after -> keep the second half
# open its middle -> 'salt'
# 'python' comes before -> keep the first half
# ...| Step | What you see | What you keep |
|---|---|---|
| open in the middle | 'mango' | keep the second half |
| open its middle | 'salt' | keep the first half |
| repeat | each time halving | about 20 steps |
Note what makes it work.
Why: The words are in order, so one comparison tells you which half the target must be in — that is the whole trick.
Note what each step buys.
Why: Either way, you cut the number of remaining items in half, which is why the number of steps grows with the logarithm rather than the size.
Note the scale.
Why: If the sequence has 1,000,000 items, it will take about 20 steps to find the word or conclude that it is not there.
Figure (svg): A growth chart comparing linear and logarithmic search costs
Twenty comparisons where a linear search would need up to a million. Nobody looks up a word by reading from 'aardvark', and this is why.
Verify: Check the 50,000 figure.
Why: A million divided by twenty is fifty thousand, which is the book's 'about 50,000 times faster'. And the ratio itself grows with the size — at ten million items the linear search takes ten times longer and the bisection search takes about one step more.
Prediction
A million takes about twenty.
# 1,000,000 items -> about 20 steps
# 2,000,000 items -> ?| Change | Bisection | Note |
|---|---|---|
| doubling the items | adds one halving | one more step |
| a linear search | would double | twice the work |
| that gap | widens with size |
Predict first
Roughly how many steps does a bisection search need for 2,000,000 items?
Correct: About 21 — each step cuts the number of remaining items in half, so doubling the collection adds one step.
Why: That is what O(log n) means concretely, and it is why the advantage over a linear search grows with the size: doubling the data doubles a linear search's work and adds one comparison to a bisection search's.
Worked example
It is not free.
# bisection search requires the sequence
# to be in order
t.sort() # O(n log n), once
# then many bisection searches, O(log n) each| Usage | The cost | Verdict |
|---|---|---|
| one search on unsorted data | sort then search | worse than linear |
| many searches | sort once, search often | worth it |
| data that keeps changing | re-sorting | may not be worth it |
Note the requirement.
Why: Bisection search can be much faster than linear search, but it requires the sequence to be in order, which might require extra work.
Price the extra work.
Why: Sorting is O(n log n), which is more than a single linear search — so for one lookup on unsorted data, sorting first is a loss.
Find where it pays.
Why: The sort is paid once and the searches are paid many times, so the more lookups you do, the better the trade.
Figure (svg): The state of the program after each line of Worked example what bisection costs you, drawn as a ladder with one rung per traced line
A trade rather than a free improvement, and its terms depend on how often you search. That is a different kind of question from order of growth alone.
Verify: Ask what a hashtable changes about the trade.
Why: It removes it: a hashtable can do a search in constant time and it doesn't require the items to be sorted. So there is no preparation to amortise and no re-sorting when the data changes — which is why the next idea is worth the three pages it takes.
Trap
A bisection search is applied to a list that has not been sorted.
Halve on each comparison
Why: The algorithm still runs and returns something.
It gives wrong answers silently. The comparison at the middle only tells you which half to keep if the sequence is in order; on unsorted data it discards the half containing the target as often as not.
Establish the precondition first.
Sort, or keep it sorted
Why: Which is where the extra work lives.
Then bisect
Why: And the halving argument holds.
An algorithm with a precondition is only as reliable as the precondition, and this one fails quietly rather than loudly — the worst kind of failure, and the reason lesson Ab spent a lesson on semantic errors.
Discrimination
Each has a precondition.
Sort into buckets
For each situation, which search is available?
Faded example
The precondition is the whole cost.
Fill in the blanks
# bisection search is O(log n), but it
# requires the sequence to be in order
Why: That requirement might require extra work, and sorting is O(n log n) — more than a single linear search. So bisection pays off when the sequence is already sorted, or when the sort can be amortised over many lookups.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C ignores the precondition. Bisection search can be much faster than linear search, but it requires the sequence to be in order, which might require extra work — and sorting costs more than the single linear search it would replace.
Section
Section 3
Concept
To explain how hashtables work and why their performance is so good, Downey starts with a simple implementation of a map and gradually improves it until it is a hashtable. For the rest of the lesson you have to imagine that dictionaries don't exist.
class LinearMap:
def __init__(self):
self.items = []
def add(self, k, v):
self.items.append((k, v))
def get(self, k):
for key, val in self.items:
if key == k:
return val
raise KeyError| Method | What it does | Its class |
|---|---|---|
| add | appends a key-value tuple | constant time |
| get | a for loop over the items | linear |
| on a miss | raises KeyError | after seeing everything |
The two operations to implement are add(k, v), which with a real dictionary is written d[k] = v, and get(k), which is written d[k]. add appends a key-value tuple to the list of items, which takes constant time; get uses a for loop to search the list, so get is linear.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 206-206
Picture it
The same items, differently arranged.
Figure (svg): Two columns comparing a single list of pairs with a hundred shorter lists
That is nice, and still not as good as a hashtable — the class has not changed, only the constant.
Worked example
A hundred maps, and a way to choose one.
class BetterMap:
def __init__(self, n=100):
self.maps = []
for i in range(n):
self.maps.append(LinearMap())
def find_map(self, k):
index = hash(k) % len(self.maps)
return self.maps[index]
def add(self, k, v):
m = self.find_map(k)
m.add(k, v)
def get(self, k):
m = self.find_map(k)
return m.get(k)| Method | What it does | Note |
|---|---|---|
| __init__ | makes a list of n LinearMaps | 100 by default |
| find_map | hash(k) % len(self.maps) | a legal index |
| add and get | delegate to one LinearMap | the chosen one |
Note what hash gives you.
Why: find_map uses the built-in function hash, which takes almost any Python object and returns an integer.
Note what the modulus does.
Why: It wraps the hash values into the range from 0 to len(self.maps), so the result is a legal index into the list.
Note the payoff.
Why: If the hash function spreads things out pretty evenly, we expect n/100 items per LinearMap — so BetterMap is about 100 times faster than LinearMap.
Figure (svg): A diagram showing keys distributed across a hundred smaller maps by their hash values
A hundredfold improvement that changes nothing about the class: the order of growth is still linear, but the leading coefficient is smaller.
Verify: Say why that is not enough.
Why: The previous lesson's closing line names the problem exactly: a constant factor is a fixed multiple at every size, so at ten million items BetterMap is doing a hundred thousand comparisons per lookup. A hundredfold is a lot and it is bounded, which is why the book says 'still not as good as a hashtable'.
Prediction
A hundred maps, n items.
# 100 LinearMaps, n items total
# expect n/100 items per LinearMap
# LinearMap.get is linear in its own size| Part | Cost | Note |
|---|---|---|
| find_map | constant | one hash and a modulus |
| the inner get | n/100 comparisons | linear in n |
| together | linear | smaller coefficient |
Predict first
What is the order of growth of BetterMap.get?
Correct: Linear — the order of growth is still linear, but the leading coefficient is smaller.
Why: Dividing by a hundred does not change the order of growth, because n/100 still grows in proportion to n. The number of maps is fixed while the number of items is not, so each map grows without limit — which is precisely what the next step fixes.
Worked example
A limitation the implementation exposes.
hash('apple') # fine
hash((1, 2)) # fine
hash([1, 2]) # TypeError: unhashable type: 'list'| Object | Hashable? | Why |
|---|---|---|
| a string | hashable | immutable |
| a tuple | hashable | immutable |
| a list | unhashable | mutable |
Note the limitation.
Why: A limitation of this implementation is that it only works with hashable keys. Mutable types like lists and dictionaries are unhashable.
Note the guarantee.
Why: Hashable objects that are considered equivalent return the same hash value — which is what makes find_map find the item again.
Note the non-guarantee.
Why: But the converse is not necessarily true: two objects with different values can return the same hash value.
Figure (svg): The state of the program after each line of Worked example what hashable means, drawn as a ladder with one rung per traced line
The rule from chapter 11 explained. Dictionary keys have to be hashable because the implementation computes an index from them, and a mutable object's index would go stale.
Verify: Connect it to chapter 12's tuples as keys.
Why: Lesson 12b used tuples as dictionary keys, and this is why it worked and a list would not have: a tuple is immutable, so its hash is stable, and an item stored under it stays findable. A list's contents can change after it is stored, and then hash(k) points somewhere else.
Trap
BetterMap is a hundred times faster, so the problem is treated as solved.
Measure the improvement
Why: A hundredfold is a large improvement.
The order of growth for get is still linear. As the number of items grows, each of the hundred LinearMaps grows with it, so the lookup cost still grows in proportion to n.
Ask what happens as n grows.
The number of maps is fixed at 100
Why: So each one grows without limit.
Which means the search inside one grows too
Why: n/100 is still linear in n.
The fix is to let the number of maps grow with the data, which is exactly the crucial idea of the next section — and the reason BetterMap is described as a step on the path rather than a destination.
Discrimination
Two classes, four methods.
Sort into buckets
For each method, what is its order of growth?
Faded example
A hash, wrapped into range.
Fill in the blanks
def find_map(self, k):
index = hash(k) % len(self.maps)
return self.maps[index]
Why: The modulus operator wraps the hash values into the range from 0 to len(self.maps), so the result is always a legal index. Many different hash values will wrap onto the same index, which is fine — that map holds a few items and is searched linearly.
Explain it
Chapter 11 stated the rule; this lesson explains it.
Discussion prompt
A classmate asks why a list cannot be a dictionary key. Answer using find_map.
Hint: Where does the item get stored?
Answer:
Because the position is computed from the key: index is hash(k) modulo the number of maps. The key decides where the item lives.
If the key can change after the item is stored, the computed position changes with it, and the item is no longer where the lookup will look. It would still be in the structure and permanently unfindable.
So mutable types like lists and dictionaries are unhashable — Python refuses to compute a hash for them rather than let you build a broken dictionary. Tuples are fine, which is why lesson 12b could use them as keys.
Section
Section 4
Concept
Here, finally, is the crucial idea that makes hashtables fast: if you can keep the maximum length of the LinearMaps bounded, LinearMap.get is constant time.
class HashMap:
def __init__(self):
self.maps = BetterMap(2)
self.num = 0
def get(self, k):
return self.maps.get(k)
def add(self, k, v):
if self.num == len(self.maps.maps):
self.resize()
self.maps.add(k, v)
self.num += 1| Method | What it does | Note |
|---|---|---|
| __init__ | a BetterMap of 2, and num = 0 | num counts the items |
| get | dispatches to BetterMap | no work of its own |
| add | resizes when num equals the map count | then adds |
All you have to do is keep track of the number of items, and when the number of items per LinearMap exceeds a threshold, resize the hashtable by adding more LinearMaps. Here the threshold is an average of one item per map.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 207-208
Picture it
Which keeps each one short.
Figure (svg): A chart showing the number of maps growing alongside the number of items so the ratio stays constant
And a search of a list with about one item in it takes about the same time however large the whole structure has become.
Worked example
You cannot just copy the maps across.
def resize(self):
new_maps = BetterMap(self.num * 2)
for m in self.maps.maps:
for k, v in m.items:
new_maps.add(k, v)
self.maps = new_maps| Step | What happens | Note |
|---|---|---|
| a new BetterMap | twice as big | geometric growth |
| every item | added again | rehashed |
| the old maps | discarded | replaced wholesale |
Note what resize builds.
Why: A new BetterMap, twice as big as the previous one, and then it rehashes the items from the old map to the new.
Note why rehashing is necessary.
Why: Changing the number of LinearMaps changes the denominator of the modulus operator in find_map — so an item's index is different in the new structure.
Note what that achieves.
Why: Some objects that used to hash into the same LinearMap will get split up, which is exactly what we wanted.
Figure (svg): A flowchart of what add does, showing the resize branch
Every item has to be re-placed because every item's address depends on the size of the table. That is the price of computing the location instead of storing it.
Verify: Check what happens if you skip the rehash.
Why: Items would sit at indices computed with the old denominator while lookups used the new one, so most lookups would search the wrong map and raise KeyError. The structure would be intact and wrong — a semantic error of exactly the kind lesson Ab described.
Prediction
Read the condition.
def add(self, k, v):
if self.num == len(self.maps.maps):
self.resize()
self.maps.add(k, v)
self.num += 1| Quantity | What it counts | Note |
|---|---|---|
| num | the number of items | |
| len(self.maps.maps) | the number of LinearMaps | |
| equal | one item per map on average | resize |
Predict first
What condition triggers a resize?
Correct: When they are equal, the average number of items per LinearMap is 1, so it calls resize.
Why: The threshold is about the average rather than any individual map, which is what makes it cheap to check — num is a single counter. Individual maps may hold two or three items, and the average is what the bound is stated over.
Worked example
A bounded list is a constant-time search.
# LinearMap.get is linear in ITS OWN length
#
# BetterMap: 100 maps, n items
# -> each map holds n/100 -> grows with n
#
# HashMap: maps grow with n
# -> each map holds about 1 -> bounded| Structure | Length of one map | get |
|---|---|---|
| BetterMap | map length grows with n | linear get |
| HashMap | map length stays about 1 | constant get |
| the difference | whether the map count is fixed |
State what get actually costs.
Why: LinearMap.get is linear in the length of its own list, and nothing else — not in the total number of items.
Note what BetterMap failed to do.
Why: It divided that length by a hundred and left it growing, because the number of maps was fixed.
Note what HashMap does.
Why: It keeps the maximum length of the LinearMaps bounded, and then LinearMap.get is constant time.
Figure (svg): The state of the program after each line of Worked example why the maximum length is the whole thing, drawn as a ladder with one rung per traced line
The class changed because a growing quantity was made bounded. That is a different kind of move from making something a hundred times smaller, and it is why one gives an unbounded improvement and the other does not.
Verify: Say what the hash function has to do for this to hold.
Why: Spread things out pretty evenly, which is what hash functions are designed to do. If every key hashed to the same index, one map would hold everything and get would be linear again — the bound is an expectation about typical behaviour rather than a guarantee.
Trap
The speed is attributed to the hash function, since that is the distinctive part.
Compute an index instead of searching
Why: One calculation replaces a traversal.
BetterMap does exactly that and is still linear. The hash chooses a map; what makes the lookup constant is that the chosen map is short, and it is short only because the number of maps grows.
Attribute it to the bound.
The hash distributes items evenly
Why: Necessary, and not sufficient.
The resizing keeps each map short
Why: Which is the crucial idea.
Both parts are needed: an even spread over a fixed number of buckets gives a smaller coefficient, and a growing number of buckets with a bad hash gives no improvement at all.
Faded example
One word carries it.
Fill in the blanks
# if you can keep the maximum length of the
# LinearMaps bounded, LinearMap.get is constant time
Why: That is the whole difference between BetterMap and HashMap. Dividing a growing length by a hundred leaves it growing; holding it below a fixed ceiling makes the search inside it constant, which changes the order of growth rather than the coefficient.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C ignores the rehash. Rehashing is linear, so resize is linear — every item has to be added to the new structure. What makes add constant on average is not that resize is cheap but that it happens rarely enough for its cost to average out.
Socratic
HashMap begins with BetterMap(2).
Discussion prompt
The implementation starts with two LinearMaps rather than a hundred. Why is that reasonable?
Hint: What decides how many there should be?
Answer:
Because the right number depends on how many items there are, and at the start there are none. Two costs nothing to allocate and is enough for the first two adds.
The structure then grows to fit: each resize doubles it, so after a handful of resizes it is as large as it needs to be. Guessing a hundred in advance would waste space on small maps and still be wrong for large ones.
Which is the difference from BetterMap in one sentence. BetterMap picks a size once and lives with it; HashMap lets the data decide, and that is what keeps the items-per-map ratio bounded.
Section
Section 5
Concept
Rehashing is linear, so resize is linear, which might seem bad — since add was promised to be constant time. But we don't have to resize every time, so add is usually constant time and only occasionally linear.
# starting empty, with 2 LinearMaps:
# adds 1-2 1 unit each total 2
# add 3 resize (2) + add (1)
# add 4 1 unit total 6
# add 5 5 units
# adds 6-8 1 unit each total 14
# add 9 9 units total 30 by 16
# by 32 adds total 62| Adds | Total work | Note |
|---|---|---|
| 4 adds | 6 units | |
| 8 adds | 14 units | |
| 16 adds | 30 units | |
| 32 adds | 62 units | 2n minus 2 |
After n adds, where n is a power of two, the total cost is 2n minus 2 units, so the average work per add is a little less than 2 units. For other values of n the average is a little higher, and the important thing is that it is O(1).
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 208-208
Picture it
The book's own figure, described.
Figure (svg): The cost of successive adds, showing tall rare spikes among many cheap ones
Now if you knock over the towers, spreading the cost of resizing over all adds, you can see graphically that the total cost after n adds is 2n minus 2.
Worked example
Follow the units.
# adds 1, 2 cheap 2 units
# add 3 rehash 2 + add 1 = 3 5
# add 4 1 6 (4 items)
# add 5 rehash 4 + add 1 = 5 11
# adds 6, 7, 8 1 each 14 (8 items)
# add 9 rehash 8 + add 1 = 9 23
# adds 10-16 1 each 30 (16 items)| After | Total | Average |
|---|---|---|
| 4 items | 6 units | 1.5 per add |
| 8 items | 14 units | 1.75 per add |
| 16 items | 30 units | 1.875 per add |
| n items | 2n minus 2 | under 2 per add |
Note where the spikes are.
Why: A resize happens at add 3, 5, 9, 17 — each time the table doubles, so the gaps between resizes double as well.
Note how big each spike is.
Why: A resize rehashes everything currently stored, so its cost is the number of items — which is also the number of cheap adds since the last resize.
Read the pattern.
Why: After n adds, where n is a power of two, the total cost is 2n minus 2 units, so the average work per add is a little less than 2 units.
Figure (svg): A chart contrasting the total work with the average work per add as the number of adds grows
Under two units per add, forever. The average does not creep upward as the structure grows, which is what makes it O(1) rather than merely small.
Verify: Check the trend in the average column.
Why: 1.5, then 1.75, then 1.875 — rising, and toward 2 rather than without limit. A cost that approaches a fixed ceiling is constant time; one that keeps climbing would not be.
Prediction
The pattern is in the table.
# 4 adds -> 6 units
# 8 adds -> 14 units
# 16 adds -> 30 units
# 32 adds -> ?| n | Total | The rule |
|---|---|---|
| 4 | 6 | 2n minus 2 |
| 8 | 14 | 2n minus 2 |
| 16 | 30 | 2n minus 2 |
Predict first
What is the total cost after 32 adds?
Correct: 62 units — after 32 adds, the total cost is 62 units, and after n adds where n is a power of two the total cost is 2n minus 2.
Why: Doubling the number of adds roughly doubles the total, which is what a linear total looks like. Dividing gives an average of a little under two units per add, and it stays under two however large n gets — that is the O(1) claim.
Worked example
Doubling rather than adding.
# geometric: BetterMap(self.num * 2)
# resizes at 2, 4, 8, 16, 32 ...
# -> average add is constant
# arithmetic: BetterMap(self.num + 10)
# resizes at 10, 20, 30, 40 ...
# -> average add is LINEAR| Growth | What happens | Note |
|---|---|---|
| geometric | gaps double as costs double | they cancel |
| arithmetic | gaps stay fixed as costs grow | they do not |
| the consequence | the growth rule is load-bearing |
Name the property.
Why: An important feature of this algorithm is that when we resize the HashTable it grows geometrically; that is, we multiply the size by a constant.
Note the alternative.
Why: If you increase the size arithmetically — adding a fixed number each time — the average time per add is linear.
See why.
Why: With arithmetic growth the resizes keep costing more while staying equally frequent, so their total cost outgrows the number of adds.
Figure (svg): The state of the program after each line of Worked example why the growth must be geometric, drawn as a ladder with one rung per traced line
A one-word change in resize that changes the order of growth of the whole structure. Doubling is not an arbitrary choice.
Verify: Connect it back to lists.
Why: The previous lesson said adding to the end of a list is constant on average because when it runs out of room it occasionally gets copied to a bigger location. That is this same argument, and Python's list uses geometric growth for exactly this reason.
Trap
add is called linear because one call in it can rehash everything.
Analyse the worst case
Why: Which the previous lesson said to do.
For a single call that is right, and for a sequence of calls it is misleading. The total amount of work to run add n times is proportional to n, so the average time of each add is constant.
Say which claim you are making.
A single add is occasionally linear
Why: True, and the worst case.
n adds cost O(n) in total
Why: So the average is constant.
Both statements are correct about different things. The second is the one that matters for a program that does many adds, which is nearly every program — and it is why the previous lesson's table says 'constant on average' rather than 'constant'.
Sorting
In the structures the lesson built.
Sort into buckets
For each operation, is it constant time?
Faded example
One word decides the class.
Fill in the blanks
# resizing must grow the table geometrically
# - multiplying the size by a constant.
# Adding a fixed number each time makes
# the average add linear.
Why: Doubling makes the gaps between resizes grow at the same rate as the cost of a resize, so the two cancel. Adding a fixed amount each time leaves resizes equally frequent while they get steadily more expensive, and the average climbs without limit.
Real world
Something done occasionally to keep everything else cheap.
Discussion prompt
Think of something rarely done that is a lot of work at once, and that keeps the everyday version quick. What makes it worth it?
Hint: How often, against how much?
Answer:
Reorganising a filing system, restocking a shelf, moving to a bigger space — each costs a great deal on the day and nothing on most days.
What makes it worth it is the ratio: if the big job gets twice as large but happens half as often, the average effort per day does not change. If it gets larger without getting rarer, it eventually dominates everything.
Which is exactly the geometric-growth argument. Doubling makes the resizes twice as expensive and twice as far apart, so the average add stays under two units — and arithmetic growth, which does not, makes it linear.
Comparison
Fill the blanks. One interface, three implementations.
Comparison matrix
| Question | LinearMap | BetterMap | HashMap |
|---|---|---|---|
| What holds the items? | one list of tuples | 100 LinearMaps | a BetterMap that resizes |
| How does get find the item? | a for loop over everything | hash picks a map, then a for loop | the same, over a short list |
| Order of growth of get | linear | linear, smaller coefficient | constant |
| Does the map count grow with the data? | there is only one | no — fixed at 100 | yes — it doubles |
The bottom row is the whole difference. Everything else about BetterMap and HashMap is the same code.
Pattern
Five questions, and the fourth is the one people skip.
Step 4 is BetterMap against HashMap, and step 5 is why both list append and dictionary add are described as constant on average rather than constant.
Python documentation — Built-in Types Built-in Types
Check
One requires a precondition.
Check your understanding
Which search is O(log n), and what does it require?
Answer: A
Why: Bisection search cuts the number of remaining items in half on each step, which is what makes it logarithmic — about 20 steps for a million items. It only works because one comparison at the middle tells you which half the target is in, and that is true only if the sequence is in order.
Check
A hundredfold improvement.
Check your understanding
BetterMap is about 100 times faster than LinearMap. What is its order of growth?
Answer: A
Why: The number of maps is fixed at a hundred while the number of items is not, so each LinearMap holds about n/100 items and grows without limit. Dividing by a constant does not change the order of growth — which is exactly the previous lesson's distinction between a constant factor and a better class.
Check
Occasionally linear.
Check your understanding
Why is HashMap.add constant time on average?
Answer: A
Why: After n adds the total cost is about 2n units, so the average is a little under two per add and stays there however large n becomes. The resizes get more expensive as the structure grows and correspondingly rarer, and geometric growth is what makes those two effects cancel.
Real world
Finding something without looking through everything.
Discussion prompt
Think of a place where things are stored so that you can go straight to the one you want. What decides where each thing goes, and what happens when the space runs out?
Hint: Shelves, pigeonholes, an index.
Answer:
Something about the item itself decides its place — the first letter of a name, a number, a category — so finding it means computing the location rather than searching for it.
When the space runs out, everything has to be redistributed over a larger set of places, and the rule that assigned them has to change with it. That redistribution is expensive and it is rare.
Both halves are the hashtable: find_map computes an index from the key, and resize doubles the space and rehashes everything because the assignment rule now has a different denominator.
Commit first
Answer, then rate your confidence.
Predict first
BetterMap splits the items across a hundred LinearMaps and is about a hundred times faster. Is its get constant time?
Correct: No — the order of growth is still linear, but the leading coefficient is smaller.
Why: The hash does go straight to the right map, in constant time, and then a linear search runs inside that map. Each map holds about n/100 items, and a hundredth of a growing quantity still grows in proportion to n — so the lookup cost still rises as items are added. This is the previous lesson's distinction in its sharpest form: a hundredfold speed-up is a constant factor, fixed at every size, and it does not change the class. What changes the class is the crucial idea of the next section: if you can keep the maximum length of the LinearMaps bounded, LinearMap.get is constant time. HashMap does that by counting the items and doubling the number of maps whenever the average reaches one per map, so the length being searched stops growing at all. A better coefficient and a better class look similar at one size and behave completely differently as the data grows.
Explain it
The question chapter 11 left open.
Discussion prompt
A classmate asks how a dictionary can look something up in the same time whether it holds ten items or ten million. Explain it in three sentences.
Hint: Compute, bound, resize.
Answer:
The position is computed from the key rather than searched for: hash(k) modulo the number of buckets gives an index, and that calculation costs the same however many items there are.
Each bucket holds a short list, which is searched linearly — and it stays short because the number of buckets grows with the number of items, roughly one item per bucket.
Growing means doubling and rehashing everything, which is linear but rare enough that the average add stays under two units of work. That is the minor miracle: computing the location, bounding the bucket, and paying for the growth in rare instalments.
Exit ticket
One honest answer. The course ends here, and this is for you.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The two searches are worth knowing as a pair, because the contrast between them is what a precondition buys. BetterMap is the most instructive failure in the book: a hundredfold improvement that changes nothing about the class. The crucial idea is one sentence and it is the whole of chapter B.4. And the averaging argument is the one that also explains list append, so it is worth being able to reconstruct rather than recall.
Connect it up
One page, from memory.
Draw it
Draw the three search strategies with their orders of growth and what each requires. Underneath, draw the three maps in order — one list of tuples, a hundred lists chosen by a hash, and a growing number of lists — and write the order of growth of get beside each. Mark the one place where the class changes, and write the sentence that makes it change. Finish with the accounting: 4 adds, 8 adds, 16, 32, and the formula.
Recap
Four pages, and the answer to a question the course has carried since chapter 11.
| If you remember one thing | It is this |
|---|---|
| From linear search | Absence can only be established by looking at everything. |
| From bisection | A precondition can buy you a better order of growth. |
| From BetterMap | A hundredfold improvement is still a constant factor. |
| From HashMap | Bounding a growing quantity is what changes the class. |
| From the accounting | Rare and expensive averages away if the rarity grows with the cost. |
And Downey's own last word on the code: you wouldn't write this in Python — if you want a map, just use a dictionary. The value of building one is knowing what the dictionary is doing, which is what makes it possible to predict when it will be fast and when a list will not.
Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-208 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.