Bb Search Algorithms and Hashtables

This lesson analyses linear and bisection search, then builds a hashtable from a list of tuples in three steps to explain why dictionary operations are constant time.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson Bb Search Algorithms and Hashtables

Title

Python · Appendix B — Analysis of Algorithms

§B.3-B.4, pp. 205-208

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-208 — the pages these objectives are drawn from

3. Before we start: the claim from chapter 11

Warm-up

Two lessons have now repeated it without proving it.

Discussion prompt

Lesson 11a said most dictionary operations, including in, are constant time, and the previous lesson called that 'one of the minor miracles of computer science'. What would have to be true of the implementation for that to work?

Hint: A linear search looks at items one at a time. What would let it look at fewer?

Answer:

Something would have to take you straight to the right place rather than making you look through everything — a way of computing where an item should be, from the item itself.

And that computation would have to cost the same whatever the size, or the saving would evaporate as the dictionary grew.

This lesson builds exactly that, in three steps, from a list of tuples. The book's word for the result is a hashtable, and Python dictionaries are implemented using hashtables.

4. The one idea behind this lesson: compute the location

Concept

A search is an algorithm that takes a collection and a target item and determines whether the target is in the collection, often returning the index of the target.

Three strategies and three orders of growth, and the third one is what a dictionary does.

Figure (svg): Three search strategies with their orders of growth

Each row does strictly less work than the one above, and asks for something in return.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205

5. Linear search, and what uses it

Section

Section 1

6. Traverse in order, stop when you find it

Concept

The simplest search algorithm is a linear search, which traverses the items of the collection in order, stopping if it finds the target.

linear search — A search that traverses the items of a collection in order, stopping if it finds the target.

def linear_search(t, target):
    for i, x in enumerate(t):
        if x == target:
            return i
    return None
SituationWork doneNote
target is first1 comparisonthe best case
target is in the middleabout n/2 comparisonsthe average
target is absentn comparisonsthe worst case

In the worst case it has to traverse the entire collection, so the run time is linear. The in operator for sequences uses a linear search; so do string methods like find and count.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205

7. Picture it: one at a time until the end

Picture it

The worst case is the whole collection.

Figure (svg): A sequence being examined element by element from the start

Absence is the expensive answer: it can only be established by looking at everything.

Which is why the in operator gets slower as a list grows, and why the same operator on a dictionary does not.

8. Worked example: where linear search is hiding

Worked example

In operators you have used all course.

'z' in ['a', 'b', 'c']      # linear
'ell' in 'hello world'      # linear
'hello'.find('l')           # linear
'banana'.count('a')         # linear
OperationWhat it doesNote
in, on a sequencelinear searchlist, tuple or string
str.findlinear search
str.countlinearit cannot stop early

Note the operator.

Why: The in operator for sequences uses a linear search, so its cost grows with the length of the sequence.

Note the string methods.

Why: So do string methods like find and count — the same traversal under a different name.

Note which can stop early.

Why: in and find stop at the first hit; count cannot, because it has to see them all.

Figure (svg): The state of the program after each line of Worked example where linear search is hiding, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Four familiar operations, all the same algorithm. Recognising it is what lets you predict which parts of a program will slow down as its data grows.

Verify: Check it against the word-list exercises.

Why: Chapter 13's has_duplicates and the word-search exercises used in on lists, and the ones that felt slow were the ones running that linear search once per word — a linear operation inside a linear loop, which the previous lesson's rule makes quadratic.

9. Predict: which case is the worst?

Prediction

For a linear search.

'z' in ['a', 'b', 'c', 'd', 'e']
TargetComparisonsNote
'a'1 comparisonfound immediately
'e'5 comparisonsfound at the end
'z'5 comparisonsnot found at all

Predict first

Which target makes this linear search do the most work?

  • 'z' or 'e' — both require traversing the entire collection
  • 'a', because it is checked first
  • 'c', because it is in the middle
  • They all cost the same

Correct: 'z' or 'e' — in the worst case it has to traverse the entire collection.

Why: A linear search stops if it finds the target, so a hit near the front is cheap and a hit at the end or a miss is not. Analysis uses the worst case, which is why the in operator on a sequence is classified as linear rather than by its average.

10. Worked example: the linear search inside a loop

Worked example

The pattern that makes programs quadratic.

# for each word, is it in the list?
found = []
for word in words:
    if word in wordlist:
        found.append(word)
PartIts classNote
the loopn passeslinear
the bodya linear searchlinear
togetherquadraticO(n squared)

Count the loop.

Why: One pass per word in words, which is linear in the number of words.

Count the body.

Why: The in operator on a list is a linear search, so each pass is linear in the length of the wordlist.

Compose them.

Why: The previous lesson's rule applies: a linear body inside a linear loop is quadratic.

Figure (svg): Two columns comparing a list and a dictionary as the container searched inside a loop

The same five lines, and an unbounded difference in behaviour as the data grows.

Quadratic, from two innocent-looking lines. The fix is to make the inner search constant, which is what a dictionary or a set does.

Verify: Check what changing the container does.

Why: Making wordlist a dictionary or a set changes the inner search from linear to constant, which drops the whole loop to linear. Nothing about the loop changed — the improvement came entirely from the data structure, which is chapter 13's point stated in the previous lesson's vocabulary.

11. Trap: reading in as a single cheap operation

Trap

The trap

A membership test is treated as one step because it is one operator.

Count operators

Why: One line, one operation, one unit of cost.

The in operator for sequences uses a linear search, so it is n steps rather than one. Written inside a loop it silently makes the program quadratic.

The fix

Ask what the operator has to look at.

On a sequence it looks at every item

Why: Until it finds one, or runs out.

On a dictionary or a set it does not

Why: Because those are hashtables.

Which is why the same operator has two different costs, and why the previous lesson's operation table is worth knowing rather than deriving each time.

12. Sort: does this use a linear search?

Sorting

Some of these are hashtable operations.

Sort into buckets

For each operation, is it a linear search?

a linear search
'z' in ['a', 'b', 'c']; 'hello'.find('l'); 'banana'.count('a'); 'z' in ('a', 'b', 'c')
not a linear search
'z' in {'a': 1, 'b': 2}; d['key']
yes
The in operator for sequences uses a linear search, and so do string methods like find and count — a tuple is a sequence too, so it searches the same way a list does.
no
Both are dictionary operations, and Python dictionaries are implemented using hashtables, which is why most dictionary operations, including the in operator, are constant time.

13. Complete it: the worst case

Faded example

What makes a linear search linear.

Fill in the blanks

# in the worst case a linear search has to
# traverse the entire collection

Why: That is the case analysis uses, because the point of the classification is to give a guarantee rather than a typical figure. A miss is always the worst case, since absence can only be established by looking at everything.

14. Think it through: why is a miss the expensive answer?

Socratic

A hit can be cheap.

Discussion prompt

A linear search can find an item in one step. Why can it never report an absence in one step?

Hint: What would have to be checked?

Answer:

Because 'not here' is a claim about every item. One comparison tells you about one item, so establishing that none of them matches takes as many comparisons as there are items.

A hit can stop early because one matching item is enough to settle the question. There is no equivalent single piece of evidence for a miss.

Which is why the two structures that beat it change the question. A bisection search rules out half the collection with one comparison, and a hashtable computes where the item would be, so a miss is answered by looking in one place.

15. Bisection search

Section

Section 2

16. Halve the remaining items each time

Concept

If the elements of the sequence are in order, you can use a bisection search, which is O(log n).

Bisection search is similar to the algorithm you might use to look a word up in a dictionary — a paper dictionary, not the data structure. If the sequence has 1,000,000 items, it will take about 20 steps to find the word or conclude that it is not there. So that is about 50,000 times faster than a linear search.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-205

17. Picture it: a million items in twenty steps

Picture it

Each step throws away half of what is left.

Figure (svg): The remaining search space halving on each step from a million down to one

Twenty steps against a million, which is about 50,000 times faster than a linear search.

And doubling the collection adds one step rather than doubling the work — which is what O(log n) means in practice.

18. Worked example: the paper dictionary

Worked example

The analogy is exact.

# looking up 'python' in a paper dictionary
#   open in the middle -> 'mango'
#   'python' comes after -> keep the second half
#   open its middle -> 'salt'
#   'python' comes before -> keep the first half
#   ...
StepWhat you seeWhat you keep
open in the middle'mango'keep the second half
open its middle'salt'keep the first half
repeateach time halvingabout 20 steps

Note what makes it work.

Why: The words are in order, so one comparison tells you which half the target must be in — that is the whole trick.

Note what each step buys.

Why: Either way, you cut the number of remaining items in half, which is why the number of steps grows with the logarithm rather than the size.

Note the scale.

Why: If the sequence has 1,000,000 items, it will take about 20 steps to find the word or conclude that it is not there.

Figure (svg): A growth chart comparing linear and logarithmic search costs

The lower curve is almost flat, which is what a logarithm looks like at this scale.

Twenty comparisons where a linear search would need up to a million. Nobody looks up a word by reading from 'aardvark', and this is why.

Verify: Check the 50,000 figure.

Why: A million divided by twenty is fifty thousand, which is the book's 'about 50,000 times faster'. And the ratio itself grows with the size — at ten million items the linear search takes ten times longer and the bisection search takes about one step more.

19. Predict: how many steps for two million?

Prediction

A million takes about twenty.

# 1,000,000 items -> about 20 steps
# 2,000,000 items -> ?
ChangeBisectionNote
doubling the itemsadds one halvingone more step
a linear searchwould doubletwice the work
that gapwidens with size

Predict first

Roughly how many steps does a bisection search need for 2,000,000 items?

  • About 21 — doubling the collection adds a single step
  • About 40 — the work doubles
  • About 20 — exactly the same
  • About 2,000,000

Correct: About 21 — each step cuts the number of remaining items in half, so doubling the collection adds one step.

Why: That is what O(log n) means concretely, and it is why the advantage over a linear search grows with the size: doubling the data doubles a linear search's work and adds one comparison to a bisection search's.

20. Worked example: what bisection costs you

Worked example

It is not free.

# bisection search requires the sequence
# to be in order

t.sort()          # O(n log n), once
# then many bisection searches, O(log n) each
UsageThe costVerdict
one search on unsorted datasort then searchworse than linear
many searchessort once, search oftenworth it
data that keeps changingre-sortingmay not be worth it

Note the requirement.

Why: Bisection search can be much faster than linear search, but it requires the sequence to be in order, which might require extra work.

Price the extra work.

Why: Sorting is O(n log n), which is more than a single linear search — so for one lookup on unsorted data, sorting first is a loss.

Find where it pays.

Why: The sort is paid once and the searches are paid many times, so the more lookups you do, the better the trade.

Figure (svg): The state of the program after each line of Worked example what bisection costs you, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A trade rather than a free improvement, and its terms depend on how often you search. That is a different kind of question from order of growth alone.

Verify: Ask what a hashtable changes about the trade.

Why: It removes it: a hashtable can do a search in constant time and it doesn't require the items to be sorted. So there is no preparation to amortise and no re-sorting when the data changes — which is why the next idea is worth the three pages it takes.

21. Trap: bisecting an unsorted sequence

Trap

The trap

A bisection search is applied to a list that has not been sorted.

Halve on each comparison

Why: The algorithm still runs and returns something.

It gives wrong answers silently. The comparison at the middle only tells you which half to keep if the sequence is in order; on unsorted data it discards the half containing the target as often as not.

The fix

Establish the precondition first.

Sort, or keep it sorted

Why: Which is where the extra work lives.

Then bisect

Why: And the halving argument holds.

An algorithm with a precondition is only as reliable as the precondition, and this one fails quietly rather than loudly — the worst kind of failure, and the reason lesson Ab spent a lesson on semantic errors.

22. Discriminate: which search applies?

Discrimination

Each has a precondition.

Sort into buckets

For each situation, which search is available?

linear search
an unsorted list, one lookup; a string, looking for a substring
bisection search
a sorted list, many lookups; a sorted list of a million words
a hashtable
a dictionary; a set
lin
Neither is in order, so there is nothing to bisect — and the in operator for sequences and string methods like find both use a linear search.
bis
Both are in order, which is the precondition, and with many lookups the sorting cost is amortised over all of them.
hash
Both are implemented using hashtables, which do a search in constant time and don't require the items to be sorted.

23. Complete it: what bisection needs

Faded example

The precondition is the whole cost.

Fill in the blanks

# bisection search is O(log n), but it
# requires the sequence to be in order

Why: That requirement might require extra work, and sorting is O(n log n) — more than a single linear search. So bisection pays off when the sequence is already sorted, or when the sort can be amortised over many lookups.

24. Two truths and a lie: bisection search

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. About 20 steps suffice to search a million sorted items
  • B. It works by cutting the remaining items in half at each step
  • C. It is always faster than a linear search

Survives elimination: C

Why: C ignores the precondition. Bisection search can be much faster than linear search, but it requires the sequence to be in order, which might require extra work — and sorting costs more than the single linear search it would replace.

25. LinearMap, then BetterMap

Section

Section 3

26. Two implementations of the same interface

Concept

To explain how hashtables work and why their performance is so good, Downey starts with a simple implementation of a map and gradually improves it until it is a hashtable. For the rest of the lesson you have to imagine that dictionaries don't exist.

class LinearMap:

    def __init__(self):
        self.items = []

    def add(self, k, v):
        self.items.append((k, v))

    def get(self, k):
        for key, val in self.items:
            if key == k:
                return val
        raise KeyError
MethodWhat it doesIts class
addappends a key-value tupleconstant time
geta for loop over the itemslinear
on a missraises KeyErrorafter seeing everything

The two operations to implement are add(k, v), which with a real dictionary is written d[k] = v, and get(k), which is written d[k]. add appends a key-value tuple to the list of items, which takes constant time; get uses a for loop to search the list, so get is linear.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 206-206

27. Picture it: one long list against a hundred short ones

Picture it

The same items, differently arranged.

Figure (svg): Two columns comparing a single list of pairs with a hundred shorter lists

About 100 times faster, and the same order of growth — which is the previous lesson's distinction made concrete.

That is nice, and still not as good as a hashtable — the class has not changed, only the constant.

28. Worked example: BetterMap and the hash function

Worked example

A hundred maps, and a way to choose one.

class BetterMap:

    def __init__(self, n=100):
        self.maps = []
        for i in range(n):
            self.maps.append(LinearMap())

    def find_map(self, k):
        index = hash(k) % len(self.maps)
        return self.maps[index]

    def add(self, k, v):
        m = self.find_map(k)
        m.add(k, v)

    def get(self, k):
        m = self.find_map(k)
        return m.get(k)
MethodWhat it doesNote
__init__makes a list of n LinearMaps100 by default
find_maphash(k) % len(self.maps)a legal index
add and getdelegate to one LinearMapthe chosen one

Note what hash gives you.

Why: find_map uses the built-in function hash, which takes almost any Python object and returns an integer.

Note what the modulus does.

Why: It wraps the hash values into the range from 0 to len(self.maps), so the result is a legal index into the list.

Note the payoff.

Why: If the hash function spreads things out pretty evenly, we expect n/100 items per LinearMap — so BetterMap is about 100 times faster than LinearMap.

Figure (svg): A diagram showing keys distributed across a hundred smaller maps by their hash values

Two keys landing in the same map is expected and harmless — that map is searched linearly, and it is short.

A hundredfold improvement that changes nothing about the class: the order of growth is still linear, but the leading coefficient is smaller.

Verify: Say why that is not enough.

Why: The previous lesson's closing line names the problem exactly: a constant factor is a fixed multiple at every size, so at ten million items BetterMap is doing a hundred thousand comparisons per lookup. A hundredfold is a lot and it is bounded, which is why the book says 'still not as good as a hashtable'.

29. Predict: what class is BetterMap.get?

Prediction

A hundred maps, n items.

# 100 LinearMaps, n items total
# expect n/100 items per LinearMap
# LinearMap.get is linear in its own size
PartCostNote
find_mapconstantone hash and a modulus
the inner getn/100 comparisonslinear in n
togetherlinearsmaller coefficient

Predict first

What is the order of growth of BetterMap.get?

  • Linear — the same class as before, with a smaller leading coefficient
  • Constant, because the hash finds the map directly
  • Logarithmic
  • Quadratic

Correct: Linear — the order of growth is still linear, but the leading coefficient is smaller.

Why: Dividing by a hundred does not change the order of growth, because n/100 still grows in proportion to n. The number of maps is fixed while the number of items is not, so each map grows without limit — which is precisely what the next step fixes.

30. Worked example: what hashable means

Worked example

A limitation the implementation exposes.

hash('apple')      # fine
hash((1, 2))       # fine
hash([1, 2])       # TypeError: unhashable type: 'list'
ObjectHashable?Why
a stringhashableimmutable
a tuplehashableimmutable
a listunhashablemutable

Note the limitation.

Why: A limitation of this implementation is that it only works with hashable keys. Mutable types like lists and dictionaries are unhashable.

Note the guarantee.

Why: Hashable objects that are considered equivalent return the same hash value — which is what makes find_map find the item again.

Note the non-guarantee.

Why: But the converse is not necessarily true: two objects with different values can return the same hash value.

Figure (svg): The state of the program after each line of Worked example what hashable means, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The rule from chapter 11 explained. Dictionary keys have to be hashable because the implementation computes an index from them, and a mutable object's index would go stale.

Verify: Connect it to chapter 12's tuples as keys.

Why: Lesson 12b used tuples as dictionary keys, and this is why it worked and a list would not have: a tuple is immutable, so its hash is stable, and an item stored under it stays findable. A list's contents can change after it is stored, and then hash(k) points somewhere else.

31. Trap: reading BetterMap as a solution

Trap

The trap

BetterMap is a hundred times faster, so the problem is treated as solved.

Measure the improvement

Why: A hundredfold is a large improvement.

The order of growth for get is still linear. As the number of items grows, each of the hundred LinearMaps grows with it, so the lookup cost still grows in proportion to n.

The fix

Ask what happens as n grows.

The number of maps is fixed at 100

Why: So each one grows without limit.

Which means the search inside one grows too

Why: n/100 is still linear in n.

The fix is to let the number of maps grow with the data, which is exactly the crucial idea of the next section — and the reason BetterMap is described as a step on the path rather than a destination.

32. Discriminate: which method is which class?

Discrimination

Two classes, four methods.

Sort into buckets

For each method, what is its order of growth?

constant
LinearMap.add; BetterMap.find_map; hash(k)
linear
LinearMap.get; BetterMap.get; searching one of the 100 LinearMaps
con
Each does a fixed amount of work: appending to a list is constant on average, and computing a hash and a modulus takes the same time however many items are stored.
lin
Each searches a list whose length grows with the number of items — and a hundredth of a growing quantity still grows in proportion to it.

33. Complete it: find_map

Faded example

A hash, wrapped into range.

Fill in the blanks

def find_map(self, k):
index = hash(k) % len(self.maps)
return self.maps[index]

Why: The modulus operator wraps the hash values into the range from 0 to len(self.maps), so the result is always a legal index. Many different hash values will wrap onto the same index, which is fine — that map holds a few items and is searched linearly.

34. Explain it: why must dictionary keys be immutable?

Explain it

Chapter 11 stated the rule; this lesson explains it.

Discussion prompt

A classmate asks why a list cannot be a dictionary key. Answer using find_map.

Hint: Where does the item get stored?

Answer:

Because the position is computed from the key: index is hash(k) modulo the number of maps. The key decides where the item lives.

If the key can change after the item is stored, the computed position changes with it, and the item is no longer where the lookup will look. It would still be in the structure and permanently unfindable.

So mutable types like lists and dictionaries are unhashable — Python refuses to compute a hash for them rather than let you build a broken dictionary. Tuples are fine, which is why lesson 12b could use them as keys.

35. The crucial idea: bound the length

Section

Section 4

36. Grow the number of maps with the data

Concept

Here, finally, is the crucial idea that makes hashtables fast: if you can keep the maximum length of the LinearMaps bounded, LinearMap.get is constant time.

class HashMap:

    def __init__(self):
        self.maps = BetterMap(2)
        self.num = 0

    def get(self, k):
        return self.maps.get(k)

    def add(self, k, v):
        if self.num == len(self.maps.maps):
            self.resize()
        self.maps.add(k, v)
        self.num += 1
MethodWhat it doesNote
__init__a BetterMap of 2, and num = 0num counts the items
getdispatches to BetterMapno work of its own
addresizes when num equals the map countthen adds

All you have to do is keep track of the number of items, and when the number of items per LinearMap exceeds a threshold, resize the hashtable by adding more LinearMaps. Here the threshold is an average of one item per map.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 207-208

37. Picture it: the maps grow with the items

Picture it

Which keeps each one short.

Figure (svg): A chart showing the number of maps growing alongside the number of items so the ratio stays constant

The flat line is the one that matters: about one item per map, however many items there are.

And a search of a list with about one item in it takes about the same time however large the whole structure has become.

38. Worked example: resize, and why rehashing is needed

Worked example

You cannot just copy the maps across.

def resize(self):
    new_maps = BetterMap(self.num * 2)
    for m in self.maps.maps:
        for k, v in m.items:
            new_maps.add(k, v)
    self.maps = new_maps
StepWhat happensNote
a new BetterMaptwice as biggeometric growth
every itemadded againrehashed
the old mapsdiscardedreplaced wholesale

Note what resize builds.

Why: A new BetterMap, twice as big as the previous one, and then it rehashes the items from the old map to the new.

Note why rehashing is necessary.

Why: Changing the number of LinearMaps changes the denominator of the modulus operator in find_map — so an item's index is different in the new structure.

Note what that achieves.

Why: Some objects that used to hash into the same LinearMap will get split up, which is exactly what we wanted.

Figure (svg): A flowchart of what add does, showing the resize branch

The expensive branch is the one taken rarely, which is what makes the average constant.

Every item has to be re-placed because every item's address depends on the size of the table. That is the price of computing the location instead of storing it.

Verify: Check what happens if you skip the rehash.

Why: Items would sit at indices computed with the old denominator while lookups used the new one, so most lookups would search the wrong map and raise KeyError. The structure would be intact and wrong — a semantic error of exactly the kind lesson Ab described.

39. Predict: when does add call resize?

Prediction

Read the condition.

def add(self, k, v):
    if self.num == len(self.maps.maps):
        self.resize()
    self.maps.add(k, v)
    self.num += 1
QuantityWhat it countsNote
numthe number of items
len(self.maps.maps)the number of LinearMaps
equalone item per map on averageresize

Predict first

What condition triggers a resize?

  • When the number of items equals the number of LinearMaps — an average of one item each
  • When any single LinearMap has more than one item
  • On every add
  • When a hash collision occurs

Correct: When they are equal, the average number of items per LinearMap is 1, so it calls resize.

Why: The threshold is about the average rather than any individual map, which is what makes it cheap to check — num is a single counter. Individual maps may hold two or three items, and the average is what the bound is stated over.

40. Worked example: why the maximum length is the whole thing

Worked example

A bounded list is a constant-time search.

# LinearMap.get is linear in ITS OWN length
#
# BetterMap: 100 maps, n items
#   -> each map holds n/100  -> grows with n
#
# HashMap: maps grow with n
#   -> each map holds about 1 -> bounded
StructureLength of one mapget
BetterMapmap length grows with nlinear get
HashMapmap length stays about 1constant get
the differencewhether the map count is fixed

State what get actually costs.

Why: LinearMap.get is linear in the length of its own list, and nothing else — not in the total number of items.

Note what BetterMap failed to do.

Why: It divided that length by a hundred and left it growing, because the number of maps was fixed.

Note what HashMap does.

Why: It keeps the maximum length of the LinearMaps bounded, and then LinearMap.get is constant time.

Figure (svg): The state of the program after each line of Worked example why the maximum length is the whole thing, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The class changed because a growing quantity was made bounded. That is a different kind of move from making something a hundred times smaller, and it is why one gives an unbounded improvement and the other does not.

Verify: Say what the hash function has to do for this to hold.

Why: Spread things out pretty evenly, which is what hash functions are designed to do. If every key hashed to the same index, one map would hold everything and get would be linear again — the bound is an expectation about typical behaviour rather than a guarantee.

41. Trap: thinking the hash is what makes it fast

Trap

The trap

The speed is attributed to the hash function, since that is the distinctive part.

Compute an index instead of searching

Why: One calculation replaces a traversal.

BetterMap does exactly that and is still linear. The hash chooses a map; what makes the lookup constant is that the chosen map is short, and it is short only because the number of maps grows.

The fix

Attribute it to the bound.

The hash distributes items evenly

Why: Necessary, and not sufficient.

The resizing keeps each map short

Why: Which is the crucial idea.

Both parts are needed: an even spread over a fixed number of buckets gives a smaller coefficient, and a growing number of buckets with a bad hash gives no improvement at all.

42. Complete it: the crucial idea

Faded example

One word carries it.

Fill in the blanks

# if you can keep the maximum length of the
# LinearMaps bounded, LinearMap.get is constant time

Why: That is the whole difference between BetterMap and HashMap. Dividing a growing length by a hundred leaves it growing; holding it below a fixed ceiling makes the search inside it constant, which changes the order of growth rather than the coefficient.

43. Two truths and a lie: resizing

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. Rehashing is necessary because changing the number of maps changes the modulus denominator
  • B. resize makes a new BetterMap twice as big as the previous one
  • C. resize is constant time because it only allocates a bigger list

Survives elimination: C

Why: C ignores the rehash. Rehashing is linear, so resize is linear — every item has to be added to the new structure. What makes add constant on average is not that resize is cheap but that it happens rarely enough for its cost to average out.

44. Think it through: why start with only two maps?

Socratic

HashMap begins with BetterMap(2).

Discussion prompt

The implementation starts with two LinearMaps rather than a hundred. Why is that reasonable?

Hint: What decides how many there should be?

Answer:

Because the right number depends on how many items there are, and at the start there are none. Two costs nothing to allocate and is enough for the first two adds.

The structure then grows to fit: each resize doubles it, so after a handful of resizes it is as large as it needs to be. Guessing a hundred in advance would waste space on small maps and still be wrong for large ones.

Which is the difference from BetterMap in one sentence. BetterMap picks a size once and lives with it; HashMap lets the data decide, and that is what keeps the items-per-map ratio bounded.

45. Why add is constant on average

Section

Section 5

46. Rare expensive operations average away

Concept

Rehashing is linear, so resize is linear, which might seem bad — since add was promised to be constant time. But we don't have to resize every time, so add is usually constant time and only occasionally linear.

# starting empty, with 2 LinearMaps:
#   adds 1-2      1 unit each          total  2
#   add 3         resize (2) + add (1)
#   add 4         1 unit               total  6
#   add 5         5 units
#   adds 6-8      1 unit each          total 14
#   add 9         9 units              total 30 by 16
#   by 32 adds                         total 62
AddsTotal workNote
4 adds6 units
8 adds14 units
16 adds30 units
32 adds62 units2n minus 2

After n adds, where n is a power of two, the total cost is 2n minus 2 units, so the average work per add is a little less than 2 units. For other values of n the average is a little higher, and the important thing is that it is O(1).

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 208-208

47. Picture it: towers, then knock them over

Picture it

The book's own figure, described.

Figure (svg): The cost of successive adds, showing tall rare spikes among many cheap ones

Increasingly tall towers with increasing space between them — and the spacing grows as fast as the height.

Now if you knock over the towers, spreading the cost of resizing over all adds, you can see graphically that the total cost after n adds is 2n minus 2.

48. Worked example: the accounting

Worked example

Follow the units.

# adds 1, 2      cheap                    2 units
# add 3          rehash 2 + add 1 = 3     5
# add 4          1                        6  (4 items)
# add 5          rehash 4 + add 1 = 5    11
# adds 6, 7, 8   1 each                  14  (8 items)
# add 9          rehash 8 + add 1 = 9    23
# adds 10-16     1 each                  30  (16 items)
AfterTotalAverage
4 items6 units1.5 per add
8 items14 units1.75 per add
16 items30 units1.875 per add
n items2n minus 2under 2 per add

Note where the spikes are.

Why: A resize happens at add 3, 5, 9, 17 — each time the table doubles, so the gaps between resizes double as well.

Note how big each spike is.

Why: A resize rehashes everything currently stored, so its cost is the number of items — which is also the number of cheap adds since the last resize.

Read the pattern.

Why: After n adds, where n is a power of two, the total cost is 2n minus 2 units, so the average work per add is a little less than 2 units.

Figure (svg): A chart contrasting the total work with the average work per add as the number of adds grows

A linear total over n operations is a constant average — the same argument that made list append constant.

Under two units per add, forever. The average does not creep upward as the structure grows, which is what makes it O(1) rather than merely small.

Verify: Check the trend in the average column.

Why: 1.5, then 1.75, then 1.875 — rising, and toward 2 rather than without limit. A cost that approaches a fixed ceiling is constant time; one that keeps climbing would not be.

49. Predict: the total after 32 adds

Prediction

The pattern is in the table.

# 4 adds  ->  6 units
# 8 adds  -> 14 units
# 16 adds -> 30 units
# 32 adds -> ?
nTotalThe rule
462n minus 2
8142n minus 2
16302n minus 2

Predict first

What is the total cost after 32 adds?

  • 62 units — twice n, less two
  • 60 units
  • 1024 units
  • 32 units

Correct: 62 units — after 32 adds, the total cost is 62 units, and after n adds where n is a power of two the total cost is 2n minus 2.

Why: Doubling the number of adds roughly doubles the total, which is what a linear total looks like. Dividing gives an average of a little under two units per add, and it stays under two however large n gets — that is the O(1) claim.

50. Worked example: why the growth must be geometric

Worked example

Doubling rather than adding.

# geometric: BetterMap(self.num * 2)
#   resizes at 2, 4, 8, 16, 32 ...
#   -> average add is constant

# arithmetic: BetterMap(self.num + 10)
#   resizes at 10, 20, 30, 40 ...
#   -> average add is LINEAR
GrowthWhat happensNote
geometricgaps double as costs doublethey cancel
arithmeticgaps stay fixed as costs growthey do not
the consequencethe growth rule is load-bearing

Name the property.

Why: An important feature of this algorithm is that when we resize the HashTable it grows geometrically; that is, we multiply the size by a constant.

Note the alternative.

Why: If you increase the size arithmetically — adding a fixed number each time — the average time per add is linear.

See why.

Why: With arithmetic growth the resizes keep costing more while staying equally frequent, so their total cost outgrows the number of adds.

Figure (svg): The state of the program after each line of Worked example why the growth must be geometric, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A one-word change in resize that changes the order of growth of the whole structure. Doubling is not an arbitrary choice.

Verify: Connect it back to lists.

Why: The previous lesson said adding to the end of a list is constant on average because when it runs out of room it occasionally gets copied to a bigger location. That is this same argument, and Python's list uses geometric growth for exactly this reason.

51. Trap: reporting the worst case as the answer

Trap

The trap

add is called linear because one call in it can rehash everything.

Analyse the worst case

Why: Which the previous lesson said to do.

For a single call that is right, and for a sequence of calls it is misleading. The total amount of work to run add n times is proportional to n, so the average time of each add is constant.

The fix

Say which claim you are making.

A single add is occasionally linear

Why: True, and the worst case.

n adds cost O(n) in total

Why: So the average is constant.

Both statements are correct about different things. The second is the one that matters for a program that does many adds, which is nearly every program — and it is why the previous lesson's table says 'constant on average' rather than 'constant'.

52. Sort: constant or not?

Sorting

In the structures the lesson built.

Sort into buckets

For each operation, is it constant time?

constant
HashMap.get; HashMap.add, on average; LinearMap.add
linear
HashMap.resize; LinearMap.get; BetterMap.get
yes
Each does a bounded amount of work: the maps stay short, appending to a list is constant, and add's occasional rehash averages out to a little under two units.
no
Each traverses something that grows with the number of items — rehashing touches every item, and both of the earlier get methods search a list whose length grows with n.

53. Complete it: the growth rule

Faded example

One word decides the class.

Fill in the blanks

# resizing must grow the table geometrically
# - multiplying the size by a constant.
# Adding a fixed number each time makes
# the average add linear.

Why: Doubling makes the gaps between resizes grow at the same rate as the cost of a resize, so the two cancel. Adding a fixed amount each time leaves resizes equally frequent while they get steadily more expensive, and the average climbs without limit.

54. Where a rare expensive step averages away

Real world

Something done occasionally to keep everything else cheap.

Discussion prompt

Think of something rarely done that is a lot of work at once, and that keeps the everyday version quick. What makes it worth it?

Hint: How often, against how much?

Answer:

Reorganising a filing system, restocking a shelf, moving to a bigger space — each costs a great deal on the day and nothing on most days.

What makes it worth it is the ratio: if the big job gets twice as large but happens half as often, the average effort per day does not change. If it gets larger without getting rarer, it eventually dominates everything.

Which is exactly the geometric-growth argument. Doubling makes the resizes twice as expensive and twice as far apart, so the average add stays under two units — and arithmetic growth, which does not, makes it linear.

55. Compare: the three maps

Comparison

Fill the blanks. One interface, three implementations.

Comparison matrix

QuestionLinearMapBetterMapHashMap
What holds the items?one list of tuples100 LinearMapsa BetterMap that resizes
How does get find the item?a for loop over everythinghash picks a map, then a for loopthe same, over a short list
Order of growth of getlinearlinear, smaller coefficientconstant
Does the map count grow with the data?there is only oneno — fixed at 100yes — it doubles

The bottom row is the whole difference. Everything else about BetterMap and HashMap is the same code.

56. The procedure: reading a data structure's cost

Pattern

Five questions, and the fourth is the one people skip.

  1. Ask what the operation has to look at — one item, or all of them.
  2. If it searches, ask whether the collection is ordered: unordered means linear, ordered permits bisection.
  3. If the position is computed rather than searched for, ask how long the list at that position is.
  4. Ask whether that length is bounded, or whether it grows with the total — this is what separates a smaller coefficient from a better class.
  5. For an operation that is occasionally expensive, ask whether the total over n operations is proportional to n.

Step 4 is BetterMap against HashMap, and step 5 is why both list append and dictionary add are described as constant on average rather than constant.

Python documentation — Built-in Types Built-in Types

57. Check yourself 1 of 3: search algorithms

Check

One requires a precondition.

Check your understanding

Which search is O(log n), and what does it require?

  • A. Bisection search — it requires the sequence to be in order (correct)
  • B. Linear search — it requires nothing
  • C. A hashtable lookup — it requires hashable keys
  • D. All three are O(log n)

Answer: A

Why: Bisection search cuts the number of remaining items in half on each step, which is what makes it logarithmic — about 20 steps for a million items. It only works because one comparison at the middle tells you which half the target is in, and that is true only if the sequence is in order.

Why B tempts people
A linear search is O(n): in the worst case it has to traverse the entire collection.
Why C tempts people
A hashtable does a search in constant time, which is better than logarithmic, and it doesn't require the items to be sorted.
Why D tempts people
The three are the lesson's three different classes, which is the point of comparing them.

58. Check yourself 2 of 3: BetterMap

Check

A hundredfold improvement.

Check your understanding

BetterMap is about 100 times faster than LinearMap. What is its order of growth?

  • A. Still linear — with a smaller leading coefficient (correct)
  • B. Constant, because of the hash function
  • C. Logarithmic
  • D. It depends on the hash function

Answer: A

Why: The number of maps is fixed at a hundred while the number of items is not, so each LinearMap holds about n/100 items and grows without limit. Dividing by a constant does not change the order of growth — which is exactly the previous lesson's distinction between a constant factor and a better class.

Why B tempts people
The hash chooses a map in constant time, and then a linear search runs inside a list that grows with n.
Why C tempts people
Nothing here halves the remaining items; that is bisection search.
Why D tempts people
A good hash function gives the hundredfold; it cannot make a growing list stop growing.

59. Check yourself 3 of 3: the average cost of add

Check

Occasionally linear.

Check your understanding

Why is HashMap.add constant time on average?

  • A. The total work for n adds is proportional to n, so the average per add is constant (correct)
  • B. Because resize is constant time
  • C. Because resizing never actually happens
  • D. Because the hash function is fast

Answer: A

Why: After n adds the total cost is about 2n units, so the average is a little under two per add and stays there however large n becomes. The resizes get more expensive as the structure grows and correspondingly rarer, and geometric growth is what makes those two effects cancel.

Why B tempts people
Rehashing is linear, so resize is linear — that is why the argument is needed at all.
Why C tempts people
It happens at adds 3, 5, 9, 17 and so on; it just gets rarer.
Why D tempts people
A fast hash is a constant factor and would not answer the objection about resizing.

60. Where this shows up outside this course

Real world

Finding something without looking through everything.

Discussion prompt

Think of a place where things are stored so that you can go straight to the one you want. What decides where each thing goes, and what happens when the space runs out?

Hint: Shelves, pigeonholes, an index.

Answer:

Something about the item itself decides its place — the first letter of a name, a number, a category — so finding it means computing the location rather than searching for it.

When the space runs out, everything has to be redistributed over a larger set of places, and the rule that assigned them has to change with it. That redistribution is expensive and it is rare.

Both halves are the hashtable: find_map computes an index from the key, and resize doubles the space and rehashes everything because the assignment rule now has a different denominator.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

BetterMap splits the items across a hundred LinearMaps and is about a hundred times faster. Is its get constant time?

  • No — the order of growth is still linear, with a smaller leading coefficient
  • Yes — the hash goes straight to the right map
  • Yes, as long as the hash spreads items evenly
  • It is logarithmic

Correct: No — the order of growth is still linear, but the leading coefficient is smaller.

Why: The hash does go straight to the right map, in constant time, and then a linear search runs inside that map. Each map holds about n/100 items, and a hundredth of a growing quantity still grows in proportion to n — so the lookup cost still rises as items are added. This is the previous lesson's distinction in its sharpest form: a hundredfold speed-up is a constant factor, fixed at every size, and it does not change the class. What changes the class is the crucial idea of the next section: if you can keep the maximum length of the LinearMaps bounded, LinearMap.get is constant time. HashMap does that by counting the items and doubling the number of maps whenever the average reaches one per map, so the length being searched stops growing at all. A better coefficient and a better class look similar at one size and behave completely differently as the data grows.

62. Explain it to someone else

Explain it

The question chapter 11 left open.

Discussion prompt

A classmate asks how a dictionary can look something up in the same time whether it holds ten items or ten million. Explain it in three sentences.

Hint: Compute, bound, resize.

Answer:

The position is computed from the key rather than searched for: hash(k) modulo the number of buckets gives an index, and that calculation costs the same however many items there are.

Each bucket holds a short list, which is searched linearly — and it stays short because the number of buckets grows with the number of items, roughly one item per bucket.

Growing means doubling and rehashing everything, which is linear but rare enough that the average add stays under two units of work. That is the minor miracle: computing the location, bounding the bucket, and paying for the growth in rare instalments.

63. Exit ticket

Exit ticket

One honest answer. The course ends here, and this is for you.

Predict first

Which of these is still least solid for you?

  • Linear and bisection search, and what each requires
  • LinearMap and BetterMap, and why BetterMap is still linear
  • The crucial idea — bounding the length of each map
  • The averaging argument that makes add constant

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The two searches are worth knowing as a pair, because the contrast between them is what a precondition buys. BetterMap is the most instructive failure in the book: a hundredfold improvement that changes nothing about the class. The crucial idea is one sentence and it is the whole of chapter B.4. And the averaging argument is the one that also explains list append, so it is worth being able to reconstruct rather than recall.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the three search strategies with their orders of growth and what each requires. Underneath, draw the three maps in order — one list of tuples, a hundred lists chosen by a hash, and a growing number of lists — and write the order of growth of get beside each. Mark the one place where the class changes, and write the sentence that makes it change. Finish with the accounting: 4 adds, 8 adds, 16, 32, and the formula.

65. What you can do now

Recap

Four pages, and the answer to a question the course has carried since chapter 11.

If you remember one thingIt is this
From linear searchAbsence can only be established by looking at everything.
From bisectionA precondition can buy you a better order of growth.
From BetterMapA hundredfold improvement is still a constant factor.
From HashMapBounding a growing quantity is what changes the class.
From the accountingRare and expensive averages away if the rarity grows with the cost.

And Downey's own last word on the code: you wouldn't write this in Python — if you want a map, just use a dictionary. The value of building one is knowing what the dictionary is doing, which is what makes it possible to predict when it will be fast and when a list will not.

Think Python, 2nd edition — Allen B. Downey §B.3-B.4, pp. 205-208 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §B.3-B.4, pp. 205-208
  2. Python documentation — Built-in Types
  3. Python documentation — Data Structures
  4. Open Data Structures (Python edition) — Pat Morin — Pat Morin, Open Data Structures (Python edition), opendatastructures.org

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108