Ba Order of Growth and the Cost of Python Operations

This lesson introduces order of growth as a way of comparing algorithms, works through why the leading term dominates, and catalogues the run time of the Python operations the course has been using.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson Ba Order of Growth and the Cost of Python Operations

Title

Python · Appendix B — Analysis of Algorithms

§B.1-B.2, pp. 201-204

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-204 — the pages these objectives are drawn from

3. Before we start: why was the dictionary faster?

Warm-up

Chapter 11 said so without explaining it.

Discussion prompt

Lesson 11a said the in operator takes about the same time for a dictionary however many items it has, and gets slower for a list as the list grows. What kind of claim is that, and what would it take to justify it?

Hint: It is not a claim about one program on one machine.

Answer:

It is a claim about how the cost behaves as the size grows, rather than about any particular timing — a shape rather than a number.

Justifying it needs a way of talking about that shape which does not depend on the machine, the data or the size. That is what this appendix provides.

And it names something the course has demonstrated repeatedly: the list search in chapter 13, the memo in chapter 11, and the list-and-join in chapter 18 were all this idea without the vocabulary.

4. The one idea behind this appendix: compare how the cost grows

Concept

The practical goal of algorithm analysis is to predict the performance of different algorithms in order to guide design decisions.

That third response is the one this appendix builds on: group functions into categories depending on how quickly they grow as problem size increases.

Figure (svg): Two columns pairing each obstacle to comparing algorithms with the standard response

Three ways a comparison could be meaningless, and three ways to keep it meaningful.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-202

5. The leading term dominates

Section

Section 1

6. Two algorithms, and a table

Concept

Suppose Algorithm A takes 100n + 1 steps to solve a problem with size n, and Algorithm B takes n squared plus n plus 1.

leading term — In an expression describing run time, the term with the highest exponent.

#  size        A: 100n+1      B: n^2+n+1
#    10           1 001            111
#   100          10 001         10 101
#  1000         100 001      1 001 001
# 10000       1 000 001    100 010 001
SizeWhat happensNote
n = 10A takes almost 10 times longerB looks better
n = 100about the samethe crossover
larger nA is much betterand the gap widens

The fundamental reason is that for large values of n, any function that contains an n-squared term will grow faster than a function whose leading term is n — regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n, for any a and b.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 202-202

7. Picture it: the crossover

Picture it

One algorithm wins at small sizes and loses at large ones.

Figure (svg): A growth chart showing a linear and a quadratic run time crossing over

At n = 10 the quadratic one is ten times faster. At n = 10,000 it is a hundred times slower.

Which is why the analysis is about large problems: for small ones, order of growth is irrelevant, and the book says so.

8. Worked example: why the coefficient cannot save it

Worked example

Even a very large one.

# A: 100n + 1        - a big coefficient
# B: n^2 + n + 1     - a small one

# and even if A were n + 1000000,
# it would still be better than B
# for sufficiently large n
RangeWhat decidesNote
at small nthe coefficient dominatesB wins
at large nthe exponent dominatesA wins
alwayssome n exists where it flipsfor any coefficients

Note what the coefficient does.

Why: For Algorithm A, the leading term has a large coefficient, 100, which is why B does better than A for small n.

Note what it cannot do.

Why: Regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n, for any values of a and b.

Note that non-leading terms are the same.

Why: Even if the run time of Algorithm A were n + 1000000, it would still be better than Algorithm B for sufficiently large n.

Figure (svg): The state of the program after each line of Worked example why the coefficient cannot save it, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The exponent wins eventually, whatever the constants. That is what makes the leading term the thing worth comparing.

Verify: Check where the crossover moves.

Why: A bigger coefficient on A pushes the crossover to a larger n without preventing it — so the constants decide where the change happens rather than whether. That is why they are ignored for classification and still matter in practice.

9. Predict: which algorithm wins?

Prediction

The table's smallest size.

# A: 100n + 1        at n = 10 -> 1001
# B: n^2 + n + 1     at n = 10 ->  111
AlgorithmStepsNote
A1001 stepsthe big coefficient
B111 stepsthe small n
at n = 10B is fasterby almost ten times

Predict first

At n = 10, which algorithm is faster?

  • B, by almost ten times — despite its worse order of growth
  • A, because linear beats quadratic
  • They are the same
  • It cannot be determined

Correct: B, by almost ten times — at n = 10, Algorithm A looks pretty bad.

Why: Order of growth describes behaviour for large n, and ten is not large. The crossover is around n = 100, above which A wins by a widening margin — which is why the book says that for small problems, order of growth is irrelevant.

10. Worked example: the crossover point

Worked example

Ignored for analysis, and not for engineering.

# 'In general, we expect an algorithm with a smaller
#  leading term to be a better algorithm for large
#  problems, but for smaller problems, there may be a
#  crossover point where another algorithm is better.'

# 'it is usually ignored for purposes of algorithmic
#  analysis. But that doesn't mean you can forget
#  about it.'
AspectWhat is trueNote
where it isdepends on the algorithms, inputs and hardwarenot general
for analysisignoredit is not a property of the class
for a real programmay be exactly your range

Note that it exists.

Why: For smaller problems there may be a crossover point where another algorithm is better — which the table shows at around n = 100.

Note why analysis ignores it.

Why: The location depends on the details of the algorithms, the inputs and the hardware, so it is not a property of the growth class.

Note the caveat.

Why: But that doesn't mean you can forget about it — if your problem sizes sit below the crossover, the theoretically worse algorithm is the faster one.

Figure (svg): Two columns separating what order of growth captures from what it discards

The right-hand column is real and is not part of the class — which is why the caveats matter.

A real quantity that the classification deliberately discards. Knowing that it was discarded is what stops the analysis being misapplied.

Verify: Connect it to the course's own examples.

Why: Lesson 19b's has_duplicates is exactly this: the set version is shorter and the loop version stops at the first duplicate, so for a long list with an early repeat the loop wins. Order of growth would not distinguish them, and the details do.

11. Trap: choosing by order of growth alone

Trap

The trap

An algorithm is replaced by one with a better order of growth, on a problem with a few dozen items.

Prefer the better class

Why: Which is right for large problems, and this is the standard advice.

For small problems, order of growth is irrelevant — the book says so — and the crossover may be far above the sizes involved. The replacement can be slower and is certainly more work.

The fix

Check the sizes you actually have.

Below the crossover, the constants decide

Why: And they are what the classification discarded.

Above it, the order of growth decides

Why: And the difference can be unbounded.

This is chapter 13's advice again: choose the structure that is easiest to implement and see whether it is fast enough. Order of growth tells you what to do when it is not.

12. Discriminate: what is the leading term?

Discrimination

The term with the highest exponent.

Sort into buckets

For each run time, what is the leading term?

linear — n
100n + 1; n + 1000000; 2n
cubic — n cubed
1000000n cubed + n squared; n cubed + 1000000n squared
quadratic — n squared
n squared + n + 1
lin
Each has n as its highest power, whatever the coefficient or the constant — even a constant of a million does not change the leading term.
cub
Both have n cubed as their highest power, and the enormous coefficients on the other terms are irrelevant to the classification.
quad
The highest power is two, so it is quadratic — and the n and the 1 do not affect the class.

13. Complete it: the leading term

Faded example

The term with the highest exponent.

Fill in the blanks

# run time: 1000000n^3 + n^2
# leading term: n^3

Why: The leading term is the term with the highest exponent, so the coefficient of a million is irrelevant to the classification — this is cubic. Coefficients decide where the crossover is, and exponents decide which algorithm eventually wins.

14. Think it through: why ignore the coefficients?

Socratic

They are real, and the analysis discards them.

Discussion prompt

A coefficient of 100 makes a real difference to how long a program takes. Why does the classification throw it away?

Hint: What does the coefficient depend on?

Answer:

Because it depends on things the analysis is trying to abstract away: the hardware, the language, the implementation. The same algorithm has a different coefficient on a different machine.

The exponent does not. It is a property of the algorithm itself, so it is the part that transfers between machines and implementations — which is what makes a general comparison possible at all.

And the book's closing line puts the trade precisely: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.

15. The orders of growth

Section

Section 2

16. A set of functions considered equivalent

Concept

An order of growth is a set of functions whose growth behaviour is considered equivalent. For example, 2n, 100n and n + 1 belong to the same order of growth, written O(n) and often called linear.

# O(1)         constant
# O(log n)     logarithmic
# O(n)         linear
# O(n log n)   linearithmic
# O(n^2)       quadratic
# O(n^3)       cubic
# O(c^n)       exponential
AspectWhat is trueNote
the listin increasing order of badnessthe book's phrase
the base of a logarithmdoes not matterchanging bases multiplies by a constant
the base of an exponentlikewiseall exponentials are one class

For the logarithmic terms, the base doesn't matter — changing bases is the equivalent of multiplying by a constant, which doesn't change the order of growth. Exponential functions grow very quickly, so exponential algorithms are only useful for small problems.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 203-203

17. Picture it: seven classes, worst last

Picture it

In increasing order of badness.

Figure (svg): The seven common orders of growth listed from best to worst

Each step down the list is a qualitative change rather than a constant factor.

And the top two are worth noticing: constant time means the size does not matter at all, and logarithmic means it barely does.

18. Worked example: what each class means concretely

Worked example

A million items, and how much work each class implies.

# n = 1,000,000
#   O(1)        about 1 step
#   O(log n)    about 20 steps
#   O(n)        a million steps
#   O(n log n)  about 20 million
#   O(n^2)      a million million
ClassAt a million itemsNote
constantunaffected by size1
logarithmicbarely affected20
quadratica trillion stepsimpractical

Note the top of the list.

Why: Constant and logarithmic barely respond to size at all — twenty steps for a million items is the bisection search of the next lesson.

Note the middle.

Why: Linear and linearithmic scale usefully: a million steps is fast, and twenty million is still fine.

Note the bottom.

Why: Quadratic at a million items is a million million steps, which is why the difference between classes is not a constant factor.

Figure (svg): A growth chart comparing constant, logarithmic, linear and quadratic growth

Three shapes rather than three speeds, which is the point of the classification.

Five classes and five wildly different amounts of work for one problem size. That spread is what the classification is for.

Verify: Check the claim about unboundedness.

Why: Between O(n) and O(n squared) at a million items the ratio is a million — and at ten million it is ten million. So the difference between a good algorithm and a bad algorithm is unbounded, exactly as the book says, where two algorithms in the same class differ by a fixed factor forever.

19. Sort: which class is better?

Sorting

In increasing order of badness.

Sort into buckets

For each pair, which order of growth is better for large problems?

the first
O(1) or O(n); O(log n) or O(n); O(n) or O(n squared); O(n log n) or O(n squared); O(n squared) or O(2 to the n); O(n) or O(n log n)
first
In every pair the first is earlier in the book's table, which lists the orders in increasing order of badness — so it grows more slowly and wins for sufficiently large n.
second
None of these pairs favours the second, since each is listed as worse. For small n the constants could reverse any of them, which is what the crossover point is about.

20. Worked example: why the base does not matter

Worked example

For logarithms and for exponentials, differently.

# log base 2 of n, and log base 10 of n,
# differ by a constant factor
#   -> the same order of growth

# 2^n and 3^n are both exponential
#   -> the same order of growth
ChangeEffect on the classNote
changing log basesmultiplies by a constantno change of class
changing exponent basesall exponentialone class
what that meansthe classes are coarsedeliberately

Note the logarithm rule.

Why: Changing bases is the equivalent of multiplying by a constant, which doesn't change the order of growth — so O(log n) needs no base.

Note the exponential rule.

Why: All exponential functions belong to the same order of growth regardless of the base of the exponent.

Note the coarseness.

Why: The classification is deliberately loose, because its purpose is to distinguish shapes rather than to rank precisely.

Figure (svg): The state of the program after each line of Worked example why the base does not matter, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Two rules that keep the classes small in number. The coarseness is the feature: seven classes cover most of what you need to distinguish.

Verify: Ask whether the coarseness ever misleads.

Why: It can: two O(n log n) sorts can differ by a large constant, and the book says that if two algorithms have the same leading order term, it is hard to say which is better — the answer depends on the details. So the classification is a first filter rather than a final answer.

21. Trap: treating the class as a speed

Trap

The trap

An O(n log n) algorithm is assumed to be faster than an O(n) one on some particular input.

Read the classes as a ranking

Why: They are listed in increasing order of badness.

The classes describe growth rather than speed, and for a specific size the constants can reverse the order. Two algorithms in the same class can also differ by a large constant factor.

The fix

Read the class as a shape.

It says how the cost responds to size

Why: Not how long anything takes.

And a comparison at one size needs measurement

Why: Which is chapter 13's benchmarking.

The book's own hedge applies: sometimes the coefficients and the non-leading terms make a real difference, and for small problems order of growth is irrelevant. It is a tool with caveats rather than a verdict.

22. Predict: are these the same order of growth?

Prediction

Different coefficients, same exponent.

# 2n, 100n, and n + 1
FunctionLeading termClass
all threeleading term nlinear
the coefficientsdifferand do not matter
the classO(n)

Predict first

Do 2n, 100n and n + 1 belong to the same order of growth?

  • Yes — all are O(n), because every function in the set grows linearly with n
  • No — the coefficients differ
  • No — n + 1 has an extra term
  • Only the first two

Correct: Yes — all are O(n), and every function in the set grows linearly with n.

Why: An order of growth is a set of functions whose growth behaviour is considered equivalent, and functions with the same leading term are considered equivalent even if they have different coefficients. The differences are real and are constant factors, which the classification deliberately discards.

23. Complete it: name the class

Faded example

The leading term names it.

Fill in the blanks

# run time: n^2 + n + 1
# order of growth: O(n^2), called quadratic

Why: All functions with the leading term n squared belong to O(n squared), and they are called quadratic. The n and the 1 are non-leading terms, which do not affect the class — only the term with the highest exponent does.

24. Two truths and a lie: orders of growth

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. For logarithmic terms, the base of the logarithm doesn't matter
  • B. The difference between a good algorithm and a bad algorithm is unbounded
  • C. An algorithm with a better order of growth is always faster

Survives elimination: C

Why: C ignores the crossover point. For smaller problems there may be a point where another algorithm is better — the book's own table shows the quadratic algorithm winning by almost tenfold at n = 10. Order of growth is about behaviour for large n, and for small problems it is irrelevant.

25. The cost of Python's operations

Section

Section 3

26. What you have been using, and what it costs

Concept

Most arithmetic operations are constant time, and indexing operations — reading or writing elements in a sequence or dictionary — are also constant time, regardless of the size of the data structure.

Very large integers are an exception to the arithmetic rule; in that case the run time increases with the number of digits.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 204-204

27. Picture it: the operations you have used, by class

Picture it

Three groups, and the exceptions are where the interest is.

Figure (svg): Two columns separating constant-time operations from linear ones

Indexing is constant regardless of the size of the data structure, which is why it appears so often in fast code.

And two exceptions are worth their own slides: appending to a list, which is constant on average, and dictionary operations, which are the minor miracle of the next lesson.

28. Worked example: the rule of thumb for loops

Worked example

A loop multiplies the class of its body.

total = 0
for x in t:
    total += x
# the body is constant, so the loop is linear

# rule of thumb: if the body is in O(n^a),
# the whole loop is in O(n^(a+1))
PartIts classNote
a constant bodyO(1)one addition
the loopn passesO(n)
a linear bodywould give O(n squared)a nested loop

Note the base case.

Why: A for loop that traverses a sequence or dictionary is usually linear, as long as all of the operations in the body are constant time.

Note the rule.

Why: If the body of a loop is in O(n^a) then the whole loop is in O(n^(a+1)).

Note the exceptions.

Why: Unless the loop exits after a constant number of iterations — if a loop runs k times regardless of n, the loop is in O(n^a), even for large k.

Figure (svg): A flowchart showing how a loop's class is derived from its body's

Each level of nesting adds one to the exponent, unless the loop count is bounded.

A rule that composes: each level of nesting adds one to the exponent, which is why a nested loop over the same sequence is quadratic.

Verify: Check the division case.

Why: Multiplying by k doesn't change the order of growth, and neither does dividing — so a loop whose body is O(n^a) running n/k times is still O(n^(a+1)), even for large k. Halving the number of passes is a constant factor, not a change of class.

29. Sort: constant or linear?

Sorting

Ask what the operation touches.

Sort into buckets

For each operation, which class is it?

constant
t[i]; len(t); t.append(x)
linear
sum(t); max(t); '-'.join(t)
con
Each touches one element or a stored size — indexing is constant regardless of the size of the data structure, and appending is constant on average.
lin
Each has to look at every element: summing, finding a maximum, and joining all traverse the whole sequence, and join's run time depends on the total length of the strings.

30. Worked example: why append is constant on average

Worked example

An exception worth understanding.

# 'Adding an element to the end of a list is constant
#  time on average; when it runs out of room it
#  occasionally gets copied to a bigger location, but
#  the total time for n operations is O(n), so the
#  average time for each operation is O(1).'
CaseWhat it costsNote
most appendsconstantthere is room
occasionallya copy to a bigger locationlinear, that once
over n operationsO(n) in totalso O(1) each on average

Note the usual case.

Why: Adding to the end costs a constant amount when there is room, which is nearly always.

Note the occasional case.

Why: When it runs out of room it occasionally gets copied to a bigger location, which is proportional to the current size.

Note how they average.

Why: The total time for n operations is O(n), so the average time for each operation is O(1) — the expensive copies are rare enough to disappear into the average.

Figure (svg): The state of the program after each line of Worked example why append is constant on average, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Constant time on average, which is a different claim from constant time always. That distinction is why the book states it carefully rather than just listing append as O(1).

Verify: Connect it to lesson 18b's string building.

Why: That lesson's list-and-join relies on this: appending is cheap on average, where repeated string concatenation copies everything every time and is genuinely linear per operation. The two techniques differ precisely because one has this amortisation and the other does not.

31. Trap: assuming a slice costs what the sequence costs

Trap

The trap

A slice of two elements from a million-element list is assumed to be expensive.

Scale the cost with the container

Why: Most operations on a large structure cost more.

The run time of a slice operation is proportional to the length of the output, but independent of the size of the input — so a two-element slice is cheap however large the list.

The fix

Ask what the operation actually touches.

A slice touches the output

Why: So its cost is the output's length.

An index touches one element

Why: So it is constant regardless of the size of the data structure.

The rule that generalises is to count what the operation reads or writes rather than what it is applied to. Indexing, slicing and len all illustrate it differently.

32. Predict: what class is this loop?

Prediction

A constant-time body.

total = 0
for x in t:
    total += x
PartIts classNote
the bodyone additionconstant
the loopn passes
the whole thingO(n)linear

Predict first

What is the order of growth of this loop?

  • O(n) — linear
  • O(1) — constant
  • O(n squared) — quadratic
  • O(n log n)

Correct: O(n) — a for loop that traverses a sequence is usually linear, as long as the operations in the body are constant time.

Why: The rule of thumb is that if the body is in O(n^a), the whole loop is in O(n^(a+1)) — a constant body gives a linear loop. The built-in sum is also linear because it does the same thing, though it tends to be faster: in the language of algorithmic analysis, it has a smaller leading coefficient.

33. Complete it: the loop rule of thumb

Faded example

Nesting adds to the exponent.

Fill in the blanks

# if the body of a loop is in O(n^a),
# the whole loop is in O(n^(a+1))

Why: Each level of looping multiplies the work by n, which adds one to the exponent — so a constant body gives a linear loop and a linear body gives a quadratic one. The exception is a loop that runs a constant number of times regardless of n, which does not change the class at all.

34. Explain it: why is sum faster than my loop?

Explain it

Both are linear.

Discussion prompt

A classmate replaced their summing loop with sum and it got noticeably faster. Explain, given that both are linear.

Hint: What does the class not capture?

Answer:

They are in the same order of growth: the built-in sum is also linear because it does the same thing. So the class does not distinguish them, and the class is not the whole story.

What differs is the constant. In the language of algorithmic analysis, sum has a smaller leading coefficient — it is a more efficient implementation of the same algorithm.

Which is exactly what the classification discards. Two algorithms with the same order of growth differ by a constant factor, and constant factors are real: this one is measurable, and it is not what order of growth is for.

35. The exceptions worth knowing

Section

Section 4

36. Where the simple rules do not hold

Concept

The general rules cover most operations, and the exceptions are the ones that decide real programs.

That last one is lesson 12b's iterator point in cost terms: producing the iterator is cheap because nothing is computed, and the work happens when it is walked.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 204-205

37. Picture it: four exceptions and what each tells you

Picture it

Each contradicts a reasonable guess.

Figure (svg): Four operations whose cost differs from what the general rule would suggest

Each is worth knowing because the obvious guess is wrong.

And the fourth explains something from chapter 12: an iterator is cheap to make because it has computed nothing yet.

38. Worked example: why iterators are constant time

Worked example

The cost has been deferred rather than avoided.

d.keys()                 # constant: nothing computed

for k in d.keys():       # linear: now it walks
    ...
OperationIts classNote
making the iteratorconstantit knows how, and has not
walking itlinearone step per item
the totallinearthe work was deferred

Note why creation is cheap.

Why: keys, values and items are constant time because they return iterators — no values have been produced.

Note where the work happens.

Why: But if you loop through the iterators, the loop will be linear, because that is when each item is produced.

Connect it to chapter 12.

Why: This is why a zip object printed as an object description rather than as pairs: it had computed nothing, which is also why it was cheap.

Figure (svg): The state of the program after each line of Worked example why iterators are constant time, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Deferred rather than avoided. The total work is the same; what changes is when it happens, and whether it happens at all if the iterator is abandoned.

Verify: Ask when the deferral is a genuine saving.

Why: When you stop early — any over a generator that finds a True value immediately does one step rather than n, which is lesson 19a's point about the brackets. If the whole thing is walked, the deferral saves the memory rather than the time.

39. Predict: what does a small slice of a huge list cost?

Prediction

The list has a million elements.

t = list(range(1000000))
s = t[0:2]
AspectWhat mattersNote
the inputa million elementsirrelevant
the outputtwo elementsthe cost
the classproportional to the output

Predict first

How expensive is this slice?

  • Cheap — proportional to the length of the output, not the input
  • Expensive — proportional to the size of the list
  • Constant, like indexing
  • It depends on the values

Correct: Cheap — the run time of a slice operation is proportional to the length of the output, but independent of the size of the input.

Why: The operation only has to copy the elements it produces, so a two-element slice costs the same from a million-element list as from a ten-element one. Counting what the operation reads or writes is the habit that gets these right.

40. Worked example: sorting, and where it sits

Worked example

Between linear and quadratic.

t.sort()                 # O(n log n)

# at n = 1,000,000:
#   linear:      1,000,000 steps
#   n log n:    ~20,000,000 steps
#   quadratic:  1,000,000,000,000 steps
ComparisonThe ratioNote
against linearabout twenty times moreat a million
against quadraticfifty thousand times less
in practicesorting a million items is fast

Place it on the list.

Why: O(n log n) is linearithmic, between linear and quadratic in the book's table.

Note what that means at scale.

Why: Twenty times a linear pass at a million items, which is a small multiple rather than a different order of difficulty.

Note the contrast.

Why: A quadratic sort at the same size is a million times worse — which is why bubble sort is conceptually simple but slow for large datasets.

Figure (svg): A growth chart placing linearithmic between linear and quadratic

Much nearer the bottom line than the top, which is why sorting large collections is practical.

A class much closer to linear than to quadratic in practice. That is why sorting a large collection is a reasonable thing to do and a quadratic algorithm on one is not.

Verify: Connect it to the Obama anecdote.

Why: The appendix opens with a candidate being asked for the most efficient way to sort a million 32-bit integers and replying that bubble sort would be the wrong way to go — which is true, because bubble sort is quadratic and a million squared is a million million.

41. Trap: assuming update costs what the dictionary costs

Trap

The trap

A small dictionary is updated into a very large one, and the operation is assumed to be expensive.

Scale the cost with the receiver

Why: It is the big one being changed.

The run time of update is proportional to the size of the dictionary passed as a parameter, not the dictionary being updated — so merging a small one into a large one is cheap.

The fix

Count what the operation reads.

update reads the argument

Why: One insertion per item in it.

And each insertion is constant

Why: Because dictionary operations are.

The general habit is the useful one: ask what the operation has to look at. A slice looks at its output, an index looks at one element, and update looks at its argument.

42. Discriminate: which cost applies?

Discrimination

Four exceptions to the general rules.

Sort into buckets

For each operation, what is its cost proportional to?

the output or the argument
a slice t[i:j]; d1.update(d2)
the size of the collection
looping over d.keys(); t.sort()
nothing — constant
d.keys(); t[i]
out
A slice costs the length of its output, and update costs the size of the dictionary passed as a parameter — in both cases the thing being read rather than the thing being operated on.
n
Both traverse the whole collection: walking an iterator produces every item, and sorting is O(n log n), which grows with n.
c
Indexing is constant regardless of the size of the data structure, and returning an iterator computes nothing at all.

43. Complete it: the cost of a slice

Faded example

Count what it produces.

Fill in the blanks

# a slice's run time is proportional to the
# length of the output, not of the input

Why: The operation copies only the elements it produces, so a short slice of a very long sequence is cheap. That is an instance of the general habit worth forming: count what an operation actually reads or writes rather than what it is applied to.

44. Where deferring work changes the cost

Real world

Producing something on demand rather than in advance.

Discussion prompt

Think of something prepared only when asked for rather than in advance. What does that change, and what does it not?

Hint: The total work, or when it happens?

Answer:

It changes when the work happens and how much is stored, and it does not change the total — if everything is eventually asked for, the same amount gets done.

What it can change is whether everything is asked for. Stopping early means the unmade parts are never made, which is a genuine saving rather than a rescheduling.

Which is exactly the iterator: keys, values and items are constant time because they return iterators, and the loop over them is linear. The deferral saves memory always and time only if you stop.

45. What the analysis is worth

Section

Section 5

46. Useful, with caveats the book states itself

Concept

Programmers who care about performance often find this kind of analysis hard to swallow. They have a point.

But if you keep those caveats in mind, algorithmic analysis is a useful tool — at least for large problems, the better algorithm is usually better, and sometimes it is much better.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 203-203

47. Picture it: the claim, and its scope

Picture it

The strongest sentence and the three qualifications on it.

Figure (svg): Two columns setting the value of the analysis against its stated limits

Programmers who care about performance often find this hard to swallow. They have a point.

Which is the same shape as the reservations in chapter 19: the book recommends the tool and states where it does not apply.

48. Worked example: the sentence that justifies the whole appendix

Worked example

One comparison that decides when to care.

# 'The difference between two algorithms with the same
#  order of growth is usually a constant factor, but the
#  difference between a good algorithm and a bad
#  algorithm is unbounded!'
ComparisonThe differenceNote
same classa constant factortwice as fast, forever
different classunboundedthe ratio grows with n
the consequencethe class matters more than the constantat scale

Note the same-class case.

Why: Two algorithms with the same order of growth differ by a constant factor — one may be twice as fast, and it stays twice as fast at every size.

Note the different-class case.

Why: The difference between a good algorithm and a bad one is unbounded: the ratio itself grows as the problem does.

Draw the conclusion.

Why: Which is why the class is worth optimising before the constant, and why the constant is worth optimising once the class is right.

Figure (svg): A growth chart contrasting a constant-factor difference with an unbounded one

One gap stays proportional forever; the other widens without limit.

A bounded difference against an unbounded one, which is what makes the classification worth the abstraction it costs.

Verify: Check it against chapter 11's memo.

Why: Memoizing fibonacci changed an exponential algorithm into a linear one — an unbounded improvement, which is why it turned an unusable program into an instant one. Replacing a loop with sum is a constant-factor improvement, which is measurable and never transformative.

49. Predict: which improvement is unbounded?

Prediction

Two kinds of speed-up.

# A: replacing a loop with a faster equivalent loop
# B: replacing an O(n^2) algorithm with an O(n) one
ChangeThe kind of improvementNote
Athe same classa constant factor
Ba better classthe ratio grows
at large nonly B keeps improving

Predict first

Which change gives an unbounded improvement?

  • B — the difference between a good algorithm and a bad one is unbounded
  • A — a faster implementation always wins
  • Both are unbounded
  • Neither; both are constant factors

Correct: B — the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.

Why: A faster implementation of the same algorithm stays a fixed multiple faster at every size; a better order of growth means the ratio itself grows with the problem. That is why the class is worth improving first, and why chapter 11's memo transformed a program where a faster loop would only have sped it up.

50. Worked example: the book's own advice about sorting

Worked example

From the footnote on the opening page.

# 'The fastest way to sort a million integers is to use
#  whatever sort function is provided by the language
#  I'm using. Its performance is good enough for the
#  vast majority of applications, but if it turned out
#  that my application was too slow, I would use a
#  profiler to see where the time was being spent.'
StageWhat to doNote
the first answeruse the built-ingood enough almost always
if too slowprofilefind where the time goes
only thenconsider a better algorithm

Start with what exists.

Why: Its performance is good enough for the vast majority of applications — which is chapter 13's advice about choosing the easiest implementation.

Measure before optimising.

Why: If it turned out that my application was too slow, I would use a profiler to see where the time was being spent.

Then act on the measurement.

Why: If it looked like a faster sort algorithm would have a significant effect on performance, then I would look around for a good implementation.

Figure (svg): The state of the program after each line of Worked example the book's own advice about sorting, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A procedure rather than a preference, and the same one chapter 13 gave: easiest first, measure, then improve what the measurement points at.

Verify: Note where the analysis fits into it.

Why: Not at the first step and not at the second — it tells you what to reach for once the profiler has identified the bottleneck. So the appendix is a tool for the third stage, which is exactly why it comes at the end of the book.

51. Trap: optimising the class before measuring

Trap

The trap

A program is rewritten to improve an algorithm's order of growth before anyone has timed it.

Apply what the appendix teaches

Why: A better class is better, and unboundedly so.

Not if that part is not where the time goes, and not if the sizes are below the crossover. The book's own advice is to use a profiler to see where the time is being spent, and only then to look for a faster algorithm.

The fix

Measure, then use the analysis.

Choose the easiest implementation first

Why: Which is chapter 13's rule.

Profile if it is too slow

Why: To find where the time actually goes.

The analysis is what tells you which replacement to look for once you know what to replace. Used before the measurement, it optimises a guess.

52. Sort: when does the analysis help?

Sorting

The book states its own limits.

Sort into buckets

For each situation, is order of growth the right tool?

the analysis helps
choosing between algorithms for a very large dataset; explaining why a program is unusable at scale; deciding whether a dictionary or a list suits repeated lookups
it does not
deciding between two O(n log n) sorts; a list of twenty items; predicting exact run times in seconds
yes
Each is a question about how cost behaves as size grows, which is exactly what an order of growth describes — and at least for large problems, the better algorithm is usually better.
no
Each falls under a stated caveat: same-class algorithms are not distinguished, small problems make order of growth irrelevant, and the classification discards the constants that would give a time in seconds.

53. Complete it: the closing comparison

Faded example

Two kinds of difference.

Fill in the blanks

# same order of growth: usually a constant factor
# good algorithm vs bad: unbounded

Why: That contrast is what justifies the abstraction: a constant factor is a fixed multiple at every size, and an unbounded difference means the ratio itself grows with the problem. It is why the order of growth is worth improving before the coefficient.

54. Explain it: why bother with all this abstraction?

Explain it

The analysis discards a lot.

Discussion prompt

A classmate objects that ignoring coefficients and hardware makes the analysis useless. Give them the book's own answer.

Hint: The book agrees with them, partly.

Answer:

They have a point, and the book says so: programmers who care about performance often find this kind of analysis hard to swallow, because sometimes the coefficients and the details really do make a difference, and for small problems order of growth is irrelevant.

What it buys is that at least for large problems, the better algorithm is usually better, and sometimes much better — and it gives a simple classification that transfers between machines and languages, which a timing does not.

The decisive line is the comparison: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded. Constants are worth a lot; classes are worth more.

55. Compare: constant, linear and quadratic

Comparison

Fill the blanks. Three shapes rather than three speeds.

Comparison matrix

QuestionO(1)O(n)O(n squared)
How does cost respond to size?not at allproportionallyas the square
An example operationindexing a list or dictionarysumming a lista loop inside a loop over the same list
At a million itemsabout one stepa million stepsa million million
Is the gap to the next class fixed?no — it grows without limitno — likewiseno

The bottom row is why classes matter more than constants: within a class the difference is fixed, and between classes it is unbounded.

56. The procedure: working out an order of growth

Pattern

Five steps, and the third is the one that composes.

  1. Express the run time as a function of the problem size rather than as a time.
  2. Identify the leading term — the one with the highest exponent — and discard the rest.
  3. For a loop, add one to the exponent of its body's class, unless the loop runs a constant number of times.
  4. Look up the operations you use: indexing is constant, traversal is linear, sorting is O(n log n).
  5. Then check the caveats: your sizes may be below the crossover, and the constants may matter.

Step 5 is not optional. For small problems order of growth is irrelevant, and the location of the crossover depends on details the classification threw away.

Python documentation — Built-in Types Built-in Types

57. Check yourself 1 of 3: the leading term

Check

Coefficients do not count.

Check your understanding

What is the order of growth of 1000000n cubed + n squared?

  • A. O(n cubed) — cubic (correct)
  • B. O(n squared) — quadratic
  • C. O(1000000n cubed)
  • D. It cannot be classified

Answer: A

Why: The leading term is the term with the highest exponent, and all functions with that leading term belong to the same class regardless of their coefficients. The million is real and decides where the crossover with another algorithm falls, which is not part of the classification.

Why B tempts people
n squared is a non-leading term here, and non-leading terms do not affect the class.
Why C tempts people
Coefficients are not written in Big-Oh notation, because functions with the same leading term are considered equivalent.
Why D tempts people
It classifies easily; the classification is designed to be coarse.

58. Check yourself 2 of 3: operation costs

Check

One of these is not constant.

Check your understanding

Which of these is linear rather than constant time?

  • A. sum(t) (correct)
  • B. t[i]
  • C. len(t)
  • D. t.append(x)

Answer: A

Why: Summing has to look at every element, so it is linear — the built-in is faster than an equivalent loop because it has a smaller leading coefficient, not because it is in a different class. Indexing and len are constant regardless of the size, and appending is constant on average.

Why B tempts people
Indexing operations are constant time, regardless of the size of the data structure.
Why C tempts people
len reads a stored size rather than counting.
Why D tempts people
Constant on average: occasional copying to a bigger location averages out over n operations.

59. Check yourself 3 of 3: the caveats

Check

The book states its own limits.

Check your understanding

When is order of growth NOT a useful guide?

  • A. For small problems, and between two algorithms in the same class (correct)
  • B. For large problems
  • C. When comparing different languages
  • D. It is always useful

Answer: A

Why: For small problems, order of growth is irrelevant — the book's own table shows the quadratic algorithm winning at n = 10. And if two algorithms have the same leading order term, it is hard to say which is better; the answer depends on the details.

Why B tempts people
Large problems are exactly where it is most useful: the better algorithm is usually better, and sometimes much better.
Why C tempts people
Language differences are constant factors, which is what the classification abstracts away — so comparisons still transfer.
Why D tempts people
The book lists three caveats explicitly, and says that programmers who care about performance have a point.

60. Where this shows up outside this course

Real world

A method that works at one scale and not at another.

Discussion prompt

Think of a way of doing something that is fine for a few items and impossible for many. What changes, and at roughly what point?

Hint: Checking each against each.

Answer:

Comparing every item with every other, checking a list by reading it from the start, or handling each case by hand — each is fine for a handful and impossible for thousands.

What changes is not the difficulty of one step but how many steps there are, and the point where it becomes impossible is the crossover: below it the simple method wins, above it nothing can rescue it.

Which is exactly the appendix's subject. The classification exists to name the shape of that growth, and the caveat exists because the crossover is real and its location depends on details the classification discards.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

Algorithm A takes 100n + 1 steps and Algorithm B takes n squared + n + 1. At n = 10, which is faster?

  • B, by almost ten times — order of growth describes large n, and ten is not large
  • A, because linear always beats quadratic
  • They are equal at every size
  • It depends on the hardware

Correct: B, by almost ten times — at n = 10, Algorithm A looks pretty bad.

Why: This is the book's own table and it is there to make exactly this point. At n = 10 the quadratic algorithm takes 111 steps and the linear one takes 1,001; at n = 100 they are about the same; and beyond that A is much better and the gap widens without limit. The reason is the large coefficient on A's leading term, and the reason it does not save B is that regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n. That point is the crossover, and its location depends on the details of the algorithms, the inputs and the hardware — so it is usually ignored for purposes of algorithmic analysis. The book's warning is the part worth keeping: that doesn't mean you can forget about it. For small problems, order of growth is irrelevant, and if your problem sizes sit below the crossover, the theoretically worse algorithm is the faster one.

62. Explain it to someone else

Explain it

Why the coefficients are thrown away.

Discussion prompt

A classmate objects that ignoring a coefficient of a hundred is absurd, since it is a hundredfold difference. Answer them.

Hint: What does the coefficient depend on?

Answer:

A hundredfold difference is real and it is fixed — it stays a hundredfold at every size. The exponent's effect is not fixed: the ratio between n and n squared grows without limit as n does.

And the coefficient depends on the machine, the language and the implementation, so it is not a property of the algorithm. The exponent is, which is what lets a comparison transfer between machines.

The book's own sentence is the summary: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • The leading term, and why coefficients are discarded
  • The orders of growth, and their relative badness
  • The costs of the Python operations you have been using
  • The caveats, and when the analysis does not apply

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The leading term is the whole mechanism, and the argument for discarding coefficients — that the exponent is a property of the algorithm and the coefficient is a property of the machine — is worth being able to make. The seven classes are worth memorising in order, since the ordering is the useful part. The operation costs retrospectively explain several choices the course made much earlier. And the caveats matter because the book states them itself: for small problems this is irrelevant, and the crossover point is real.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the two run-time curves from the book's table and mark the crossover. Beside them, list the seven orders of growth in increasing order of badness. Underneath, make two columns of Python operations — constant and linear — and add the four exceptions with what each one's cost is actually proportional to. Finally write the sentence about constant factors and unbounded differences.

65. What you can do now

Recap

Four pages, and a name for something the course kept demonstrating.

If you remember one thingIt is this
From the tableAt small n the worse algorithm can win by tenfold.
From the leading termThe exponent belongs to the algorithm; the coefficient belongs to the machine.
From the loop ruleEach level of nesting adds one to the exponent.
From the exceptionsCount what the operation reads, not what it is applied to.
From the closing lineSame class: a constant factor. Different class: unbounded.

The last lesson applies all of this to searching: why the in operator on a list is linear, how a bisection search finds a word among a million in twenty steps, and how a hashtable makes dictionary lookup take about the same time however large the dictionary gets.

Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-204 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §B.1-B.2, pp. 201-204
  2. Python documentation — Built-in Types
  3. Python documentation — Data Structures
  4. Open Data Structures (Python edition) — Pat Morin — Pat Morin, Open Data Structures (Python edition), opendatastructures.org

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108