This lesson introduces order of growth as a way of comparing algorithms, works through why the leading term dominates, and catalogues the run time of the Python operations the course has been using.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Appendix B — Analysis of Algorithms
§B.1-B.2, pp. 201-204
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-204 — the pages these objectives are drawn from
Warm-up
Chapter 11 said so without explaining it.
Discussion prompt
Lesson 11a said the in operator takes about the same time for a dictionary however many items it has, and gets slower for a list as the list grows. What kind of claim is that, and what would it take to justify it?
Hint: It is not a claim about one program on one machine.
Answer:
It is a claim about how the cost behaves as the size grows, rather than about any particular timing — a shape rather than a number.
Justifying it needs a way of talking about that shape which does not depend on the machine, the data or the size. That is what this appendix provides.
And it names something the course has demonstrated repeatedly: the list search in chapter 13, the memo in chapter 11, and the list-and-join in chapter 18 were all this idea without the vocabulary.
Concept
The practical goal of algorithm analysis is to predict the performance of different algorithms in order to guide design decisions.
That third response is the one this appendix builds on: group functions into categories depending on how quickly they grow as problem size increases.
Figure (svg): Two columns pairing each obstacle to comparing algorithms with the standard response
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-202
Section
Section 1
Concept
Suppose Algorithm A takes 100n + 1 steps to solve a problem with size n, and Algorithm B takes n squared plus n plus 1.
leading term — In an expression describing run time, the term with the highest exponent.
# size A: 100n+1 B: n^2+n+1
# 10 1 001 111
# 100 10 001 10 101
# 1000 100 001 1 001 001
# 10000 1 000 001 100 010 001| Size | What happens | Note |
|---|---|---|
| n = 10 | A takes almost 10 times longer | B looks better |
| n = 100 | about the same | the crossover |
| larger n | A is much better | and the gap widens |
The fundamental reason is that for large values of n, any function that contains an n-squared term will grow faster than a function whose leading term is n — regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n, for any a and b.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 202-202
Picture it
One algorithm wins at small sizes and loses at large ones.
Figure (svg): A growth chart showing a linear and a quadratic run time crossing over
Which is why the analysis is about large problems: for small ones, order of growth is irrelevant, and the book says so.
Worked example
Even a very large one.
# A: 100n + 1 - a big coefficient
# B: n^2 + n + 1 - a small one
# and even if A were n + 1000000,
# it would still be better than B
# for sufficiently large n| Range | What decides | Note |
|---|---|---|
| at small n | the coefficient dominates | B wins |
| at large n | the exponent dominates | A wins |
| always | some n exists where it flips | for any coefficients |
Note what the coefficient does.
Why: For Algorithm A, the leading term has a large coefficient, 100, which is why B does better than A for small n.
Note what it cannot do.
Why: Regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n, for any values of a and b.
Note that non-leading terms are the same.
Why: Even if the run time of Algorithm A were n + 1000000, it would still be better than Algorithm B for sufficiently large n.
Figure (svg): The state of the program after each line of Worked example why the coefficient cannot save it, drawn as a ladder with one rung per traced line
The exponent wins eventually, whatever the constants. That is what makes the leading term the thing worth comparing.
Verify: Check where the crossover moves.
Why: A bigger coefficient on A pushes the crossover to a larger n without preventing it — so the constants decide where the change happens rather than whether. That is why they are ignored for classification and still matter in practice.
Prediction
The table's smallest size.
# A: 100n + 1 at n = 10 -> 1001
# B: n^2 + n + 1 at n = 10 -> 111| Algorithm | Steps | Note |
|---|---|---|
| A | 1001 steps | the big coefficient |
| B | 111 steps | the small n |
| at n = 10 | B is faster | by almost ten times |
Predict first
At n = 10, which algorithm is faster?
Correct: B, by almost ten times — at n = 10, Algorithm A looks pretty bad.
Why: Order of growth describes behaviour for large n, and ten is not large. The crossover is around n = 100, above which A wins by a widening margin — which is why the book says that for small problems, order of growth is irrelevant.
Worked example
Ignored for analysis, and not for engineering.
# 'In general, we expect an algorithm with a smaller
# leading term to be a better algorithm for large
# problems, but for smaller problems, there may be a
# crossover point where another algorithm is better.'
# 'it is usually ignored for purposes of algorithmic
# analysis. But that doesn't mean you can forget
# about it.'| Aspect | What is true | Note |
|---|---|---|
| where it is | depends on the algorithms, inputs and hardware | not general |
| for analysis | ignored | it is not a property of the class |
| for a real program | may be exactly your range |
Note that it exists.
Why: For smaller problems there may be a crossover point where another algorithm is better — which the table shows at around n = 100.
Note why analysis ignores it.
Why: The location depends on the details of the algorithms, the inputs and the hardware, so it is not a property of the growth class.
Note the caveat.
Why: But that doesn't mean you can forget about it — if your problem sizes sit below the crossover, the theoretically worse algorithm is the faster one.
Figure (svg): Two columns separating what order of growth captures from what it discards
A real quantity that the classification deliberately discards. Knowing that it was discarded is what stops the analysis being misapplied.
Verify: Connect it to the course's own examples.
Why: Lesson 19b's has_duplicates is exactly this: the set version is shorter and the loop version stops at the first duplicate, so for a long list with an early repeat the loop wins. Order of growth would not distinguish them, and the details do.
Trap
An algorithm is replaced by one with a better order of growth, on a problem with a few dozen items.
Prefer the better class
Why: Which is right for large problems, and this is the standard advice.
For small problems, order of growth is irrelevant — the book says so — and the crossover may be far above the sizes involved. The replacement can be slower and is certainly more work.
Check the sizes you actually have.
Below the crossover, the constants decide
Why: And they are what the classification discarded.
Above it, the order of growth decides
Why: And the difference can be unbounded.
This is chapter 13's advice again: choose the structure that is easiest to implement and see whether it is fast enough. Order of growth tells you what to do when it is not.
Discrimination
The term with the highest exponent.
Sort into buckets
For each run time, what is the leading term?
Faded example
The term with the highest exponent.
Fill in the blanks
# run time: 1000000n^3 + n^2
# leading term: n^3
Why: The leading term is the term with the highest exponent, so the coefficient of a million is irrelevant to the classification — this is cubic. Coefficients decide where the crossover is, and exponents decide which algorithm eventually wins.
Socratic
They are real, and the analysis discards them.
Discussion prompt
A coefficient of 100 makes a real difference to how long a program takes. Why does the classification throw it away?
Hint: What does the coefficient depend on?
Answer:
Because it depends on things the analysis is trying to abstract away: the hardware, the language, the implementation. The same algorithm has a different coefficient on a different machine.
The exponent does not. It is a property of the algorithm itself, so it is the part that transfers between machines and implementations — which is what makes a general comparison possible at all.
And the book's closing line puts the trade precisely: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.
Section
Section 2
Concept
An order of growth is a set of functions whose growth behaviour is considered equivalent. For example, 2n, 100n and n + 1 belong to the same order of growth, written O(n) and often called linear.
# O(1) constant
# O(log n) logarithmic
# O(n) linear
# O(n log n) linearithmic
# O(n^2) quadratic
# O(n^3) cubic
# O(c^n) exponential| Aspect | What is true | Note |
|---|---|---|
| the list | in increasing order of badness | the book's phrase |
| the base of a logarithm | does not matter | changing bases multiplies by a constant |
| the base of an exponent | likewise | all exponentials are one class |
For the logarithmic terms, the base doesn't matter — changing bases is the equivalent of multiplying by a constant, which doesn't change the order of growth. Exponential functions grow very quickly, so exponential algorithms are only useful for small problems.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 203-203
Picture it
In increasing order of badness.
Figure (svg): The seven common orders of growth listed from best to worst
And the top two are worth noticing: constant time means the size does not matter at all, and logarithmic means it barely does.
Worked example
A million items, and how much work each class implies.
# n = 1,000,000
# O(1) about 1 step
# O(log n) about 20 steps
# O(n) a million steps
# O(n log n) about 20 million
# O(n^2) a million million| Class | At a million items | Note |
|---|---|---|
| constant | unaffected by size | 1 |
| logarithmic | barely affected | 20 |
| quadratic | a trillion steps | impractical |
Note the top of the list.
Why: Constant and logarithmic barely respond to size at all — twenty steps for a million items is the bisection search of the next lesson.
Note the middle.
Why: Linear and linearithmic scale usefully: a million steps is fast, and twenty million is still fine.
Note the bottom.
Why: Quadratic at a million items is a million million steps, which is why the difference between classes is not a constant factor.
Figure (svg): A growth chart comparing constant, logarithmic, linear and quadratic growth
Five classes and five wildly different amounts of work for one problem size. That spread is what the classification is for.
Verify: Check the claim about unboundedness.
Why: Between O(n) and O(n squared) at a million items the ratio is a million — and at ten million it is ten million. So the difference between a good algorithm and a bad algorithm is unbounded, exactly as the book says, where two algorithms in the same class differ by a fixed factor forever.
Sorting
In increasing order of badness.
Sort into buckets
For each pair, which order of growth is better for large problems?
Worked example
For logarithms and for exponentials, differently.
# log base 2 of n, and log base 10 of n,
# differ by a constant factor
# -> the same order of growth
# 2^n and 3^n are both exponential
# -> the same order of growth| Change | Effect on the class | Note |
|---|---|---|
| changing log bases | multiplies by a constant | no change of class |
| changing exponent bases | all exponential | one class |
| what that means | the classes are coarse | deliberately |
Note the logarithm rule.
Why: Changing bases is the equivalent of multiplying by a constant, which doesn't change the order of growth — so O(log n) needs no base.
Note the exponential rule.
Why: All exponential functions belong to the same order of growth regardless of the base of the exponent.
Note the coarseness.
Why: The classification is deliberately loose, because its purpose is to distinguish shapes rather than to rank precisely.
Figure (svg): The state of the program after each line of Worked example why the base does not matter, drawn as a ladder with one rung per traced line
Two rules that keep the classes small in number. The coarseness is the feature: seven classes cover most of what you need to distinguish.
Verify: Ask whether the coarseness ever misleads.
Why: It can: two O(n log n) sorts can differ by a large constant, and the book says that if two algorithms have the same leading order term, it is hard to say which is better — the answer depends on the details. So the classification is a first filter rather than a final answer.
Trap
An O(n log n) algorithm is assumed to be faster than an O(n) one on some particular input.
Read the classes as a ranking
Why: They are listed in increasing order of badness.
The classes describe growth rather than speed, and for a specific size the constants can reverse the order. Two algorithms in the same class can also differ by a large constant factor.
Read the class as a shape.
It says how the cost responds to size
Why: Not how long anything takes.
And a comparison at one size needs measurement
Why: Which is chapter 13's benchmarking.
The book's own hedge applies: sometimes the coefficients and the non-leading terms make a real difference, and for small problems order of growth is irrelevant. It is a tool with caveats rather than a verdict.
Prediction
Different coefficients, same exponent.
# 2n, 100n, and n + 1| Function | Leading term | Class |
|---|---|---|
| all three | leading term n | linear |
| the coefficients | differ | and do not matter |
| the class | O(n) |
Predict first
Do 2n, 100n and n + 1 belong to the same order of growth?
Correct: Yes — all are O(n), and every function in the set grows linearly with n.
Why: An order of growth is a set of functions whose growth behaviour is considered equivalent, and functions with the same leading term are considered equivalent even if they have different coefficients. The differences are real and are constant factors, which the classification deliberately discards.
Faded example
The leading term names it.
Fill in the blanks
# run time: n^2 + n + 1
# order of growth: O(n^2), called quadratic
Why: All functions with the leading term n squared belong to O(n squared), and they are called quadratic. The n and the 1 are non-leading terms, which do not affect the class — only the term with the highest exponent does.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C ignores the crossover point. For smaller problems there may be a point where another algorithm is better — the book's own table shows the quadratic algorithm winning by almost tenfold at n = 10. Order of growth is about behaviour for large n, and for small problems it is irrelevant.
Section
Section 3
Concept
Most arithmetic operations are constant time, and indexing operations — reading or writing elements in a sequence or dictionary — are also constant time, regardless of the size of the data structure.
Very large integers are an exception to the arithmetic rule; in that case the run time increases with the number of digits.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 204-204
Picture it
Three groups, and the exceptions are where the interest is.
Figure (svg): Two columns separating constant-time operations from linear ones
And two exceptions are worth their own slides: appending to a list, which is constant on average, and dictionary operations, which are the minor miracle of the next lesson.
Worked example
A loop multiplies the class of its body.
total = 0
for x in t:
total += x
# the body is constant, so the loop is linear
# rule of thumb: if the body is in O(n^a),
# the whole loop is in O(n^(a+1))| Part | Its class | Note |
|---|---|---|
| a constant body | O(1) | one addition |
| the loop | n passes | O(n) |
| a linear body | would give O(n squared) | a nested loop |
Note the base case.
Why: A for loop that traverses a sequence or dictionary is usually linear, as long as all of the operations in the body are constant time.
Note the rule.
Why: If the body of a loop is in O(n^a) then the whole loop is in O(n^(a+1)).
Note the exceptions.
Why: Unless the loop exits after a constant number of iterations — if a loop runs k times regardless of n, the loop is in O(n^a), even for large k.
Figure (svg): A flowchart showing how a loop's class is derived from its body's
A rule that composes: each level of nesting adds one to the exponent, which is why a nested loop over the same sequence is quadratic.
Verify: Check the division case.
Why: Multiplying by k doesn't change the order of growth, and neither does dividing — so a loop whose body is O(n^a) running n/k times is still O(n^(a+1)), even for large k. Halving the number of passes is a constant factor, not a change of class.
Sorting
Ask what the operation touches.
Sort into buckets
For each operation, which class is it?
Worked example
An exception worth understanding.
# 'Adding an element to the end of a list is constant
# time on average; when it runs out of room it
# occasionally gets copied to a bigger location, but
# the total time for n operations is O(n), so the
# average time for each operation is O(1).'| Case | What it costs | Note |
|---|---|---|
| most appends | constant | there is room |
| occasionally | a copy to a bigger location | linear, that once |
| over n operations | O(n) in total | so O(1) each on average |
Note the usual case.
Why: Adding to the end costs a constant amount when there is room, which is nearly always.
Note the occasional case.
Why: When it runs out of room it occasionally gets copied to a bigger location, which is proportional to the current size.
Note how they average.
Why: The total time for n operations is O(n), so the average time for each operation is O(1) — the expensive copies are rare enough to disappear into the average.
Figure (svg): The state of the program after each line of Worked example why append is constant on average, drawn as a ladder with one rung per traced line
Constant time on average, which is a different claim from constant time always. That distinction is why the book states it carefully rather than just listing append as O(1).
Verify: Connect it to lesson 18b's string building.
Why: That lesson's list-and-join relies on this: appending is cheap on average, where repeated string concatenation copies everything every time and is genuinely linear per operation. The two techniques differ precisely because one has this amortisation and the other does not.
Trap
A slice of two elements from a million-element list is assumed to be expensive.
Scale the cost with the container
Why: Most operations on a large structure cost more.
The run time of a slice operation is proportional to the length of the output, but independent of the size of the input — so a two-element slice is cheap however large the list.
Ask what the operation actually touches.
A slice touches the output
Why: So its cost is the output's length.
An index touches one element
Why: So it is constant regardless of the size of the data structure.
The rule that generalises is to count what the operation reads or writes rather than what it is applied to. Indexing, slicing and len all illustrate it differently.
Prediction
A constant-time body.
total = 0
for x in t:
total += x| Part | Its class | Note |
|---|---|---|
| the body | one addition | constant |
| the loop | n passes | |
| the whole thing | O(n) | linear |
Predict first
What is the order of growth of this loop?
Correct: O(n) — a for loop that traverses a sequence is usually linear, as long as the operations in the body are constant time.
Why: The rule of thumb is that if the body is in O(n^a), the whole loop is in O(n^(a+1)) — a constant body gives a linear loop. The built-in sum is also linear because it does the same thing, though it tends to be faster: in the language of algorithmic analysis, it has a smaller leading coefficient.
Faded example
Nesting adds to the exponent.
Fill in the blanks
# if the body of a loop is in O(n^a),
# the whole loop is in O(n^(a+1))
Why: Each level of looping multiplies the work by n, which adds one to the exponent — so a constant body gives a linear loop and a linear body gives a quadratic one. The exception is a loop that runs a constant number of times regardless of n, which does not change the class at all.
Explain it
Both are linear.
Discussion prompt
A classmate replaced their summing loop with sum and it got noticeably faster. Explain, given that both are linear.
Hint: What does the class not capture?
Answer:
They are in the same order of growth: the built-in sum is also linear because it does the same thing. So the class does not distinguish them, and the class is not the whole story.
What differs is the constant. In the language of algorithmic analysis, sum has a smaller leading coefficient — it is a more efficient implementation of the same algorithm.
Which is exactly what the classification discards. Two algorithms with the same order of growth differ by a constant factor, and constant factors are real: this one is measurable, and it is not what order of growth is for.
Section
Section 4
Concept
The general rules cover most operations, and the exceptions are the ones that decide real programs.
That last one is lesson 12b's iterator point in cost terms: producing the iterator is cheap because nothing is computed, and the work happens when it is walked.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 204-205
Picture it
Each contradicts a reasonable guess.
Figure (svg): Four operations whose cost differs from what the general rule would suggest
And the fourth explains something from chapter 12: an iterator is cheap to make because it has computed nothing yet.
Worked example
The cost has been deferred rather than avoided.
d.keys() # constant: nothing computed
for k in d.keys(): # linear: now it walks
...| Operation | Its class | Note |
|---|---|---|
| making the iterator | constant | it knows how, and has not |
| walking it | linear | one step per item |
| the total | linear | the work was deferred |
Note why creation is cheap.
Why: keys, values and items are constant time because they return iterators — no values have been produced.
Note where the work happens.
Why: But if you loop through the iterators, the loop will be linear, because that is when each item is produced.
Connect it to chapter 12.
Why: This is why a zip object printed as an object description rather than as pairs: it had computed nothing, which is also why it was cheap.
Figure (svg): The state of the program after each line of Worked example why iterators are constant time, drawn as a ladder with one rung per traced line
Deferred rather than avoided. The total work is the same; what changes is when it happens, and whether it happens at all if the iterator is abandoned.
Verify: Ask when the deferral is a genuine saving.
Why: When you stop early — any over a generator that finds a True value immediately does one step rather than n, which is lesson 19a's point about the brackets. If the whole thing is walked, the deferral saves the memory rather than the time.
Prediction
The list has a million elements.
t = list(range(1000000))
s = t[0:2]| Aspect | What matters | Note |
|---|---|---|
| the input | a million elements | irrelevant |
| the output | two elements | the cost |
| the class | proportional to the output |
Predict first
How expensive is this slice?
Correct: Cheap — the run time of a slice operation is proportional to the length of the output, but independent of the size of the input.
Why: The operation only has to copy the elements it produces, so a two-element slice costs the same from a million-element list as from a ten-element one. Counting what the operation reads or writes is the habit that gets these right.
Worked example
Between linear and quadratic.
t.sort() # O(n log n)
# at n = 1,000,000:
# linear: 1,000,000 steps
# n log n: ~20,000,000 steps
# quadratic: 1,000,000,000,000 steps| Comparison | The ratio | Note |
|---|---|---|
| against linear | about twenty times more | at a million |
| against quadratic | fifty thousand times less | |
| in practice | sorting a million items is fast |
Place it on the list.
Why: O(n log n) is linearithmic, between linear and quadratic in the book's table.
Note what that means at scale.
Why: Twenty times a linear pass at a million items, which is a small multiple rather than a different order of difficulty.
Note the contrast.
Why: A quadratic sort at the same size is a million times worse — which is why bubble sort is conceptually simple but slow for large datasets.
Figure (svg): A growth chart placing linearithmic between linear and quadratic
A class much closer to linear than to quadratic in practice. That is why sorting a large collection is a reasonable thing to do and a quadratic algorithm on one is not.
Verify: Connect it to the Obama anecdote.
Why: The appendix opens with a candidate being asked for the most efficient way to sort a million 32-bit integers and replying that bubble sort would be the wrong way to go — which is true, because bubble sort is quadratic and a million squared is a million million.
Trap
A small dictionary is updated into a very large one, and the operation is assumed to be expensive.
Scale the cost with the receiver
Why: It is the big one being changed.
The run time of update is proportional to the size of the dictionary passed as a parameter, not the dictionary being updated — so merging a small one into a large one is cheap.
Count what the operation reads.
update reads the argument
Why: One insertion per item in it.
And each insertion is constant
Why: Because dictionary operations are.
The general habit is the useful one: ask what the operation has to look at. A slice looks at its output, an index looks at one element, and update looks at its argument.
Discrimination
Four exceptions to the general rules.
Sort into buckets
For each operation, what is its cost proportional to?
Faded example
Count what it produces.
Fill in the blanks
# a slice's run time is proportional to the
# length of the output, not of the input
Why: The operation copies only the elements it produces, so a short slice of a very long sequence is cheap. That is an instance of the general habit worth forming: count what an operation actually reads or writes rather than what it is applied to.
Real world
Producing something on demand rather than in advance.
Discussion prompt
Think of something prepared only when asked for rather than in advance. What does that change, and what does it not?
Hint: The total work, or when it happens?
Answer:
It changes when the work happens and how much is stored, and it does not change the total — if everything is eventually asked for, the same amount gets done.
What it can change is whether everything is asked for. Stopping early means the unmade parts are never made, which is a genuine saving rather than a rescheduling.
Which is exactly the iterator: keys, values and items are constant time because they return iterators, and the loop over them is linear. The deferral saves memory always and time only if you stop.
Section
Section 5
Concept
Programmers who care about performance often find this kind of analysis hard to swallow. They have a point.
But if you keep those caveats in mind, algorithmic analysis is a useful tool — at least for large problems, the better algorithm is usually better, and sometimes it is much better.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 203-203
Picture it
The strongest sentence and the three qualifications on it.
Figure (svg): Two columns setting the value of the analysis against its stated limits
Which is the same shape as the reservations in chapter 19: the book recommends the tool and states where it does not apply.
Worked example
One comparison that decides when to care.
# 'The difference between two algorithms with the same
# order of growth is usually a constant factor, but the
# difference between a good algorithm and a bad
# algorithm is unbounded!'| Comparison | The difference | Note |
|---|---|---|
| same class | a constant factor | twice as fast, forever |
| different class | unbounded | the ratio grows with n |
| the consequence | the class matters more than the constant | at scale |
Note the same-class case.
Why: Two algorithms with the same order of growth differ by a constant factor — one may be twice as fast, and it stays twice as fast at every size.
Note the different-class case.
Why: The difference between a good algorithm and a bad one is unbounded: the ratio itself grows as the problem does.
Draw the conclusion.
Why: Which is why the class is worth optimising before the constant, and why the constant is worth optimising once the class is right.
Figure (svg): A growth chart contrasting a constant-factor difference with an unbounded one
A bounded difference against an unbounded one, which is what makes the classification worth the abstraction it costs.
Verify: Check it against chapter 11's memo.
Why: Memoizing fibonacci changed an exponential algorithm into a linear one — an unbounded improvement, which is why it turned an unusable program into an instant one. Replacing a loop with sum is a constant-factor improvement, which is measurable and never transformative.
Prediction
Two kinds of speed-up.
# A: replacing a loop with a faster equivalent loop
# B: replacing an O(n^2) algorithm with an O(n) one| Change | The kind of improvement | Note |
|---|---|---|
| A | the same class | a constant factor |
| B | a better class | the ratio grows |
| at large n | only B keeps improving |
Predict first
Which change gives an unbounded improvement?
Correct: B — the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.
Why: A faster implementation of the same algorithm stays a fixed multiple faster at every size; a better order of growth means the ratio itself grows with the problem. That is why the class is worth improving first, and why chapter 11's memo transformed a program where a faster loop would only have sped it up.
Worked example
From the footnote on the opening page.
# 'The fastest way to sort a million integers is to use
# whatever sort function is provided by the language
# I'm using. Its performance is good enough for the
# vast majority of applications, but if it turned out
# that my application was too slow, I would use a
# profiler to see where the time was being spent.'| Stage | What to do | Note |
|---|---|---|
| the first answer | use the built-in | good enough almost always |
| if too slow | profile | find where the time goes |
| only then | consider a better algorithm |
Start with what exists.
Why: Its performance is good enough for the vast majority of applications — which is chapter 13's advice about choosing the easiest implementation.
Measure before optimising.
Why: If it turned out that my application was too slow, I would use a profiler to see where the time was being spent.
Then act on the measurement.
Why: If it looked like a faster sort algorithm would have a significant effect on performance, then I would look around for a good implementation.
Figure (svg): The state of the program after each line of Worked example the book's own advice about sorting, drawn as a ladder with one rung per traced line
A procedure rather than a preference, and the same one chapter 13 gave: easiest first, measure, then improve what the measurement points at.
Verify: Note where the analysis fits into it.
Why: Not at the first step and not at the second — it tells you what to reach for once the profiler has identified the bottleneck. So the appendix is a tool for the third stage, which is exactly why it comes at the end of the book.
Trap
A program is rewritten to improve an algorithm's order of growth before anyone has timed it.
Apply what the appendix teaches
Why: A better class is better, and unboundedly so.
Not if that part is not where the time goes, and not if the sizes are below the crossover. The book's own advice is to use a profiler to see where the time is being spent, and only then to look for a faster algorithm.
Measure, then use the analysis.
Choose the easiest implementation first
Why: Which is chapter 13's rule.
Profile if it is too slow
Why: To find where the time actually goes.
The analysis is what tells you which replacement to look for once you know what to replace. Used before the measurement, it optimises a guess.
Sorting
The book states its own limits.
Sort into buckets
For each situation, is order of growth the right tool?
Faded example
Two kinds of difference.
Fill in the blanks
# same order of growth: usually a constant factor
# good algorithm vs bad: unbounded
Why: That contrast is what justifies the abstraction: a constant factor is a fixed multiple at every size, and an unbounded difference means the ratio itself grows with the problem. It is why the order of growth is worth improving before the coefficient.
Explain it
The analysis discards a lot.
Discussion prompt
A classmate objects that ignoring coefficients and hardware makes the analysis useless. Give them the book's own answer.
Hint: The book agrees with them, partly.
Answer:
They have a point, and the book says so: programmers who care about performance often find this kind of analysis hard to swallow, because sometimes the coefficients and the details really do make a difference, and for small problems order of growth is irrelevant.
What it buys is that at least for large problems, the better algorithm is usually better, and sometimes much better — and it gives a simple classification that transfers between machines and languages, which a timing does not.
The decisive line is the comparison: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded. Constants are worth a lot; classes are worth more.
Comparison
Fill the blanks. Three shapes rather than three speeds.
Comparison matrix
| Question | O(1) | O(n) | O(n squared) |
|---|---|---|---|
| How does cost respond to size? | not at all | proportionally | as the square |
| An example operation | indexing a list or dictionary | summing a list | a loop inside a loop over the same list |
| At a million items | about one step | a million steps | a million million |
| Is the gap to the next class fixed? | no — it grows without limit | no — likewise | no |
The bottom row is why classes matter more than constants: within a class the difference is fixed, and between classes it is unbounded.
Pattern
Five steps, and the third is the one that composes.
Step 5 is not optional. For small problems order of growth is irrelevant, and the location of the crossover depends on details the classification threw away.
Python documentation — Built-in Types Built-in Types
Check
Coefficients do not count.
Check your understanding
What is the order of growth of 1000000n cubed + n squared?
Answer: A
Why: The leading term is the term with the highest exponent, and all functions with that leading term belong to the same class regardless of their coefficients. The million is real and decides where the crossover with another algorithm falls, which is not part of the classification.
Check
One of these is not constant.
Check your understanding
Which of these is linear rather than constant time?
Answer: A
Why: Summing has to look at every element, so it is linear — the built-in is faster than an equivalent loop because it has a smaller leading coefficient, not because it is in a different class. Indexing and len are constant regardless of the size, and appending is constant on average.
Check
The book states its own limits.
Check your understanding
When is order of growth NOT a useful guide?
Answer: A
Why: For small problems, order of growth is irrelevant — the book's own table shows the quadratic algorithm winning at n = 10. And if two algorithms have the same leading order term, it is hard to say which is better; the answer depends on the details.
Real world
A method that works at one scale and not at another.
Discussion prompt
Think of a way of doing something that is fine for a few items and impossible for many. What changes, and at roughly what point?
Hint: Checking each against each.
Answer:
Comparing every item with every other, checking a list by reading it from the start, or handling each case by hand — each is fine for a handful and impossible for thousands.
What changes is not the difficulty of one step but how many steps there are, and the point where it becomes impossible is the crossover: below it the simple method wins, above it nothing can rescue it.
Which is exactly the appendix's subject. The classification exists to name the shape of that growth, and the caveat exists because the crossover is real and its location depends on details the classification discards.
Commit first
Answer, then rate your confidence.
Predict first
Algorithm A takes 100n + 1 steps and Algorithm B takes n squared + n + 1. At n = 10, which is faster?
Correct: B, by almost ten times — at n = 10, Algorithm A looks pretty bad.
Why: This is the book's own table and it is there to make exactly this point. At n = 10 the quadratic algorithm takes 111 steps and the linear one takes 1,001; at n = 100 they are about the same; and beyond that A is much better and the gap widens without limit. The reason is the large coefficient on A's leading term, and the reason it does not save B is that regardless of the coefficients, there will always be some value of n where a times n squared exceeds b times n. That point is the crossover, and its location depends on the details of the algorithms, the inputs and the hardware — so it is usually ignored for purposes of algorithmic analysis. The book's warning is the part worth keeping: that doesn't mean you can forget about it. For small problems, order of growth is irrelevant, and if your problem sizes sit below the crossover, the theoretically worse algorithm is the faster one.
Explain it
Why the coefficients are thrown away.
Discussion prompt
A classmate objects that ignoring a coefficient of a hundred is absurd, since it is a hundredfold difference. Answer them.
Hint: What does the coefficient depend on?
Answer:
A hundredfold difference is real and it is fixed — it stays a hundredfold at every size. The exponent's effect is not fixed: the ratio between n and n squared grows without limit as n does.
And the coefficient depends on the machine, the language and the implementation, so it is not a property of the algorithm. The exponent is, which is what lets a comparison transfer between machines.
The book's own sentence is the summary: the difference between two algorithms with the same order of growth is usually a constant factor, but the difference between a good algorithm and a bad algorithm is unbounded.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The leading term is the whole mechanism, and the argument for discarding coefficients — that the exponent is a property of the algorithm and the coefficient is a property of the machine — is worth being able to make. The seven classes are worth memorising in order, since the ordering is the useful part. The operation costs retrospectively explain several choices the course made much earlier. And the caveats matter because the book states them itself: for small problems this is irrelevant, and the crossover point is real.
Connect it up
One page, from memory.
Draw it
Draw the two run-time curves from the book's table and mark the crossover. Beside them, list the seven orders of growth in increasing order of badness. Underneath, make two columns of Python operations — constant and linear — and add the four exceptions with what each one's cost is actually proportional to. Finally write the sentence about constant factors and unbounded differences.
Recap
Four pages, and a name for something the course kept demonstrating.
| If you remember one thing | It is this |
|---|---|
| From the table | At small n the worse algorithm can win by tenfold. |
| From the leading term | The exponent belongs to the algorithm; the coefficient belongs to the machine. |
| From the loop rule | Each level of nesting adds one to the exponent. |
| From the exceptions | Count what the operation reads, not what it is applied to. |
| From the closing line | Same class: a constant factor. Different class: unbounded. |
The last lesson applies all of this to searching: why the in operator on a list is linear, how a bisection search finds a word among a million in twenty steps, and how a hashtable makes dictionary lookup take about the same time however large the dictionary gets.
Think Python, 2nd edition — Allen B. Downey §B.1-B.2, pp. 201-204 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.