11c Memos, Global Variables, and Debugging Large Data

This lesson explains why the recursive Fibonacci is so slow, fixes it with a memo stored in a dictionary, introduces global variables and the global statement, distinguishes modifying a global from reassigning one, and gives three techniques for debugging large datasets.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 11c Memos, Global Variables, and Debugging Large Data

Title

Python · Chapter 11 — Dictionaries

§11.6-11.8, pp. 109-111

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 109-111 — the pages these objectives are drawn from

3. Before we start: why is it so slow?

Warm-up

You wrote the recursive Fibonacci in chapter 6. Try it on a larger argument.

Discussion prompt

fibonacci(20) returns instantly and fibonacci(35) takes several seconds. The function is five lines long and does nothing but add. Where is all that time going?

Hint: How many times does it compute fibonacci(2)?

Answer:

Every call makes two more, so the number of calls roughly doubles for each step up in n. fibonacci(35) makes many millions of calls.

And almost all of them are recomputations. fibonacci(2) is computed from scratch again and again, in different branches, with no memory that the answer was already found.

This lesson stores the answers as they are computed, so each one is worked out once. The change is four lines and the speedup is enormous.

4. The one idea behind this lesson: remember what you worked out

Concept

One solution to the slow Fibonacci is to keep track of values that have already been computed by storing them in a dictionary.

memo — A previously computed value stored for later use.

The dictionary is doing something new here. It is not the program's data — it is machinery, a place to keep results so they need not be found twice. That use of a dictionary is one of the reasons the book calls them the building blocks of many efficient algorithms.

Figure (svg): Two columns contrasting recomputing values with looking them up in a memo

Same algorithm, same arithmetic. The only change is remembering.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 109-109

5. The call graph, and where the time goes

Section

Section 1

6. A picture of every call

Concept

A call graph shows a set of function frames, with lines connecting each frame to the frames of the functions it calls. At the top of the graph, fibonacci with n=4 calls fibonacci with n=3 and n=2. In turn, fibonacci with n=3 calls fibonacci with n=2 and n=1.

def fibonacci(n):
    if n == 0:
        return 0
    elif n == 1:
        return 1
    else:
        return fibonacci(n-1) + fibonacci(n-2)
CallWhat it callsNote
fibonacci(4)calls 3 and 22 calls
fibonacci(3)calls 2 and 1and 2 is computed again
fibonacci(2)called from two placeseach time from scratch
fibonacci(1)called three times in this graphfor n = 4 alone

Count how many times fibonacci(0) and fibonacci(1) are called. This is an inefficient solution to the problem, and it gets worse as the argument gets bigger.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 109-109

7. Picture it: figure 11.2, the call graph for n = 4

Picture it

Nine frames for an argument of four, and only five distinct values among them.

Figure (svg): A call graph for fibonacci with n equal to 4 showing repeated subcalls

The book's figure 11.2. Two whole subtrees compute the same thing.

The right-hand fibonacci(2) is an exact copy of a computation already done on the left. Nothing in the function remembers, so it is done again in full.

8. Worked example: counting the repeated work

Worked example

Count the frames for a small argument and the pattern is obvious.

# frames in the call graph for n = 4
# fibonacci(4)  1 time
# fibonacci(3)  1 time
# fibonacci(2)  2 times
# fibonacci(1)  3 times
# fibonacci(0)  2 times
#               9 calls, 5 distinct values
ArgumentCalls madeValues that exist
n = 49 calls5 distinct
n = 515 calls6 distinct
n = 10177 calls11 distinct
n = 30over 2.7 million calls31 distinct

Count the frames in the graph.

Why: Nine calls to compute a number that depends on only five distinct sub-answers.

Notice which values repeat.

Why: fibonacci(1) appears three times and fibonacci(2) twice, in different branches that know nothing about each other.

See how the gap grows.

Why: The number of distinct values grows by one per step, and the number of calls roughly doubles — so the waste grows without limit.

Figure (svg): A growth chart contrasting the doubling call count with the linear number of distinct values

The gap between the two lines is the wasted work, and it widens at every step.

Nine calls for five values at n = 4, and millions of calls for thirty values at n = 30. The number of things worth computing is tiny; the number of times they are computed is not.

Verify: Check the shape of the growth.

Why: Each step up in n roughly doubles the call count while adding one distinct value, so the ratio between work done and work needed doubles too. That is why the run time increases so quickly, and it is why no amount of faster arithmetic would save this version.

9. Predict: how many times is fibonacci(1) called?

Prediction

Look at the call graph for n = 4.

# the graph for n = 4:
#   4 -> 3, 2
#   3 -> 2, 1
#   2 -> 1, 0     (twice: once under 3, once under 4)
WhereCalls to fibonacci(1)Running total
under fibonacci(3)directly, once1 call
under the first fibonacci(2)once2 calls
under the second fibonacci(2)once3 calls

Predict first

How many times is fibonacci(1) called when computing fibonacci(4)?

  • once
  • twice
  • three times
  • four times

Correct: three times — once directly under fibonacci(3), and once under each of the two fibonacci(2) calls.

Why: The book asks you to count exactly this, because the repetition is the whole diagnosis. Three calls to compute one value that never changes, for an argument as small as four — and the count roughly doubles with every step up in n. The memo turns all three into one computation and two lookups.

10. Worked example: why recursion is not the problem

Worked example

The recursion is fine. The forgetting is the problem.

# the same recursion, computing each value once
known = {0: 0, 1: 1}

def fibonacci(n):
    if n in known:
        return known[n]
    res = fibonacci(n-1) + fibonacci(n-2)
    known[n] = res
    return res
AspectWhat is trueNote
the recursionunchangedstill calls itself twice
what changedresults are rememberedand looked up
the effecteach value computed oncethe graph collapses

Compare the two functions.

Why: The recursive structure is identical — the same two calls, the same addition, the same base cases.

Identify the only change.

Why: Results are stored as they are found and checked for before any work is done.

See what that does to the graph.

Why: The second fibonacci(2) is no longer a subtree; it is a lookup, so the whole branch below it disappears.

Figure (svg): The state of the program after each line of Worked example why recursion is not the problem, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The recursion was never the problem. It was that each call started from nothing, and the fix is to give it something to start from.

Verify: Ask whether the answers change.

Why: They do not — the function computes the same Fibonacci numbers by the same definition. That is what makes this a pure optimisation: the observable behaviour is identical and only the time taken differs, which is the safest kind of change you can make to a program.

11. Trap: assuming a slow program has a slow line in it

Trap

The trap

A student times each line of fibonacci looking for the expensive one, and finds nothing but an addition.

Look for the slow operation

Why: Which works when a program is slow because it does something expensive.

Every individual operation here is instant. The program is slow because of how many times the fast operations run, which no amount of staring at a single line will reveal.

The fix

Count the calls, not the cost of each one.

Draw or count the call graph

Why: Which is exactly what the book does, and the repetition is visible immediately.

Ask how the count grows with the input

Why: Doubling per step is the signature of this problem.

This is the same distinction as the previous lesson's list search: nothing about one comparison is slow, and doing it fifty thousand times is. Performance problems are usually about counts rather than costs.

12. Watch the graph collapse

Invariant

The same computation, with and without a memo.

Step through it

What is preserved across all four frames, and what shrinks?

  1. The left branch expands all the way to the base cases, computing 2, 1 and 0 along the way.
  2. The right branch expands identically, repeating every one of those computations from scratch.
  3. With the memo, the left branch is the same — the values have to be computed a first time.
  4. The right branch is now a single lookup, and the entire subtree beneath it never exists.

The answers are identical in every frame — the memo changes nothing about what is computed, only how often. What shrinks is the number of frames, from nine to five, and the saving compounds at every level as n grows.

13. Think it through: which recursive functions benefit?

Socratic

Not every recursive function is helped by a memo.

Discussion prompt

The factorial function is also recursive and it is not slow. What is different about it, and what does that tell you about when a memo helps?

Hint: Draw its call graph.

Answer:

factorial(n) makes exactly one recursive call, so its call graph is a straight line of n frames with no branching and no repetition.

fibonacci makes two calls, and the two subtrees overlap heavily — the same arguments appear in both. That overlap is what a memo eliminates.

So a memo helps when the same sub-problem is solved more than once. If every call has a distinct argument, there is nothing to remember and the dictionary is pure overhead. The name for the property is overlapping sub-problems, and recognising it is what tells you a memo is worth writing.

14. Where repeated work hides

Real world

Recomputing something you already worked out is not only a programming failure.

Discussion prompt

Think of a process that repeatedly works out something it already knew. What would the equivalent of a memo be, and why do people not always keep one?

Hint: Anything answered by looking it up the same way each time.

Answer:

Looking up the same fact repeatedly, recalculating a total after each small change, re-deriving a route you take every week — each one is work done afresh that could have been recorded.

The memo is a written note, a saved total, a bookmark. Cheap to keep and it removes the work entirely on every repetition after the first.

People do not always keep one for the same reason programs do not: the note has to be kept correct. A memo that is stale is worse than no memo, which is why memoization is safe here — Fibonacci numbers never change — and needs care whenever the underlying answer can.

15. The memoized fibonacci

Section

Section 2

16. Check first, compute once, remember

Concept

Here is a memoized version of fibonacci. known is a dictionary that keeps track of the Fibonacci numbers we already know, and it starts with two items: 0 maps to 0 and 1 maps to 1.

known = {0: 0, 1: 1}

def fibonacci(n):
    if n in known:
        return known[n]
    res = fibonacci(n-1) + fibonacci(n-2)
    known[n] = res
    return res
LineWhat it doesNote
if n in knownthe checkreturn immediately if so
res = ...the recursive computationonly when not known
known[n] = resremember itso no one computes it again
return reshand it backas before

Whenever fibonacci is called, it checks known. If the result is already there, it can return immediately. Otherwise it has to compute the new value, add it to the dictionary, and return it.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 109-110

17. Picture it: the three-step shape of a memoized function

Picture it

Every memoized function has this shape, whatever it computes.

Figure (svg): A flowchart showing the check, compute, and store steps of a memoized function

Check, compute, store. The store is what makes the check worth doing next time.

Leave out the store and the check never succeeds; leave out the check and the store is never consulted. Both halves are needed and neither is useful alone.

18. Worked example: tracing the memoized version

Worked example

Follow fibonacci(4) and watch known fill up.

>>> known = {0: 0, 1: 1}
>>> fibonacci(4)
3
>>> known
{0: 0, 1: 1, 2: 1, 3: 2, 4: 3}
CallWhat happensThe memo after
fibonacci(4)not known, so it recursesknown unchanged so far
fibonacci(2)computed from 1 and 0known gains 2: 1
fibonacci(3)computed from 2 and 1known gains 3: 2
fibonacci(2) againfound in knownreturns at once, no recursion
fibonacci(4)computed from 3 and 2known gains 4: 3

Start with the base cases already stored.

Why: known begins with 0 mapping to 0 and 1 mapping to 1, so the recursion has somewhere to stop without a conditional for it.

Follow the first descent.

Why: Each new value is computed once and immediately stored, so the dictionary grows as the recursion unwinds.

Find the second fibonacci(2).

Why: It is in known, so it returns immediately and its entire subtree — three more calls — never happens.

Figure (svg): The state of the program after each line of Worked example tracing the memoized version, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

3, with known holding all five values. Five computations instead of nine calls, and the saving grows with n.

Verify: Call fibonacci(4) a second time.

Why: It returns immediately from the memo without any recursion at all, because known persists between calls. That persistence is the subject of the next idea, and it is the reason known was defined outside the function rather than inside it.

19. Predict: what does known hold afterwards?

Prediction

The memo starts with two items.

known = {0: 0, 1: 1}
fibonacci(3)
print(len(known))
StageWhat is addedSize
start0 and 12 items
fibonacci(2)computed and stored3 items
fibonacci(3)computed and stored4 items

Predict first

What does this print?

  • 2
  • 3
  • 4
  • 5

Correct: 4 — the two base cases plus the newly computed values for 2 and 3.

Why: Every value computed on the way to the answer is stored, so known ends up holding one item per Fibonacci number from 0 to n. That is why printing len(known) is a good check: if it comes back as 2 after a large call, the store line is missing and the memo is doing nothing.

20. Worked example: what goes wrong if you forget the store

Worked example

A check with nothing to find is a check that always fails.

known = {0: 0, 1: 1}

def fibonacci(n):
    if n in known:
        return known[n]
    res = fibonacci(n-1) + fibonacci(n-2)
    return res            # WRONG: nothing is remembered
PartWhat happensConsequence
the checkruns every calland only ever matches 0 and 1
the storemissingknown never grows
the effectthe original slow versionplus a wasted lookup

Notice the function is still correct.

Why: It returns the right Fibonacci numbers, so no test of its answers would catch this.

Notice known never grows.

Why: It keeps the two base cases it started with and gains nothing, so the check succeeds only for n of 0 or 1.

Conclude what you have.

Why: The original exponential version with an extra dictionary lookup per call — very slightly slower than not trying at all.

Figure (svg): A panel contrasting a memo that grows with one that never does

A correct function with no speedup whatsoever. The bug is invisible to any test of the results, and shows only as time.

Verify: Check known after a call.

Why: It still holds exactly two items, which is the diagnostic. Printing the memo's length after a run is the fastest way to confirm a memo is actually being filled — a summary check of the kind idea 5 recommends.

21. Trap: creating the memo inside the function

Trap

The trap

A student writes known = {0: 0, 1: 1} as the first line of fibonacci, to keep it tidy.

Keep a function's data inside the function

Why: Which is normally the right instinct, and avoids a global.

Now every call starts with a fresh, empty memo, so nothing is ever remembered across calls — and worse, each recursive call resets it. The function is correct and no faster than before.

The fix

The memo has to outlive the call.

Define known outside the function

Why: Local variables disappear when their function ends; global ones persist from one call to the next.

Recognise the symptom

Why: A memoized function that is not faster is nearly always a memo with the wrong lifetime.

This is the whole reason the next section is about global variables. The memo needs to survive between calls, and that requirement is what pushes it out of the function — which then raises the question of how a function may and may not touch it.

22. Complete it: the two halves of a memo

Faded example

One line checks and one line remembers.

Fill in the blanks

def fibonacci(n):
if n in known:
return known[n]
res = fibonacci(n-1) + fibonacci(n-2)
known[n] = res
return res

Why: Without this line the check at the top never matches anything beyond the two base cases, so the function is exactly as slow as the original with a wasted lookup added. The check and the store are two halves of one mechanism and neither does anything on its own.

23. Rank: what a memoized call does, in order

Ranking

Four steps, in the order they happen for a value not yet known.

Put in order

  1. check whether n is already in the memo
  2. compute the value recursively
  3. store the result under n
  4. return the result

Why: Check, compute, store, return. The check has to come first or the computation happens anyway; the store has to come before the return or it never happens at all. For a value that IS known, the sequence stops after the first step and returns immediately.

24. Explain it: my memo did not make it faster

Explain it

The commonest memoization bug, and it has two forms.

Discussion prompt

A classmate added a memo to a slow recursive function and it runs at the same speed. Give them two things to check and how to tell which one it is.

Hint: Print the size of the memo.

Answer:

First: is the store line there? Without it the memo never grows and the check never matches. Printing len(the memo) after a run tells you immediately — it should be roughly the number of distinct arguments.

Second: is the memo defined inside the function? Then every call gets a fresh one and the size will be small but non-zero on the way down, resetting each time.

Both produce the same symptom — a correct function with no speedup — and the memo's size distinguishes them: unchanged from its initial contents means the store is missing; growing and resetting means the lifetime is wrong.

25. Global variables

Section

Section 3

26. Variables that outlive a call

Concept

In the previous example, known is created outside the function, so it belongs to the special frame called __main__. Variables in __main__ are sometimes called global because they can be accessed from any function.

global variable — A variable defined outside a function, which can be accessed from any function.

verbose = True

def example1():
    if verbose:
        print('Running example1')
PartWhat is trueNote
verbose = Truecreated in __main__a global
inside example1read without any declarationreading is always allowed
after the callverbose still existsglobals persist

Unlike local variables, which disappear when their function ends, global variables persist from one function call to the next. It is common to use them for flags — boolean variables that indicate whether a condition is true, like a verbose flag controlling the level of detail in the output.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 110-110

27. Picture it: __main__ underneath every call

Picture it

The function's frame comes and goes. The one below it does not.

Figure (svg): A stack diagram showing __main__ holding globals beneath a temporary function frame

Globals live in the frame at the bottom, which is there for the whole program.

That is exactly why the memo works: known is in __main__, so what one call stores is still there for the next.

28. Worked example: why the memo needs a global

Worked example

The lifetime is the whole point.

# global: survives between calls
known = {0: 0, 1: 1}
def fib(n):
    if n in known:
        return known[n]
    ...

# local: a fresh one every call
def fib_bad(n):
    known = {0: 0, 1: 1}
    ...
VersionHow many memos existEffect
the global versionone dictionary for the programresults accumulate
the local versiona new dictionary per callresults are discarded
the recursionmakes many callsso many dictionaries

Ask how long the memo must last.

Why: Longer than one call, because the point is for one call to benefit from another's work.

Match that to a lifetime.

Why: Local variables disappear when their function ends; global variables persist from one function call to the next. Only the second lifetime will do.

Note the recursive twist.

Why: In the local version, each recursive call makes its own memo too, so results are discarded even within a single top-level call.

Figure (svg): The state of the program after each line of Worked example why the memo needs a global, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The memo must be global because it has to outlive the call that fills it. This is one of the clearest legitimate uses of a global variable in the book.

Verify: Check the claim about persistence directly.

Why: Call fibonacci(10), then print known — it holds eleven items after the call has ended, which a local variable could not do. That surviving dictionary is the entire mechanism, and it is why the second call to fibonacci(10) does no work at all.

29. Predict: can the function see it?

Prediction

The variable is defined outside and only read inside.

verbose = True

def show():
    if verbose:
        print('yes')

show()
PartWhat happensResult
verbosein __main__a global
inside showread, not assignedno declaration needed
outputyesthe flag is True

Predict first

What does this print?

  • yes
  • nothing
  • A NameError, because verbose is not defined in show
  • True

Correct: yes — a global variable can be read from any function without any declaration.

Why: Variables in __main__ can be accessed from any function, and reading one requires nothing special. The NameError in option C is what would happen if verbose had been defined inside another function instead — locals are not visible outside their own frame. The distinction that matters comes next: reading is free, assigning is not.

30. Worked example: a flag

Worked example

The book's other example, and the commonest use of a global.

verbose = True

def example1():
    if verbose:
        print('Running example1')

def example_two():
    if verbose:
        print('Running example_two')
AspectWith a globalNote
verboseone variableread by both functions
setting it oncechanges bothno argument passing
the alternativea parameter on every functionthreaded through everything

Note what a flag is.

Why: A boolean variable that indicates — flags — whether a condition is true. Here it controls the level of detail in the output.

Note why a global suits it.

Why: Every function wants to consult it and none wants to be passed it. Threading a verbose parameter through every call would clutter every signature.

Note the limit.

Why: Reading a global is unremarkable. Changing one from inside a function is where the difficulties start, which is the next idea.

Figure (svg): Two columns contrasting reading a global with passing a parameter to every function

The convenience and the cost are the same property: the dependency is invisible.

A single flag consulted by many functions, with no parameter passing. Reading a global needs no declaration of any kind.

Verify: Check whether anything special was needed to read it.

Why: Nothing at all — example1 uses verbose exactly as if it were local. That asymmetry is worth noticing now, because the next idea is entirely about how different the situation becomes when a function tries to assign to a global.

31. Trap: reaching for a global to avoid passing a value

Trap

The trap

Two functions need the same list, so it is made global rather than passed between them.

Avoid the plumbing

Why: Passing a value through three functions to reach the fourth is genuinely tedious.

Now nothing in either signature says the functions depend on that list, and anything in the program can change it. Global variables can be useful, but if you have a lot of them and you modify them frequently, they can make programs hard to debug.

The fix

Pass values as arguments unless the lifetime genuinely requires a global.

Ask what the variable is for

Why: A memo that must outlive calls, or a flag every function consults, is a fair case. Data being moved from one function to another is not.

Keep the number small

Why: The book's warning is about quantity and frequency of modification, not about globals existing.

The test is whether removing the global would only mean adding a parameter. If so, add the parameter — the signature then says what the function needs, which is information a reader otherwise has to hunt for.

32. Compare: local and global variables

Comparison

Fill the blanks. The lifetime is the difference that matters here.

Comparison matrix

QuestionLocalGlobal
Where is it created?inside a functionoutside every function, in __main__
How long does it last?until the function endsfrom one call to the next
Who can read it?only that functionany function
What does a memo need?no — it would reset every callyes — results must survive the call

The bottom row is why §11.6 leads straight into §11.7. Memoization forces the question of scope.

33. Discriminate: global or parameter?

Discrimination

Ask whether the value must outlive the call.

Sort into buckets

For each, is a global variable the reasonable choice?

a global is reasonable
a memo of results computed so far; a verbose flag consulted by every function; a count of how many times a program has retried
should be a parameter
the list a function is being asked to sort; the string a function should search; the two numbers a function should add
glob
Each has to outlive a single call or be consulted from many places: a memo accumulates across calls, a flag is program-wide, and a retry count is meaningless if it resets. The lifetime is the justification.
param
Each is data the function is being asked to work on, different on every call. Making it global would hide the dependency and mean only one call could be in flight at a time.

34. Explain it yourself: why does the memo have to be global?

Explain it to yourself

Every other variable in the function is local and that is fine.

Discussion prompt

Explain, in terms of lifetimes, why known must be defined outside fibonacci while n and res must not.

Hint: Which of the three is still useful after the call ends?

Answer:

n and res are about one call. When it ends they have served their purpose, and a fresh call wants fresh ones — being local is exactly right.

known is about every call. Its entire value is that what one call learned is available to the next, and a local variable disappears when its function ends.

So the choice is not stylistic. Each variable's scope should match how long its contents remain meaningful, and for a memo that is the life of the program — which is the definition of a global.

35. Reassigning a global, and the global statement

Section

Section 4

36. Assignment inside a function creates a local

Concept

If you try to reassign a global variable, you might be surprised. The following example is supposed to keep track of whether the function has been called.

been_called = False

def example2():
    been_called = True     # WRONG

# and the fix:
def example2():
    global been_called
    been_called = True
VersionWhat happensNote
the first versioncreates a NEW localthe global is untouched
the localgoes away when the function endsno effect at all
global been_calledtells the interpreter which one you meannow the assignment lands

The global statement tells the interpreter something like: in this function, when I say been_called, I mean the global variable — don't create a local one.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 110-110

37. Picture it: two variables with one name

Picture it

Without the declaration, the assignment creates a second variable in the function's own frame.

Figure (svg): A stack diagram showing a local been_called shadowing the global of the same name

Two boxes, one name. The assignment filled the upper one, which then disappeared.

Nothing raises, and the value of been_called doesn't change — which makes this one of the quietest bugs in the book.

38. Worked example: the noisier failure

Worked example

Trying to update rather than replace produces an error instead of silence.

count = 0

def example3():
    count = count + 1     # WRONG

>>> example3()
UnboundLocalError: local variable 'count' referenced before assignment
PartWhat happensNote
the assignmentmakes count local for the whole functionincluding the right-hand side
the right-hand sidereads the local countwhich has no value yet
the errorreferenced before assignmentand it names the variable

Notice the assignment makes count local.

Why: Python assumes that count is local, because there is an assignment to it somewhere in the function.

Read the right-hand side under that assumption.

Why: Under that assumption you are reading it before writing it, which is not allowed.

Add the declaration.

Why: global count, then count += 1, which now reads and writes the variable in __main__.

Figure (svg): A panel comparing the silent reassignment failure with the noisy UnboundLocalError

An UnboundLocalError naming the variable. Unlike the been_called case this one is loud, because the wrong reading requires reading a local that has never been assigned.

Verify: Compare the two failures side by side.

Why: been_called = True is silent because nothing is read before it is written; count = count + 1 raises because it is. The same misunderstanding produces a silent wrong answer in one case and an error in the other, which is a useful reminder that loudness is not proportional to severity.

39. Predict: what does been_called hold?

Prediction

No global declaration, and no error either.

been_called = False

def example2():
    been_called = True

example2()
print(been_called)
StageWhat happensResult
the assignmentcreates a localthe global is untouched
the function endsthe local disappearsnothing was recorded
printthe global, still FalseFalse

Predict first

What does this print?

  • False
  • True
  • An UnboundLocalError is raised
  • None

Correct: False — the assignment created a new local variable, which went away when the function ended.

Why: The local variable has no effect on the global one, and nothing raises because nothing was read before being written. Adding global been_called as the first line of the function makes the assignment land on the global and the answer becomes True. Compare with count = count + 1, which raises UnboundLocalError precisely because it reads first.

40. Worked example: modifying without declaring

Worked example

The exception to the rule, and the reason the memo needed no global statement.

known = {0: 0, 1: 1}

def example4():
    known[2] = 1          # fine: modifies the dictionary

def example5():
    global known
    known = dict()        # reassignment: needs the declaration
StatementWhat it doesDeclaration needed?
known[2] = 1modifies the objectno declaration needed
known = dict()reassigns the nameneeds global
the rulemodify freely, reassign with a declarationand only names are scoped

Notice what the first one touches.

Why: If a global variable refers to a mutable value, you can modify the value without declaring the variable.

Notice what the second one touches.

Why: You can add, remove and replace elements of a global list or dictionary, but if you want to reassign the variable, you have to declare it.

Connect it to chapter 10.

Why: This is the modify-versus-reassign distinction again. Scope is about names, and modifying an object does not touch any name.

Figure (svg): The state of the program after each line of Worked example modifying without declaring, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Modification needs no declaration and reassignment does. The memo only ever does known[n] = res, which is a modification — which is why fibonacci needs no global statement at all.

Verify: Check the rule against the memoized fibonacci.

Why: It contains known[n] = res and never known = anything, so it modifies and never reassigns — and indeed the book's version has no global statement. Noticing that the two sections fit together is the point: §11.7 is explaining something §11.6 quietly relied on.

41. Trap: adding global out of caution

Trap

The trap

A student declares global for every global name a function touches, including ones it only reads or modifies.

Declare what you use

Why: It looks like documentation and it removes the risk of the silent shadowing bug.

It is unnecessary for reads and for modifications, and it converts a function that merely used a global into one that announces it may replace it — which is a stronger and more alarming claim than the code actually makes.

The fix

Declare only when you assign to the name itself.

Reading needs nothing

Why: example1's use of verbose is complete as written.

Modifying an object needs nothing

Why: known[n] = res, t.append(x), d['k'] = v — none of these touch a name.

Only name = something needs the declaration. Keeping global to that one case means its presence tells a reader something specific: this function replaces that variable.

42. Sort: does this need a global declaration?

Sorting

Ask whether the statement assigns to the name itself.

Sort into buckets

For each statement inside a function, is a global declaration needed?

needs global
known = dict(); count = count + 1; been_called = True
no declaration needed
known[n] = res; if verbose: print(...); items.append(x)
need
Each assigns to the name itself, which without a declaration creates a local instead. Two of them fail silently and one raises UnboundLocalError, because it reads the name before writing it.
no
Each either reads the variable or modifies the object it refers to. Neither touches the name, and scope is entirely about names — which is why the memoized fibonacci needs no global statement.

43. Complete it: update a global counter

Faded example

This one reads before it writes, so the declaration is not optional.

Fill in the blanks

count = 0

def example3():
global count
count += 1

Why: Without it, Python assumes count is local because the function assigns to it, and reading it on the right-hand side of += then raises UnboundLocalError: local variable 'count' referenced before assignment. The declaration tells the interpreter that count means the global one, so both the read and the write land there.

44. Two truths and a lie: globals

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. You can modify a global list or dictionary without declaring the variable
  • B. Assigning to a global name inside a function creates a local unless you declare it
  • C. You need a global declaration to read a global variable

Survives elimination: C

Why: C is false: reading needs nothing at all, which is why example1's use of verbose works exactly as written. The declaration is needed only when you assign to the name itself. Adding it for reads is harmless but misleading, because it announces that the function may replace the variable when it does not.

45. Debugging large datasets

Section

Section 5

46. Three techniques for data too big to read

Concept

As you work with bigger datasets it can become unwieldy to debug by printing and checking the output by hand. The book gives three suggestions.

If there is an error, you can reduce n to the smallest value that manifests the error, and then increase it gradually as you find and correct errors.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 111-111

47. Picture it: three techniques, three kinds of question

Picture it

Each answers a different question about a dataset you cannot read.

Figure (svg): Two columns pairing each debugging technique with the question it answers

Three questions you cannot answer by printing everything and looking.

The third is the only one that keeps working when nobody is watching, which is why it is the one worth building into the program.

48. Worked example: a sanity check

Worked example

The average of a list has bounds you know without computing it.

def average(t):
    result = sum(t) / len(t)
    assert result <= max(t)
    assert result >= min(t)
    return result
LineWhat it assertsNote
the computationsum divided by countthe answer
the first checkcannot exceed the largestor something is wrong
the second checkcannot be below the smallestlikewise

Identify something you know without computing.

Why: If you are computing the average of a list of numbers, you could check that the result is not greater than the largest element or less than the smallest.

Write it as a check the program runs.

Why: This is called a sanity check because it detects results that are insane — impossible rather than merely surprising.

Note what it catches.

Why: A wrong divisor, a sum over the wrong list, a stray element — all of which produce an answer outside the bounds while looking perfectly plausible on their own.

Figure (svg): A flowchart showing a computation followed by a sanity check that either passes or fails

The check runs every time, including on the run nobody is watching.

A function that refuses to return an impossible answer. The check costs two comparisons and catches a whole family of bugs that no amount of looking would find.

Verify: Ask what the check does not catch.

Why: An average that is wrong but still between the smallest and largest values — which is most wrong averages. A sanity check narrows the space of undetected bugs rather than eliminating it, and that is worth having precisely because it costs almost nothing.

49. Predict: what does this check catch?

Prediction

A consistency check on a histogram.

h = histogram(s)
total = 0
for c in h:
    total += h[c]
assert total == len(s)
QuantityWhat it measuresNote
the sumall the counts addedcharacters accounted for
len(s)characters in the stringcomputed independently
the assertionthey must agreeor something was lost

Predict first

Which bug would this check catch?

  • A histogram that initialises new counters to 0 instead of 1
  • A histogram whose keys are in the wrong order
  • A histogram that uses get instead of a conditional
  • A string containing characters the program has not seen before

Correct: A histogram that initialises new counters to 0 instead of 1 — the total would fall short by the number of distinct characters.

Why: The off-by-one makes every count one too small, so the sum is less than the length of the string and the assertion fails. The second option is not a bug at all, since the order of items in a dictionary is unpredictable; the third is a legitimate alternative implementation; and the fourth is exactly the case a dictionary handles without any advance knowledge.

50. Worked example: a consistency check

Worked example

Two ways to compute the same thing should agree.

def check_histogram(s, h):
    total = 0
    for c in h:
        total += h[c]
    assert total == len(s)
PartWhat it computesNote
the histogramcounts per characterone item per distinct character
summing the countsthe total characters counteda second computation
comparingmust equal len(s)or a character was lost

Find a second route to a known quantity.

Why: The counts in a histogram must add up to the length of the string, which you can compute independently.

Compare the two.

Why: Another kind of check compares the results of two different computations to see if they are consistent.

Note what it catches.

Why: A character skipped, double-counted, or a counter initialised to zero instead of one — the exact off-by-one from lesson 11a, caught automatically.

Figure (svg): The state of the program after each line of Worked example a consistency check, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A check that the histogram accounts for every character. It uses the data itself rather than a hand-computed expected answer, so it works on any input.

Verify: Run it against the broken histogram that initialises to zero.

Why: The total comes out short by exactly the number of distinct characters, and the check fails on the first input tried. That is the value of a consistency check: it turns a bug that produces plausible numbers into one that announces itself.

51. Trap: debugging on the full dataset

Trap

The trap

A program fails somewhere in a hundred thousand lines of input, so a print is added inside the loop and the program is run again.

Look at what the program is doing

Why: Printing is the standard technique, and it worked on every small program so far.

It produces a hundred thousand lines of output, in which the interesting one is invisible. The technique has not stopped working — the volume has made it useless.

The fix

Scale the input down first.

Modify the program to read only the first n lines

Why: Better than editing the files, because it is one place to change and it leaves the data intact.

Reduce n until the error just still appears

Why: Then you have the smallest case that manifests it, and printing becomes useful again.

Then increase n gradually as you find and correct errors. The point is not that printing is wrong; it is that printing needs a small enough case to be readable, and producing one is a step you can take deliberately.

52. Sort: which technique fits this problem?

Sorting

Three techniques, three kinds of difficulty.

Sort into buckets

For each situation, which of the book's suggestions applies most directly?

scale down the input
the program crashes somewhere in a 50,000-line file; the output is thousands of lines and you cannot see the wrong one
check summaries and types
a TypeError says an operation is not supported for these operands; you want to know whether the dictionary got all the records
write a self-check
an average comes out larger than every value in the list; two parts of the program compute the same total differently
scale
Both are volume problems: the program's behaviour is fine to inspect, there is just far too much of it. Reducing n until the error just still appears makes every other technique usable again.
summ
One is a type question, for which printing the type of a value is often enough; the other is a size question, answered by the number of items rather than by their contents.
self
Both compare a result against something independently known — bounds in one case, a second computation in the other. These are the sanity and consistency checks, and they keep working unattended.

53. Complete it: a sanity check on an average

Faded example

An average cannot be outside the range of the values.

Fill in the blanks

def average(t):
result = sum(t) / len(t)
assert result <= max(t)
return result

Why: The average of a list of numbers cannot exceed its largest element, so a result that does is impossible rather than merely surprising. The book calls this a sanity check because it detects results that are insane, and it catches a wrong divisor or a sum over the wrong data without your having to know the right answer in advance.

54. Where sanity checks are already standard

Real world

The idea is not specific to programming.

Discussion prompt

Where outside programming does someone routinely check that an answer is even possible, before checking whether it is right?

Hint: Anything where a wrong answer is expensive.

Answer:

A pharmacist checking that a dose is within a plausible range; an accountant seeing whether a total exceeds the sum of its parts; anyone noticing that a journey time came out negative.

None of these confirms the answer is correct. They rule out a class of answers that cannot be, which is much cheaper than verifying and catches the worst errors.

And the programming version has one advantage: it can be written into the program, so it runs on every input forever rather than only when someone remembers to look. That is what makes a self-check different from being careful.

55. Compare: modifying and reassigning a global

Comparison

Fill the blanks. It is chapter 10's distinction, at the level of scope.

Comparison matrix

QuestionModifying: known[n] = resReassigning: known = dict()
What does it change?the objectwhich object the name refers to
Declaration needed?noyes — global known
Without one, what happens?it works as intendeda new local is created and the global is untouched
Which does fibonacci do?this one — so it needs no declarationnever

Scope is about names. An operation that does not touch a name does not raise a scope question at all.

56. The procedure: memoizing a slow recursive function

Pattern

Five steps. The fourth is the one people leave out.

  1. Confirm the function repeats work — the same argument appearing in more than one branch of the call graph.
  2. Create a dictionary outside the function, so it survives from one call to the next.
  3. Put the known base cases in it, which removes the need for their conditionals.
  4. At the top of the function, return the stored value if the argument is already a key.
  5. Before returning a newly computed value, store it under its argument.

Step 5 is the one that gets forgotten, and its symptom is a correct function that is no faster. Printing the size of the memo after a run is the check that finds it.

Python documentation — More Control Flow Tools More Control Flow Tools

57. Check yourself 1 of 3: the memo

Check

One line has been removed from the memoized version.

known = {0: 0, 1: 1}

def fibonacci(n):
    if n in known:
        return known[n]
    res = fibonacci(n-1) + fibonacci(n-2)
    return res
PartWhat happensEffect
the checkruns every callmatches only 0 and 1
the storemissingknown never grows
the resultcorrectand just as slow

Check your understanding

What is wrong with this version?

  • A. It returns the wrong numbers
  • B. It is correct but no faster, because nothing is ever stored (correct)
  • C. It raises a KeyError on the first call
  • D. It needs a global declaration for known

Answer: B

Why: Without known[n] = res, the memo keeps only the two base cases, so the check at the top almost never matches and every value is recomputed exactly as in the original. The bug is invisible to any test of the answers and shows only as time — or as len(known) coming back as 2 after a large call.

Why A tempts people
The arithmetic is untouched, so the answers are right. That is precisely why the bug is hard to see.
Why C tempts people
The check uses in rather than a lookup, so a missing key produces False rather than an exception.
Why D tempts people
A declaration would be needed only to reassign known. This version does not assign to the name at all — which is the deeper problem.

58. Check yourself 2 of 3: the global statement

Check

One of these needs a declaration.

Check your understanding

Which statement inside a function requires a global declaration?

  • A. known[n] = res
  • B. known = dict() (correct)
  • C. if verbose: print('hello')
  • D. print(len(known))

Answer: B

Why: Only assignment to the name itself needs the declaration. Without it, known = dict() creates a new local and the global dictionary is untouched — a silent failure. The other three either read the variable or modify the object it refers to, and neither of those touches a name, which is what scope is about.

Why A tempts people
This modifies the dictionary. You can add, remove and replace elements of a global list or dictionary without declaring the variable.
Why C tempts people
Reading a global needs nothing at all, which is why the verbose flag works exactly as written.
Why D tempts people
Also a read, passed to a function. Nothing about it assigns to the name known.

59. Check yourself 3 of 3: self-checks

Check

Two kinds of automatic check.

Check your understanding

What is the difference between a sanity check and a consistency check?

  • A. A sanity check runs before the computation and a consistency check runs after
  • B. A sanity check detects an impossible result; a consistency check compares two computations of the same thing (correct)
  • C. A sanity check is written by hand and a consistency check is automatic
  • D. They are two names for the same technique

Answer: B

Why: A sanity check detects results that are insane — an average larger than every value, a negative count — using bounds you know without computing anything. A consistency check compares the results of two different computations to see if they are consistent, like summing a histogram's counts and comparing with the length of the string.

Why A tempts people
Both run after a result exists, since both examine one. Timing is not the distinction.
Why C tempts people
Both are code you write and the program runs. Being automatic is what makes either of them worth more than remembering to look.
Why D tempts people
The book names them separately because they use different evidence: known bounds in one case, a second computation in the other.

60. Where this shows up outside this course

Real world

Storing an answer so it need not be found again is one of the most general ideas in computing.

Discussion prompt

Think of something you use that is fast the second time and slow the first — a page that loads instantly on a revisit, a search that remembers. What is being stored, and what could go wrong with storing it?

Hint: What happens if the underlying thing changes?

Answer:

Browser caches, search suggestions, saved calculations — each stores an answer that was expensive to produce so the next request can be answered without producing it again. That is exactly a memo.

What can go wrong is that the stored answer becomes wrong. A cached page that has since changed is worse than no cache, because the program is confidently returning something stale.

Fibonacci is the ideal case: the answer for a given n never changes, so a memo can never be stale. Every real caching problem is about deciding when a stored answer stops being trustworthy, and that is a genuinely hard question — which is why this example is the one the book chooses to introduce the idea.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence. This one catches almost everyone once.

Predict first

A global variable holds a dictionary. A function does d['k'] = 1. Does it need a global d declaration?

  • No — it modifies the dictionary rather than reassigning the name
  • Yes — any use of a global inside a function needs the declaration
  • Yes — because it is an assignment statement
  • Only if the key is new

Correct: No — it modifies the dictionary rather than reassigning the name, and scope is about names.

Why: The book states this directly: if a global variable refers to a mutable value, you can modify the value without declaring the variable. You can add, remove and replace elements of a global list or dictionary; only reassigning the variable — d = dict() — requires the declaration. It looks like an assignment because of the equals sign, but the target is an item inside the object, not the name d. This is chapter 10's modify-versus-reassign distinction showing up at the level of scope, and it is exactly why the memoized fibonacci works without any global statement.

62. Explain it to someone else

Explain it

Four lines turned an unusable function into an instant one.

Discussion prompt

A classmate cannot see why the memoized fibonacci is so much faster, since it does the same additions. Explain it using the call graph.

Hint: Ask them to count the calls, not the additions.

Answer:

Draw the call graph for n = 4 and have them count: nine calls for five distinct values, with fibonacci(1) computed three separate times.

Then cross out the second fibonacci(2) and everything under it, because with the memo it is one lookup. Two thirds of the graph disappears for an argument as small as four.

The additions are the same, and there are far fewer of them, because each value is computed once instead of once per branch that needs it. The speedup is not in doing the work faster — it is in not doing it repeatedly.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • The call graph, and why the recursive Fibonacci is so slow
  • Writing a memo: the check, the store, and where the dictionary lives
  • Global variables, and when a global declaration is needed
  • Debugging large data: scaling down, summaries, and self-checks

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The call graph is worth drawing by hand once, because seeing the repetition is more convincing than being told about it. The memo is a three-part shape you will reuse, and the store is the part people leave out. The global rules are where nearly everyone is caught once — the silent been_called failure especially — and the modify-versus-reassign distinction is what makes them predictable rather than arbitrary. The debugging techniques matter most later, when a program stops fitting on a screen.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw the call graph for fibonacci(4) and circle every frame that a memo would replace with a lookup. Beside it, write the memoized function and label its three parts: the check, the computation, and the store. Underneath, draw two stack frames — __main__ and a function — and write next to them the three things a function can do with a global name, marking which one needs a declaration and what happens if you leave it out.

65. What you can do now

Recap

Three pages, and chapter 11 is finished: a dictionary used as machinery.

If you remember one thingIt is this
From the call graphA program can be slow because of how often fast things run.
From memosCheck, compute, store — and the store is the half that gets forgotten.
From globalsLocal variables disappear; global ones persist. A memo needs the second.
From the global statementModify freely; declare only to reassign the name.
From debuggingA self-check runs on every input, including the ones nobody looks at.

The next chapter introduces tuples: immutable sequences, which can be dictionary keys precisely because they cannot change, and which let a function return more than one value.

Think Python, 2nd edition — Allen B. Downey §11.6-11.8, pp. 109-111 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §11.6-11.8, pp. 109-111
  2. Python documentation — More Control Flow Tools
  3. Python documentation — Built-in Functions

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108