14c Databases, Pickling, Pipes, and Writing Modules

This lesson covers dbm databases that behave like dictionaries on disk, the pickle module for storing arbitrary objects, pipes for running other programs, writing importable modules with the __name__ idiom, and repr for debugging invisible whitespace.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 14c Databases, Pickling, Pipes, and Writing Modules

Title

Python · Chapter 14 — Files

§14.6-14.10, pp. 141-144

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-144 — the pages these objectives are drawn from

3. Before we start: storing a dictionary in a file

Warm-up

A text file holds characters. A dictionary does not.

Discussion prompt

You want to save {'a': 1, 'b': [2, 3]} so a later run can load it back. Writing it with str gives you a string that looks right. What is the difficulty in reading it back?

Hint: What would you have to write to turn that string into a dictionary?

Answer:

You would have to parse it: find the braces, split on commas that are not inside brackets, split each pair on its colon, and work out the type of every value.

Which is a small programming language interpreter, and it goes wrong on the first value containing a comma or a colon — the separator problem from lesson 12c, at full scale.

This lesson has two answers: a database that stores key-value pairs directly, and a module that converts any object to a string and back again without your writing any of that.

4. The one idea behind this lesson: storage that keeps the structure

Concept

A database is a file that is organised for storing data. Many databases are organised like a dictionary in the sense that they map from keys to values.

database — A file whose contents are organised like a dictionary with keys that correspond to values.

The biggest difference between a database and a dictionary is that the database is on disk — or other permanent storage — so it persists after the program ends. Everything else about using one is familiar.

Figure (svg): Two columns comparing a dictionary in memory with a database on disk

The interface is nearly the same, which is what makes it easy to use.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-141

5. A database that behaves like a dictionary

Section

Section 1

6. dbm, and mode 'c'

Concept

The module dbm provides an interface for creating and updating database files. Opening a database is similar to opening other files.

>>> import dbm
>>> db = dbm.open('captions', 'c')
>>> db['cleese.png'] = 'Photo of John Cleese.'
>>> db['cleese.png']
b'Photo of John Cleese.'
>>> db.close()
LineWhat happensNote
mode 'c'create if it does not existand open it if it does
assignmentdbm updates the database FILEimmediately
lookupdbm reads the fileand returns bytes

The mode 'c' means that the database should be created if it doesn't already exist — unlike a file's mode 'w', which destroys an existing one. The result is a database object that can be used, for most operations, like a dictionary.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-141

7. Picture it: every assignment reaches the disk

Picture it

There is no separate saving step.

Figure (svg): A pipeline showing an assignment to a database object updating the file immediately

When you create a new item, dbm updates the database file.

That is the whole point: the structure and the storage are the same thing, so nothing has to be written out at the end.

8. Worked example: what comes back is bytes

Worked example

The one visible difference from a dictionary.

>>> db['cleese.png'] = 'Photo of John Cleese.'
>>> db['cleese.png']
b'Photo of John Cleese.'
>>> db['cleese.png'] = 'Photo of John Cleese doing a silly walk.'
>>> db['cleese.png']
b'Photo of John Cleese doing a silly walk.'
OperationWhat happensNote
what you storea stringordinary
what comes backa bytes objectnote the b
a second assignmentreplaces the old valueas in a dictionary

Store a string.

Why: The assignment looks exactly like a dictionary's, and dbm writes it to the file.

Read it back.

Why: The result is a bytes object, which is why it begins with b. A bytes object is similar to a string in many ways.

Replace it.

Why: If you make another assignment to an existing key, dbm replaces the old value — the same behaviour as a dictionary.

Figure (svg): The state of the program after each line of Worked example what comes back is bytes, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Values come back as bytes rather than strings. When you get farther into Python the difference becomes important, and for now it can be ignored.

Verify: Check that the bytes compare as expected.

Why: b'abc' == 'abc' is False, so a comparison against a plain string fails even when the text matches — which is the one place the difference bites early. Decoding with .decode() gives an ordinary string back if you need one.

9. Predict: what comes back?

Prediction

A string went in.

db['a.png'] = 'a caption'
print(db['a.png'])
StepWhat it isNote
storeda stringordinary
retrieveda bytes objectbegins with b
the differencematters laterignorable for now

Predict first

What does this print?

  • b'a caption'
  • 'a caption'
  • a caption
  • A TypeError, since databases store only bytes

Correct: b'a caption' — the result is a bytes object, which is why it begins with b.

Why: A bytes object is similar to a string in many ways, and printing one shows the b prefix. Storing a plain string is fine; it is the retrieval that converts. The practical consequence is that comparing the result against a plain string fails, so .decode() is needed when the value has to be an ordinary string.

10. Worked example: what does not work

Worked example

It is like a dictionary for most operations, not all.

# some dictionary methods do not work:
#   db.keys()  and db.items()  are not available as usual

# but iteration with a for loop does:
for key in db.keys():
    print(key, db[key])

db.close()
OperationAvailable?Note
keys, itemsdo not work as they do for a dictionarythe interface is narrower
iterationworksone key per pass
closerequired, as with other fileswhen you are done

Note the missing methods.

Why: Some dictionary methods, like keys and items, don't work with database objects — the interface resembles a dictionary rather than matching it.

Note that iteration works.

Why: A for loop over the database gives its keys, which covers most of what those methods were for.

Close it when done.

Why: As with other files, you should close the database when you are finished.

Figure (svg): Two columns separating the dictionary operations a database supports from those it does not

Assignment, lookup and iteration cover nearly everything you need.

Most operations transfer and a few do not. The practical approach is to use assignment, lookup and iteration, which are all supported.

Verify: Ask why the interface is narrower.

Why: Because the data are on disk, so an operation that would be a quick pass over memory becomes a pass over a file — and some, like producing a list of every item at once, are exactly what you would want to avoid on a large database. The gaps are mostly places where the convenient dictionary operation would be an expensive database one.

11. Trap: forgetting to close the database

Trap

The trap

A program writes many items to a database and ends without closing it.

Assume the writes have happened

Why: Each assignment updated the file, so it looks finished.

As with other files, the database should be closed when you are done — and until it is, the file may not be complete on disk. A program that writes and then exits abruptly can leave it in a partial state.

The fix

Close it, as you would a file.

db.close() when you have finished

Why: The book says so directly: as with other files, you should close the database.

And treat it as the point the data are safe

Why: The same reasoning as closing a written file.

This is the same habit as lesson 14a's close, for the same reason. Persistence is only useful if the data actually reached the disk, and closing is where that is guaranteed.

12. Compare: a dictionary and a dbm database

Comparison

Fill the blanks. The interface is nearly the same.

Comparison matrix

QuestionDictionaryDatabase
Where does it live?in memoryon disk
Does it survive the program?noyes
What can the keys be?any hashable typestrings or bytes only
Do keys() and items() work?yesnot as they do for a dictionary

The third row is the limitation pickle addresses, and it is the reason the next idea exists.

13. Complete it: open a database

Faded example

Create it if it is not there.

Fill in the blanks

import dbm
db = dbm.open('captions', 'c')

Why: Mode 'c' means the database should be created if it doesn't already exist, and opened if it does — which is what you almost always want. Note that this is not the same as a file's mode 'w', which destroys an existing file; a database opened with 'c' keeps its contents.

14. Explain it yourself: why does a database look like a dictionary?

Explain it to yourself

The interface is a deliberate choice.

Discussion prompt

dbm could have offered store and fetch functions. Why present it as something you index with brackets instead?

Hint: What do you already know how to use?

Answer:

Because you already know how to use a dictionary. Presenting the database the same way means nearly everything you know transfers with no new syntax at all.

And because the underlying idea genuinely is the same: many databases are organised like a dictionary, mapping from keys to values. The interface matches the concept rather than disguising it.

The cost is that the resemblance is not exact, so the gaps surprise you — keys and items do not work as usual, and the values come back as bytes. A familiar interface that is almost the same is easier to learn and slightly harder to trust.

15. Pickling

Section

Section 2

16. Any object, as a string and back

Concept

A limitation of dbm is that the keys and values have to be strings or bytes; if you try to use any other type, you get an error. The pickle module can help: it translates almost any type of object into a string suitable for storage in a database, and then translates strings back into objects.

>>> import pickle
>>> t = [1, 2, 3]
>>> pickle.dumps(t)
b'\x80\x03]q\x00(K\x01K\x02K\x03e.'
>>> t2 = pickle.loads(pickle.dumps(t))
>>> t2
[1, 2, 3]
FunctionWhat it doesNote
dumpsdump stringan object to a string
the formatnot obvious to human readersmeant for pickle
loadsload stringreconstitutes the object

The format isn't obvious to human readers; it is meant to be easy for pickle to interpret. That is the trade against a text file, which is readable by anything and cannot restore a structure.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 142-142

17. Picture it: the round trip

Picture it

Out to a string, into storage, and back to an object.

Figure (svg): A pipeline showing an object pickled to a string, stored, and unpickled back

The structure survives the round trip, which is what a text file cannot manage.

And the object that comes back is equal to the original without being the same object — which is the next slide, and it is a familiar distinction.

18. Worked example: equal, but not the same object

Worked example

Chapter 10's distinction, in a new place.

>>> t1 = [1, 2, 3]
>>> s = pickle.dumps(t1)
>>> t2 = pickle.loads(s)
>>> t1 == t2
True
>>> t1 is t2
False
ExpressionWhat it asksResult
t1 == t2same valueTrue
t1 is t2not the same objectFalse
the effectthe same as copyinga second object with equal contents

Compare the values.

Why: Although the new object has the same value as the old, it is not in general the same object.

Compare the identities.

Why: is reports False, exactly as it did for two separately built lists in lesson 10c.

Name what happened.

Why: Pickling and then unpickling has the same effect as copying the object.

Figure (svg): A state diagram showing two equal lists produced by pickling and unpickling

Two boxes, equal contents — which is exactly the picture of a copy.

Equivalent and not identical — the distinction from chapter 10, and the neatest possible summary of what pickling does.

Verify: Ask why it could not be the same object.

Why: Because the string is all that travels: everything about the original except its contents is discarded, and loads builds something new from that description. That is also why pickling is a way of copying, and why it works when the string has been to disk and back in another program run entirely.

19. Predict: is it the same object?

Prediction

Pickle out and back again.

t1 = [1, 2, 3]
t2 = pickle.loads(pickle.dumps(t1))
print(t1 == t2, t1 is t2)
OperatorWhat it asksResult
==same contentsTrue
isa new objectFalse
the effectthe same as copying

Predict first

What does this print?

  • True False
  • True True
  • False False
  • False True

Correct: True False — the new object has the same value as the old but is not the same object.

Why: Only the contents travel through the string, so loads builds something new that happens to be equal. The book's summary is exact: pickling and then unpickling has the same effect as copying the object — which is lesson 10c's equivalent-but-not-identical distinction in a new setting.

20. Worked example: what pickling is for

Worked example

It removes dbm's limitation, and there is a module for the combination.

# dbm alone: keys and values must be strings or bytes
db['scores'] = [1, 2, 3]        # error

# with pickle:
db['scores'] = pickle.dumps([1, 2, 3])
scores = pickle.loads(db['scores'])

# and the combination has its own module: shelve
ApproachWhat happensNote
a list into dbmnot a stringerror
pickled firsta stringaccepted
unpickled on the way outthe list againstructure preserved

Note the limitation.

Why: The keys and values have to be strings or bytes, so a list cannot be stored directly.

Pickle on the way in and unpickle on the way out.

Why: You can use pickle to store non-strings in a database.

Note the shortcut.

Why: This combination is so common that it has been encapsulated in a module called shelve.

Figure (svg): The state of the program after each line of Worked example what pickling is for, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Arbitrary objects in a key-value database, with two conversions. shelve does the same thing without the explicit calls.

Verify: Ask what is given up compared with a text file.

Why: Readability. A pickled value is meant to be easy for pickle to interpret and is not obvious to human readers, so you cannot inspect the stored data with a text editor or read it from another language. That is the trade: structure preserved, legibility lost.

21. Trap: expecting to read a pickled file

Trap

The trap

A program pickles its data and someone opens the file in an editor to check it.

Assume a saved file is readable

Why: Which is true of the text files the chapter started with.

The format isn't obvious to human readers — it begins with control characters and encodes the structure in a way meant for pickle. Inspecting it tells you nothing.

The fix

Choose the format by who will read it.

A text file if a person or another program must read it

Why: Flat, legible, and it loses the types.

Pickle if only this program will read it back

Why: Structure preserved, legibility lost.

The two are for different jobs rather than better and worse. Chapter 13's histogram could go either way; a nested structure of dictionaries and lists really only has one option.

22. Discriminate: text file or pickle?

Discrimination

Ask who will read it back.

Sort into buckets

For each case, which storage format fits?

a text file
a word-count list someone will inspect; data another program in another language must read; a report a person will read
pickle
a nested dictionary of lists of tuples; a program's internal state, reloaded by itself; an object graph with types that matter
txt
Each has a human or a foreign program as the reader, so legibility is the requirement — and the data are flat enough that losing the types costs little.
pkl
Each is structured enough that encoding it as text would mean writing a parser, and the only reader is the program itself, so the unreadable format costs nothing.

23. Complete it: store a list in a database

Faded example

dbm takes strings; pickle makes one.

Fill in the blanks

db['scores'] = pickle.dumps([1, 2, 3])

Why: dumps is short for dump string: it takes an object and returns a string representation suitable for storage. loads reconstitutes it on the way out. The pair exists precisely because a dbm database's keys and values have to be strings or bytes, which a list is not.

24. Explain it: why not just use str?

Explain it

str(t) also turns a list into a string.

Discussion prompt

A classmate asks why pickle exists when str([1, 2, 3]) already gives a readable string. Explain the difference.

Hint: Try going the other way.

Answer:

str is one-way. It produces '[1, 2, 3]', and there is no general function that turns that back into a list — you would have to write a parser, and it would break on the first string containing a comma.

pickle is a matched pair: dumps produces a form that loads can reconstitute exactly, for almost any type of object, including nested structures.

The price is that pickle's output is not obvious to human readers, where str's is. So str is for showing something to a person and pickle is for getting it back — two different jobs that both happen to produce strings.

25. Pipes: running another program

Section

Section 3

26. A running program that behaves like a file

Concept

Most operating systems provide a command-line interface, also known as a shell. Any program that you can launch from the shell can also be launched from Python using a pipe object, which represents a running program.

>>> cmd = 'ls -l'
>>> fp = os.popen(cmd)
>>> res = fp.read()
>>> stat = fp.close()
>>> print(stat)
None
LineWhat happensNote
os.popen(cmd)launches the programthe argument is a shell command
the return valuebehaves like an open fileread or readline
closereturns the final statusNone means it ended normally

The argument is a string containing a shell command, and the return value is an object that behaves like an open file. You can read the output one line at a time with readline or get the whole thing at once with read.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 142-143

27. Picture it: another program's output as a file

Picture it

The familiar file interface, over something that is not a file.

Figure (svg): A pipeline showing a shell command launched and its output read like a file

open, read, close — the same three steps as a file, over a running program.

That reuse of a familiar interface is the point. Nothing new has to be learned except which command to run.

28. Worked example: the md5sum checksum

Worked example

The book's example, and it does something genuinely useful.

>>> filename = 'book.tex'
>>> cmd = 'md5sum ' + filename
>>> fp = os.popen(cmd)
>>> res = fp.read()
>>> stat = fp.close()
>>> print(res)
1e0033f0ed0656636de0d75144ba32e0  book.tex
>>> print(stat)
None
StepWhat happensNote
the commandbuilt by concatenationa string
readthe whole outputchecksum and filename
closeNoneended with no errors

Build the command string.

Why: The argument is a string that contains a shell command, so the filename is concatenated into it.

Read the output.

Why: res holds everything the program printed, as a string.

Check the status.

Why: The return value of close is the final status of the process, and None means that it ended normally, with no errors.

Figure (svg): The state of the program after each line of Worked example the md5sum checksum, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A checksum computed by another program and read back into Python. The probability that different contents yield the same checksum is very small.

Verify: Use it for what it is good at.

Why: Comparing two files: identical checksums mean the contents are almost certainly the same, and different ones mean they are definitely not — without reading either file into Python at all. That is an efficient way to check whether two files have the same contents, which is exactly what the book offers it for.

29. Predict: what does close return?

Prediction

The command ran without errors.

fp = os.popen('ls -l')
res = fp.read()
stat = fp.close()
print(stat)
StepWhat happensResult
the commandsucceededno errors
closethe final statusNone
a failurewould give something else

Predict first

What does this print?

  • None
  • 0
  • The output of ls
  • True

Correct: None — the return value is the final status of the process, and None means that it ended normally.

Why: This is worth knowing precisely, because None usually means nothing to report and here it means success. A non-None status indicates a failure — and since a failing command often produces empty output, checking the status is the only reliable way to tell a failure from a legitimately empty result.

30. Worked example: reading the exit status

Worked example

None is success, and anything else is not.

>>> fp = os.popen('ls /nonexistent')
>>> res = fp.read()
>>> stat = fp.close()
>>> stat is None
False        # the command failed
StageWhat happensNote
a failing commandproduces little or no outputon the read
closea non-None statusthe process ended with an error
checking itthe only way to knowread alone would not tell you

Notice the read may look normal.

Why: A failing command often produces empty output, which is indistinguishable from a command that legitimately found nothing.

Check the status from close.

Why: None means the process ended normally; anything else means it did not.

Act on it.

Why: Without checking, a failed command silently becomes an empty result — which the rest of the program will treat as data.

Figure (svg): A panel showing the exit status distinguishing a successful command from a failed one

The status is the only reliable indication of success. Reading the output alone cannot distinguish failure from a legitimately empty result.

Verify: Ask what happens if the command does not exist at all.

Why: The shell reports the failure and the status is non-None, so the same check catches it — a missing program and a failing one look the same from Python's side. That makes the status check worth doing every time rather than only where you expect trouble.

31. Trap: building a command by concatenating untrusted text

Trap

The trap

A program builds cmd = 'ls ' + name where name came from a user or a file.

Insert the value into the command

Why: Which is how the book's md5sum example is written.

The string is handed to a shell, which interprets characters like ; and | — so a name containing them runs whatever follows as a separate command. The book's own example is fine because the filename is a literal; the pattern is not safe with values you did not write.

The fix

Keep shell commands built from values you control.

Literals and program-generated names are fine

Why: Which covers most of what this technique is for.

For anything else, use the subprocess module

Why: Which the book's own footnote points to: popen is deprecated, and subprocess is the replacement.

The book keeps using popen because for simple cases subprocess is more complicated than necessary — and simple cases means commands you constructed yourself, which is worth making explicit.

32. Compare: a file and a pipe

Comparison

Fill the blanks. The interface is deliberately the same.

Comparison matrix

QuestionA fileA pipe
How do you get one?open(name)os.popen(command)
How do you read it?read or readlineread or readline
What does close return?nothing usefulthe process's final status
What is behind it?bytes on a diska running program

Only the first and third rows differ, which is why a pipe needs almost no new learning.

33. Complete it: run a command and read its output

Faded example

The pipe behaves like an open file.

Fill in the blanks

fp = os.popen('ls -l')
res = fp.read()
stat = fp.close()

Why: os.popen takes a string containing a shell command and returns an object that behaves like an open file. The book's footnote notes that popen is deprecated in favour of the subprocess module, and that for simple cases subprocess is more complicated than necessary — so it keeps using popen.

34. Where running another program is the right answer

Real world

Not everything needs to be written in Python.

Discussion prompt

Think of a task where an existing command-line tool would do the job better than code you could write. What makes it better?

Hint: The book's example computes a checksum.

Answer:

Checksums, image conversion, compression, searching a huge file — all cases where a specialised tool has had years of work put into being correct and fast.

What makes it better is not cleverness but maturity: the edge cases have been found, and the performance work has been done, by people who only did that.

So a pipe is a way of borrowing that. The cost is a dependency on the tool being installed and a command string that has to be right — which is why checking the exit status matters, since a missing program looks exactly like a failing one.

35. Writing modules

Section

Section 4

36. Any file of Python code can be imported

Concept

Any file that contains Python code can be imported as a module. Suppose you have a file named wc.py that defines a function and then calls it.

def linecount(filename):
    count = 0
    for line in open(filename):
        count += 1
    return count

print(linecount('wc.py'))
UsageWhat happensNote
run as a scriptprints 7it reads itself
import wcprints 7 as wellthe test code runs
wc.linecount('wc.py')7the function is available

If you run this program, it reads itself and prints the number of lines in the file, which is 7. You can also import it — and then you have a module object which provides linecount.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 143-143

37. Picture it: the same file, used two ways

Picture it

One as a program to run, one as a library to import.

Figure (svg): Two columns showing the same file run as a script and imported as a module

Importing runs the file, which is the problem the next slide solves.

The middle row is the surprise: normally when you import a module, it defines new functions but it doesn't run them — and this file runs its test code either way.

38. Worked example: the __name__ idiom

Worked example

The one line that makes a file work as both.

def linecount(filename):
    count = 0
    for line in open(filename):
        count += 1
    return count

if __name__ == '__main__':
    print(linecount('wc.py'))
UsageThe value of __name__What happens
run as a script__name__ is '__main__'the test code runs
imported__name__ is 'wc'the test code is skipped
either waylinecount is definedwhich is the point

Note the problem.

Why: The only problem with the earlier example is that when you import the module it runs the test code at the bottom.

Note the built-in variable.

Why: __name__ is a built-in variable that is set when the program starts. If the program is running as a script, __name__ has the value '__main__'.

Note the effect.

Why: In that case the test code runs; otherwise, if the module is being imported, the test code is skipped.

Figure (svg): A flowchart showing the __name__ check routing between script and module behaviour

The definitions always run; only the test code is conditional.

A file that is a usable program and a usable library at once. Programs that will be imported as modules often use this idiom, and it costs one line.

Verify: Check the value of __name__ when imported.

Why: It is the module's name — 'wc' for wc.py — which is what the book sets as an exercise. Knowing that the variable holds something rather than nothing makes the condition read as a comparison rather than a magic incantation.

39. Predict: what does __name__ hold?

Prediction

The file is being imported rather than run.

# in wc.py:
print(__name__)

# and elsewhere:
# >>> import wc
UsageThe valueNote
run as a script'__main__'the special value
importedthe module's name'wc'
the checkdistinguishes the two

Predict first

What does it print when wc is imported?

  • wc
  • __main__
  • None
  • Nothing — printing is skipped on import

Correct: wc — when a file is imported, __name__ holds the module's name rather than '__main__'.

Why: __name__ is a built-in variable set when the program starts: '__main__' if the file is running as a script, and the module's own name if it is being imported. Nothing is skipped on import — the whole file runs, which is exactly why the idiom is needed.

40. Worked example: the re-import warning

Worked example

Python does nothing the second time.

>>> import wc
7
>>> import wc        # nothing happens
>>>

# even if wc.py has been edited since
ActionWhat happensNote
the first importreads the fileand runs it
the seconddoes nothingnot even a re-read
after an editstill the old versionuntil you restart

Note the behaviour.

Why: If you import a module that has already been imported, Python does nothing — it does not re-read the file, even if it has changed.

Note when it bites.

Why: Editing a module while an interpreter session is open, and finding your changes have no effect.

Note the safe response.

Why: You can use the built-in function reload, but it can be tricky, so the safest thing to do is restart the interpreter and then import the module again.

Figure (svg): The state of the program after each line of Worked example the re-import warning, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A silent staleness. The symptom is that a fix appears not to work, and the cause is that the fixed file was never read.

Verify: Recognise the symptom before investigating the code.

Why: If a change has no effect at all — not a wrong effect, but none — and the module was imported earlier in the same session, the file has not been re-read. Restarting first costs seconds and rules out a whole category of imaginary bug.

41. Trap: leaving test code at the bottom of a module

Trap

The trap

A file defines useful functions and calls one at the bottom to check it works.

Test the function where it is defined

Why: Which is convenient while writing it.

Importing the file now runs the test — printing output, reading files, and doing work the importer never asked for. Normally when you import a module, it defines new functions but it doesn't run them.

The fix

Guard it with the __name__ check.

if __name__ == '__main__':

Why: The test code then runs only when the file is run as a script.

And the definitions still happen either way

Why: Which is what an importer wants.

This is one line, it costs nothing, and it makes every script you write importable. It is worth adopting as a default habit rather than adding when the need arises.

42. Two truths and a lie: modules

Two truths and a lie

Two are true. Keep the lie.

Eliminate the wrong options

Rule out the two true statements.

  • A. Any file that contains Python code can be imported as a module
  • B. Importing a module that has already been imported does nothing, even if the file changed
  • C. Importing a module defines its functions without running any of its other code

Survives elimination: C

Why: C describes what people expect and not what happens. Importing runs the whole file top to bottom, including any calls at the bottom — which is precisely the problem the __name__ idiom solves. What is true is the weaker statement that importing does not CALL the functions it defines.

43. Complete it: guard the test code

Faded example

Run it as a script, skip it on import.

Fill in the blanks

if __name__ == '__main__':
print(linecount('wc.py'))

Why: __name__ is a built-in variable set when the program starts, holding '__main__' when the file is run as a script and the module's name when it is imported. The guard costs one line and is what makes a file usable as both a program and a library.

44. Explain it: my fix did nothing at all

Explain it

Not a wrong result — no change whatsoever.

Discussion prompt

A classmate edited a module, ran their code again in the same interpreter session, and the behaviour is byte-for-byte identical. Diagnose it.

Hint: Was the file read again?

Answer:

The module was already imported, so the second import did nothing — Python does not re-read the file, even if it has changed. Their edit is on disk and not in the session.

The signature is that nothing changed at all. A wrong fix produces different wrong behaviour; an unread fix produces exactly what it did before.

The safest response is to restart the interpreter and import again. There is a reload function, but it can be tricky, and restarting rules out the whole question in a couple of seconds.

45. Debugging invisible characters

Section

Section 5

46. repr makes whitespace visible

Concept

When you are reading and writing files, you might run into problems with whitespace. These errors can be hard to debug because spaces, tabs and newlines are normally invisible.

>>> s = '1 2\t 3\n 4'
>>> print(s)
1 2      3
 4
>>> print(repr(s))
'1 2\t 3\n 4'
CallWhat you seeNote
print(s)the characters are renderedthe tab and newline act
print(repr(s))backslash sequencesyou can see what is there
the differencerendering against describing

The built-in function repr takes any object as an argument and returns a string representation of it. For strings, it represents whitespace characters with backslash sequences — which is exactly what you need when the problem is a character you cannot see.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 144-144

47. Picture it: two views of one string

Picture it

The same value, rendered and described.

Figure (svg): Two columns showing a string printed normally and printed through repr

One shows what the string does; the other shows what it is.

Which is why repr is the first thing to reach for when a file's contents look right and behave wrong.

48. Worked example: finding a stray character

Worked example

Two strings that print identically.

>>> a = 'total'
>>> b = 'total '
>>> print(a)
total
>>> print(b)
total
>>> a == b
False
>>> print(repr(b))
'total '
StepWhat you seeNote
printing bothidentical outputthe space is invisible
comparingFalseand the reason is not visible
reprthe trailing space is showninside the quotes

Notice the mystery.

Why: Two strings print the same and compare unequal, which looks like a bug in the comparison.

Reach for repr.

Why: It takes any object and returns a string representation, showing whitespace as it is.

Read the answer off.

Why: The quotes delimit the string, so a trailing space is plainly visible where printing hid it.

Figure (svg): The state of the program after each line of Worked example finding a stray character, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

A trailing space, found in one line. This class of bug is common when reading files, where lines carry their newline and fields may be padded.

Verify: Apply it to a line read from a file.

Why: repr(line) shows the trailing '\n' that every line carries, which explains why a comparison against a plain word fails. That is the commonest whitespace bug in file handling, and strip is the usual fix — but only once you can see that the newline is there.

49. Predict: what does repr show?

Prediction

The string contains a tab.

s = 'a\tb'
print(repr(s))
CallWhat you seeNote
print(s)a, a tab's worth of space, bthe tab acts
repr(s)the escape sequenceshown literally
the quotesdelimit the stringso the ends are visible

Predict first

What does this print?

  • 'a\tb'
  • a b
  • ab
  • a\tb without quotes

Correct: 'a\tb' — repr represents whitespace characters with backslash sequences, and includes the quotes.

Why: print(s) would render the tab as horizontal space, hiding what character produced it. repr describes the string instead of rendering it, showing the escape sequence and delimiting the whole with quotes — which is also how a trailing space becomes visible.

50. Worked example: the line-ending problem

Worked example

Different systems disagree about how a line ends.

# some systems use a newline:        \n
# others use a return character:     \r
# some use both:                     \r\n

>>> repr(line)
'total\r\n'      # two characters, not one
EndingWhat it isNote
\na newlineone convention
\ra returnanother
\r\nbotha third

Note the disagreement.

Why: Different systems use different characters to indicate the end of a line. Some use a newline; others use a return character; some use both.

Note when it causes trouble.

Why: If you move files between different systems, these inconsistencies can cause problems.

Note how to see it.

Why: repr shows exactly which characters are there, which is the only reliable way to tell \n from \r\n.

Figure (svg): Three line-ending conventions shown as the characters they consist of

A file written under one convention and read under another carries a visible extra character — visible, that is, through repr.

Three conventions, and a file from one system read on another carries the wrong ones. The symptom is usually a stray character at the end of every field.

Verify: Check what strip does about it.

Why: strip() with no argument removes whitespace including both \r and \n, so it handles all three conventions at once — which is why stripping lines read from a file is such a common habit. For most systems there are also applications to convert from one format to another, or you could write one yourself.

51. Trap: comparing a line from a file to a plain string

Trap

The trap

A program reads lines and tests if line == 'total':, which never matches.

Compare what you read to what you expect

Why: The line looks like the word when printed.

Every line carries its newline, so the value is 'total\n' and the comparison fails. Printing it shows total, which makes the comparison look correct and the bug invisible.

The fix

Strip the line, and use repr when unsure.

if line.strip() == 'total':

Why: Which removes the newline and any stray spaces.

And print(repr(line)) when a comparison fails inexplicably

Why: It shows exactly what is there.

The general rule for file work: whenever two things look the same and compare unequal, look at them through repr before looking at the comparison. The characters you cannot see are the usual explanation.

52. Error analysis: four whitespace confusions

Error analysis

Mark each and say what repr would reveal.

Annotate

  • Line 1 fails for every line read from a file, because each carries its newline: the value is 'total\n' rather than 'total'.
  • Line 2 tests against a string with a trailing space, which is legal and almost certainly a typo — repr of the literal would show it, and nothing else would.
  • Line 3 is the fix: strip removes whitespace from both ends, including \n and \r, so it handles every line-ending convention at once.
  • Line 4 is a diagnostic: a length one greater than the visible characters means a line ending is present, and two greater means \r\n.
  • So two are bugs, one is the fix, and one is a way of detecting the problem without repr — though repr shows you which character it is.
  • The general habit: strip lines as you read them, and reach for repr the moment a comparison fails inexplicably.

All four are invisible when printed, which is exactly why the book introduces repr in a chapter about files.

53. Complete it: see the whitespace

Faded example

Describe the string rather than rendering it.

Fill in the blanks

print(repr(line)) # shows \t and \n literally

Why: repr takes any object and returns a string representation of it, and for strings it represents whitespace characters with backslash sequences. Printing the string directly renders those characters instead, which is exactly what hides them — so repr is the tool when the problem is something you cannot see.

54. Where invisible differences cause trouble

Real world

Two things that look identical and are not.

Discussion prompt

Think of a case outside programming where two entries looked the same and were treated as different. What was invisible?

Hint: A space you cannot see.

Answer:

A trailing space in a spreadsheet cell that splits one category into two; a name pasted with a non-breaking space; two apparently identical codes that will not match.

What is invisible is a character that takes up no distinguishable room — and every system involved renders it rather than describing it, so nobody can see the difference.

Which is what repr is for: a view that describes rather than renders. The general lesson is that when two things look identical and behave differently, the display is hiding something, and you need a different view rather than a closer look.

55. Compare: three ways to make data persist

Comparison

Fill the blanks. Each keeps a different amount of the structure.

Comparison matrix

QuestionA text fileA dbm databasePickle
Readable by a person?yespartly — the values are textno
Keeps types?no — everything becomes charactersno — strings or bytes onlyyes, almost any object
Random access by key?noyes — like a dictionarynot by itself
When to use itflat data someone will readkey-value data that must persistprogram state, reloaded by itself

And the combination of the last two is so common it has its own module, shelve.

56. The procedure: writing a file that is also a module

Pattern

Five steps, and the fourth is the one worth making a habit.

  1. Put the definitions at the top: functions and constants, with nothing calling them.
  2. Put anything the file should do when run as a program at the bottom.
  3. Wrap that bottom part in if __name__ == '__main__':.
  4. Test both ways — run it as a script, then import it and check nothing unexpected happens.
  5. If you edit it while an interpreter is open, restart rather than re-importing, since a second import does nothing.

Step 5 is the one that wastes the most time when forgotten. A fix that has no effect at all — rather than a different wrong effect — almost always means the file was never re-read.

Python documentation — pickle — Python object serialization pickle — Python object serialization

57. Check yourself 1 of 3: pickling

Check

Out to a string and back again.

t1 = [1, 2, 3]
t2 = pickle.loads(pickle.dumps(t1))
print(t1 is t2)
StepWhat happensResult
dumpsthe contents become a stringthe object does not travel
loadsbuilds something newfrom the description
istwo objectsFalse

Check your understanding

What does this print?

  • A. False (correct)
  • B. True
  • C. None
  • D. An error, since lists cannot be pickled

Answer: A

Why: Although the new object has the same value as the old, it is not in general the same object — so == is True and is is False. The book's summary is exact: pickling and then unpickling has the same effect as copying the object.

Why B tempts people
This would require loads to return the original object, which it cannot: only the contents travelled, as a string.
Why C tempts people
Both functions return values. Nothing here produces None.
Why D tempts people
pickle translates almost any type of object, and a list of integers is among the easiest.

58. Check yourself 2 of 3: the __name__ idiom

Check

The file is imported rather than run.

Check your understanding

What does if __name__ == '__main__': achieve?

  • A. The guarded code runs when the file is run as a script and is skipped when it is imported (correct)
  • B. It stops the module's functions being defined on import
  • C. It makes the module reload if the file has changed
  • D. It prevents the module being imported twice

Answer: A

Why: __name__ is a built-in variable set when the program starts: '__main__' when the file runs as a script, and the module's name when it is imported. So the guarded test code runs only in the first case — which is what stops an import doing work the importer never asked for.

Why B tempts people
The definitions always run, which is the point of importing. Only the guarded code is conditional.
Why C tempts people
Nothing here affects reloading. Python does nothing on a second import regardless.
Why D tempts people
Python already does nothing on a repeat import, with or without the idiom.

59. Check yourself 3 of 3: repr

Check

Two strings that print identically.

Check your understanding

Why does the book introduce repr in a chapter about files?

  • A. Because whitespace errors are hard to debug when spaces, tabs and newlines are invisible (correct)
  • B. Because repr converts a string to bytes for writing
  • C. Because files can only store repr output
  • D. Because repr is faster than print

Answer: A

Why: repr takes any object and returns a string representation, showing whitespace characters as backslash sequences. That matters most for file work, where every line carries a line ending and different systems use different characters for it — differences that are invisible when printed.

Why B tempts people
repr produces a string, not bytes. Encoding is a different operation.
Why C tempts people
Files store whatever you write. repr is a debugging aid, not a storage format.
Why D tempts people
Speed is irrelevant; the two do different things. repr describes where print renders.

60. Where this shows up outside this course

Real world

Saving something in a form you can get back is a design decision with a cost.

Discussion prompt

Think of a file format you can open and read, and one you cannot. What does each buy, and what happens when the program that wrote the unreadable one is gone?

Hint: Plain text against a proprietary format.

Answer:

A readable format can be inspected, fixed by hand, read by other tools, and understood in ten years. It usually stores less structure and takes more space.

An opaque format keeps everything the program knew and is useless without that program. If it stops working or nobody has it any more, the data are effectively gone.

Which is exactly the text-file-against-pickle trade. Neither is right in general, and the question worth asking is who will need to read this, and when — because that is what decides which cost you can afford.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence.

Predict first

You import a module, edit its file, and import it again in the same session. What happens?

  • Nothing — Python does not re-read the file, even though it changed
  • The new version is loaded and replaces the old
  • An error, since the module is already imported
  • Both versions exist, and the newer one wins

Correct: Nothing — if you import a module that has already been imported, Python does nothing, and it does not re-read the file even if it has changed.

Why: The book gives this as an explicit warning, and it costs people real time because the symptom is so misleading: your fix appears to have no effect at all. That is the tell — a wrong fix produces different wrong behaviour, whereas an unread fix produces byte-for-byte what it did before. There is a built-in reload function, but the book notes it can be tricky, so the safest thing to do is restart the interpreter and import the module again. Making that the reflex whenever a change seems to do nothing rules out a whole category of imaginary bug in a couple of seconds.

62. Explain it to someone else

Explain it

Three ways to save data, and how to choose.

Discussion prompt

A classmate asks whether to save their program's data as a text file, in a dbm database, or with pickle. Give them the question that decides it.

Hint: Who reads it back?

Answer:

Ask who will read it. If a person or another program must, it has to be a text file — which means the data have to be flat enough to write as lines, and the types will be lost.

If only their own program will read it back and the structure matters, pickle: it translates almost any type of object into a string and back, and the format is not meant for human readers.

And if they want to look things up by key without loading everything, a dbm database — which behaves like a dictionary on disk, with the limitation that keys and values must be strings or bytes. Combining it with pickle is so common that shelve exists to do exactly that.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • dbm databases, and how they differ from dictionaries
  • Pickling, and what unpickling produces
  • Pipes, and reading another program's output and status
  • Writing modules, the __name__ idiom, and repr

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The database is the easiest of the four, because almost everything transfers from dictionaries — the bytes and the missing methods are the only surprises. Pickling is one idea with a memorable summary: it has the same effect as copying. Pipes reuse the file interface entirely, and the one thing worth remembering is that None means success. And the __name__ idiom has the longest reach of anything in this lesson: it is one line, and it makes every script you write importable.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw three boxes labelled text file, database and pickle, and beside each write what it preserves and what it loses. Underneath, write the round trip pickle.loads(pickle.dumps(t)) and note what == and is report about the result. Then write the __name__ guard from memory, with the value the variable holds in each of the two cases, and finish with one string shown both ways — printed, and through repr.

65. What you can do now

Recap

Four pages, and chapter 14 is finished: data that outlive the program, several ways.

If you remember one thingIt is this
From dbmA database is a dictionary on disk, with strings or bytes for keys and values.
From picklePickling and unpickling has the same effect as copying the object.
From pipesclose returns the status, and None means it ended normally.
From modulesImporting runs the whole file. The __name__ guard is what stops that mattering.
From reprprint renders; repr describes. When something looks right and behaves wrong, describe it.

The next chapter begins the object-oriented half of the book: programmer-defined types, with a class that represents a point in two-dimensional space — attributes, instances, and objects as return values.

Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-144 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §14.6-14.10, pp. 141-144
  2. Python documentation — pickle — Python object serialization
  3. Python documentation — sqlite3 — DB-API 2.0 interface for SQLite databases
  4. Python documentation — Modules

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108