This lesson covers dbm databases that behave like dictionaries on disk, the pickle module for storing arbitrary objects, pipes for running other programs, writing importable modules with the __name__ idiom, and repr for debugging invisible whitespace.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 14 — Files
§14.6-14.10, pp. 141-144
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-144 — the pages these objectives are drawn from
Warm-up
A text file holds characters. A dictionary does not.
Discussion prompt
You want to save {'a': 1, 'b': [2, 3]} so a later run can load it back. Writing it with str gives you a string that looks right. What is the difficulty in reading it back?
Hint: What would you have to write to turn that string into a dictionary?
Answer:
You would have to parse it: find the braces, split on commas that are not inside brackets, split each pair on its colon, and work out the type of every value.
Which is a small programming language interpreter, and it goes wrong on the first value containing a comma or a colon — the separator problem from lesson 12c, at full scale.
This lesson has two answers: a database that stores key-value pairs directly, and a module that converts any object to a string and back again without your writing any of that.
Concept
A database is a file that is organised for storing data. Many databases are organised like a dictionary in the sense that they map from keys to values.
database — A file whose contents are organised like a dictionary with keys that correspond to values.
The biggest difference between a database and a dictionary is that the database is on disk — or other permanent storage — so it persists after the program ends. Everything else about using one is familiar.
Figure (svg): Two columns comparing a dictionary in memory with a database on disk
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-141
Section
Section 1
Concept
The module dbm provides an interface for creating and updating database files. Opening a database is similar to opening other files.
>>> import dbm
>>> db = dbm.open('captions', 'c')
>>> db['cleese.png'] = 'Photo of John Cleese.'
>>> db['cleese.png']
b'Photo of John Cleese.'
>>> db.close()| Line | What happens | Note |
|---|---|---|
| mode 'c' | create if it does not exist | and open it if it does |
| assignment | dbm updates the database FILE | immediately |
| lookup | dbm reads the file | and returns bytes |
The mode 'c' means that the database should be created if it doesn't already exist — unlike a file's mode 'w', which destroys an existing one. The result is a database object that can be used, for most operations, like a dictionary.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-141
Picture it
There is no separate saving step.
Figure (svg): A pipeline showing an assignment to a database object updating the file immediately
That is the whole point: the structure and the storage are the same thing, so nothing has to be written out at the end.
Worked example
The one visible difference from a dictionary.
>>> db['cleese.png'] = 'Photo of John Cleese.'
>>> db['cleese.png']
b'Photo of John Cleese.'
>>> db['cleese.png'] = 'Photo of John Cleese doing a silly walk.'
>>> db['cleese.png']
b'Photo of John Cleese doing a silly walk.'| Operation | What happens | Note |
|---|---|---|
| what you store | a string | ordinary |
| what comes back | a bytes object | note the b |
| a second assignment | replaces the old value | as in a dictionary |
Store a string.
Why: The assignment looks exactly like a dictionary's, and dbm writes it to the file.
Read it back.
Why: The result is a bytes object, which is why it begins with b. A bytes object is similar to a string in many ways.
Replace it.
Why: If you make another assignment to an existing key, dbm replaces the old value — the same behaviour as a dictionary.
Figure (svg): The state of the program after each line of Worked example what comes back is bytes, drawn as a ladder with one rung per traced line
Values come back as bytes rather than strings. When you get farther into Python the difference becomes important, and for now it can be ignored.
Verify: Check that the bytes compare as expected.
Why: b'abc' == 'abc' is False, so a comparison against a plain string fails even when the text matches — which is the one place the difference bites early. Decoding with .decode() gives an ordinary string back if you need one.
Prediction
A string went in.
db['a.png'] = 'a caption'
print(db['a.png'])| Step | What it is | Note |
|---|---|---|
| stored | a string | ordinary |
| retrieved | a bytes object | begins with b |
| the difference | matters later | ignorable for now |
Predict first
What does this print?
Correct: b'a caption' — the result is a bytes object, which is why it begins with b.
Why: A bytes object is similar to a string in many ways, and printing one shows the b prefix. Storing a plain string is fine; it is the retrieval that converts. The practical consequence is that comparing the result against a plain string fails, so .decode() is needed when the value has to be an ordinary string.
Worked example
It is like a dictionary for most operations, not all.
# some dictionary methods do not work:
# db.keys() and db.items() are not available as usual
# but iteration with a for loop does:
for key in db.keys():
print(key, db[key])
db.close()| Operation | Available? | Note |
|---|---|---|
| keys, items | do not work as they do for a dictionary | the interface is narrower |
| iteration | works | one key per pass |
| close | required, as with other files | when you are done |
Note the missing methods.
Why: Some dictionary methods, like keys and items, don't work with database objects — the interface resembles a dictionary rather than matching it.
Note that iteration works.
Why: A for loop over the database gives its keys, which covers most of what those methods were for.
Close it when done.
Why: As with other files, you should close the database when you are finished.
Figure (svg): Two columns separating the dictionary operations a database supports from those it does not
Most operations transfer and a few do not. The practical approach is to use assignment, lookup and iteration, which are all supported.
Verify: Ask why the interface is narrower.
Why: Because the data are on disk, so an operation that would be a quick pass over memory becomes a pass over a file — and some, like producing a list of every item at once, are exactly what you would want to avoid on a large database. The gaps are mostly places where the convenient dictionary operation would be an expensive database one.
Trap
A program writes many items to a database and ends without closing it.
Assume the writes have happened
Why: Each assignment updated the file, so it looks finished.
As with other files, the database should be closed when you are done — and until it is, the file may not be complete on disk. A program that writes and then exits abruptly can leave it in a partial state.
Close it, as you would a file.
db.close() when you have finished
Why: The book says so directly: as with other files, you should close the database.
And treat it as the point the data are safe
Why: The same reasoning as closing a written file.
This is the same habit as lesson 14a's close, for the same reason. Persistence is only useful if the data actually reached the disk, and closing is where that is guaranteed.
Comparison
Fill the blanks. The interface is nearly the same.
Comparison matrix
| Question | Dictionary | Database |
|---|---|---|
| Where does it live? | in memory | on disk |
| Does it survive the program? | no | yes |
| What can the keys be? | any hashable type | strings or bytes only |
| Do keys() and items() work? | yes | not as they do for a dictionary |
The third row is the limitation pickle addresses, and it is the reason the next idea exists.
Faded example
Create it if it is not there.
Fill in the blanks
import dbm
db = dbm.open('captions', 'c')
Why: Mode 'c' means the database should be created if it doesn't already exist, and opened if it does — which is what you almost always want. Note that this is not the same as a file's mode 'w', which destroys an existing file; a database opened with 'c' keeps its contents.
Explain it to yourself
The interface is a deliberate choice.
Discussion prompt
dbm could have offered store and fetch functions. Why present it as something you index with brackets instead?
Hint: What do you already know how to use?
Answer:
Because you already know how to use a dictionary. Presenting the database the same way means nearly everything you know transfers with no new syntax at all.
And because the underlying idea genuinely is the same: many databases are organised like a dictionary, mapping from keys to values. The interface matches the concept rather than disguising it.
The cost is that the resemblance is not exact, so the gaps surprise you — keys and items do not work as usual, and the values come back as bytes. A familiar interface that is almost the same is easier to learn and slightly harder to trust.
Section
Section 2
Concept
A limitation of dbm is that the keys and values have to be strings or bytes; if you try to use any other type, you get an error. The pickle module can help: it translates almost any type of object into a string suitable for storage in a database, and then translates strings back into objects.
>>> import pickle
>>> t = [1, 2, 3]
>>> pickle.dumps(t)
b'\x80\x03]q\x00(K\x01K\x02K\x03e.'
>>> t2 = pickle.loads(pickle.dumps(t))
>>> t2
[1, 2, 3]| Function | What it does | Note |
|---|---|---|
| dumps | dump string | an object to a string |
| the format | not obvious to human readers | meant for pickle |
| loads | load string | reconstitutes the object |
The format isn't obvious to human readers; it is meant to be easy for pickle to interpret. That is the trade against a text file, which is readable by anything and cannot restore a structure.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 142-142
Picture it
Out to a string, into storage, and back to an object.
Figure (svg): A pipeline showing an object pickled to a string, stored, and unpickled back
And the object that comes back is equal to the original without being the same object — which is the next slide, and it is a familiar distinction.
Worked example
Chapter 10's distinction, in a new place.
>>> t1 = [1, 2, 3]
>>> s = pickle.dumps(t1)
>>> t2 = pickle.loads(s)
>>> t1 == t2
True
>>> t1 is t2
False| Expression | What it asks | Result |
|---|---|---|
| t1 == t2 | same value | True |
| t1 is t2 | not the same object | False |
| the effect | the same as copying | a second object with equal contents |
Compare the values.
Why: Although the new object has the same value as the old, it is not in general the same object.
Compare the identities.
Why: is reports False, exactly as it did for two separately built lists in lesson 10c.
Name what happened.
Why: Pickling and then unpickling has the same effect as copying the object.
Figure (svg): A state diagram showing two equal lists produced by pickling and unpickling
Equivalent and not identical — the distinction from chapter 10, and the neatest possible summary of what pickling does.
Verify: Ask why it could not be the same object.
Why: Because the string is all that travels: everything about the original except its contents is discarded, and loads builds something new from that description. That is also why pickling is a way of copying, and why it works when the string has been to disk and back in another program run entirely.
Prediction
Pickle out and back again.
t1 = [1, 2, 3]
t2 = pickle.loads(pickle.dumps(t1))
print(t1 == t2, t1 is t2)| Operator | What it asks | Result |
|---|---|---|
| == | same contents | True |
| is | a new object | False |
| the effect | the same as copying |
Predict first
What does this print?
Correct: True False — the new object has the same value as the old but is not the same object.
Why: Only the contents travel through the string, so loads builds something new that happens to be equal. The book's summary is exact: pickling and then unpickling has the same effect as copying the object — which is lesson 10c's equivalent-but-not-identical distinction in a new setting.
Worked example
It removes dbm's limitation, and there is a module for the combination.
# dbm alone: keys and values must be strings or bytes
db['scores'] = [1, 2, 3] # error
# with pickle:
db['scores'] = pickle.dumps([1, 2, 3])
scores = pickle.loads(db['scores'])
# and the combination has its own module: shelve| Approach | What happens | Note |
|---|---|---|
| a list into dbm | not a string | error |
| pickled first | a string | accepted |
| unpickled on the way out | the list again | structure preserved |
Note the limitation.
Why: The keys and values have to be strings or bytes, so a list cannot be stored directly.
Pickle on the way in and unpickle on the way out.
Why: You can use pickle to store non-strings in a database.
Note the shortcut.
Why: This combination is so common that it has been encapsulated in a module called shelve.
Figure (svg): The state of the program after each line of Worked example what pickling is for, drawn as a ladder with one rung per traced line
Arbitrary objects in a key-value database, with two conversions. shelve does the same thing without the explicit calls.
Verify: Ask what is given up compared with a text file.
Why: Readability. A pickled value is meant to be easy for pickle to interpret and is not obvious to human readers, so you cannot inspect the stored data with a text editor or read it from another language. That is the trade: structure preserved, legibility lost.
Trap
A program pickles its data and someone opens the file in an editor to check it.
Assume a saved file is readable
Why: Which is true of the text files the chapter started with.
The format isn't obvious to human readers — it begins with control characters and encodes the structure in a way meant for pickle. Inspecting it tells you nothing.
Choose the format by who will read it.
A text file if a person or another program must read it
Why: Flat, legible, and it loses the types.
Pickle if only this program will read it back
Why: Structure preserved, legibility lost.
The two are for different jobs rather than better and worse. Chapter 13's histogram could go either way; a nested structure of dictionaries and lists really only has one option.
Discrimination
Ask who will read it back.
Sort into buckets
For each case, which storage format fits?
Faded example
dbm takes strings; pickle makes one.
Fill in the blanks
db['scores'] = pickle.dumps([1, 2, 3])
Why: dumps is short for dump string: it takes an object and returns a string representation suitable for storage. loads reconstitutes it on the way out. The pair exists precisely because a dbm database's keys and values have to be strings or bytes, which a list is not.
Explain it
str(t) also turns a list into a string.
Discussion prompt
A classmate asks why pickle exists when str([1, 2, 3]) already gives a readable string. Explain the difference.
Hint: Try going the other way.
Answer:
str is one-way. It produces '[1, 2, 3]', and there is no general function that turns that back into a list — you would have to write a parser, and it would break on the first string containing a comma.
pickle is a matched pair: dumps produces a form that loads can reconstitute exactly, for almost any type of object, including nested structures.
The price is that pickle's output is not obvious to human readers, where str's is. So str is for showing something to a person and pickle is for getting it back — two different jobs that both happen to produce strings.
Section
Section 3
Concept
Most operating systems provide a command-line interface, also known as a shell. Any program that you can launch from the shell can also be launched from Python using a pipe object, which represents a running program.
>>> cmd = 'ls -l'
>>> fp = os.popen(cmd)
>>> res = fp.read()
>>> stat = fp.close()
>>> print(stat)
None| Line | What happens | Note |
|---|---|---|
| os.popen(cmd) | launches the program | the argument is a shell command |
| the return value | behaves like an open file | read or readline |
| close | returns the final status | None means it ended normally |
The argument is a string containing a shell command, and the return value is an object that behaves like an open file. You can read the output one line at a time with readline or get the whole thing at once with read.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 142-143
Picture it
The familiar file interface, over something that is not a file.
Figure (svg): A pipeline showing a shell command launched and its output read like a file
That reuse of a familiar interface is the point. Nothing new has to be learned except which command to run.
Worked example
The book's example, and it does something genuinely useful.
>>> filename = 'book.tex'
>>> cmd = 'md5sum ' + filename
>>> fp = os.popen(cmd)
>>> res = fp.read()
>>> stat = fp.close()
>>> print(res)
1e0033f0ed0656636de0d75144ba32e0 book.tex
>>> print(stat)
None| Step | What happens | Note |
|---|---|---|
| the command | built by concatenation | a string |
| read | the whole output | checksum and filename |
| close | None | ended with no errors |
Build the command string.
Why: The argument is a string that contains a shell command, so the filename is concatenated into it.
Read the output.
Why: res holds everything the program printed, as a string.
Check the status.
Why: The return value of close is the final status of the process, and None means that it ended normally, with no errors.
Figure (svg): The state of the program after each line of Worked example the md5sum checksum, drawn as a ladder with one rung per traced line
A checksum computed by another program and read back into Python. The probability that different contents yield the same checksum is very small.
Verify: Use it for what it is good at.
Why: Comparing two files: identical checksums mean the contents are almost certainly the same, and different ones mean they are definitely not — without reading either file into Python at all. That is an efficient way to check whether two files have the same contents, which is exactly what the book offers it for.
Prediction
The command ran without errors.
fp = os.popen('ls -l')
res = fp.read()
stat = fp.close()
print(stat)| Step | What happens | Result |
|---|---|---|
| the command | succeeded | no errors |
| close | the final status | None |
| a failure | would give something else |
Predict first
What does this print?
Correct: None — the return value is the final status of the process, and None means that it ended normally.
Why: This is worth knowing precisely, because None usually means nothing to report and here it means success. A non-None status indicates a failure — and since a failing command often produces empty output, checking the status is the only reliable way to tell a failure from a legitimately empty result.
Worked example
None is success, and anything else is not.
>>> fp = os.popen('ls /nonexistent')
>>> res = fp.read()
>>> stat = fp.close()
>>> stat is None
False # the command failed| Stage | What happens | Note |
|---|---|---|
| a failing command | produces little or no output | on the read |
| close | a non-None status | the process ended with an error |
| checking it | the only way to know | read alone would not tell you |
Notice the read may look normal.
Why: A failing command often produces empty output, which is indistinguishable from a command that legitimately found nothing.
Check the status from close.
Why: None means the process ended normally; anything else means it did not.
Act on it.
Why: Without checking, a failed command silently becomes an empty result — which the rest of the program will treat as data.
Figure (svg): A panel showing the exit status distinguishing a successful command from a failed one
The status is the only reliable indication of success. Reading the output alone cannot distinguish failure from a legitimately empty result.
Verify: Ask what happens if the command does not exist at all.
Why: The shell reports the failure and the status is non-None, so the same check catches it — a missing program and a failing one look the same from Python's side. That makes the status check worth doing every time rather than only where you expect trouble.
Trap
A program builds cmd = 'ls ' + name where name came from a user or a file.
Insert the value into the command
Why: Which is how the book's md5sum example is written.
The string is handed to a shell, which interprets characters like ; and | — so a name containing them runs whatever follows as a separate command. The book's own example is fine because the filename is a literal; the pattern is not safe with values you did not write.
Keep shell commands built from values you control.
Literals and program-generated names are fine
Why: Which covers most of what this technique is for.
For anything else, use the subprocess module
Why: Which the book's own footnote points to: popen is deprecated, and subprocess is the replacement.
The book keeps using popen because for simple cases subprocess is more complicated than necessary — and simple cases means commands you constructed yourself, which is worth making explicit.
Comparison
Fill the blanks. The interface is deliberately the same.
Comparison matrix
| Question | A file | A pipe |
|---|---|---|
| How do you get one? | open(name) | os.popen(command) |
| How do you read it? | read or readline | read or readline |
| What does close return? | nothing useful | the process's final status |
| What is behind it? | bytes on a disk | a running program |
Only the first and third rows differ, which is why a pipe needs almost no new learning.
Faded example
The pipe behaves like an open file.
Fill in the blanks
fp = os.popen('ls -l')
res = fp.read()
stat = fp.close()
Why: os.popen takes a string containing a shell command and returns an object that behaves like an open file. The book's footnote notes that popen is deprecated in favour of the subprocess module, and that for simple cases subprocess is more complicated than necessary — so it keeps using popen.
Real world
Not everything needs to be written in Python.
Discussion prompt
Think of a task where an existing command-line tool would do the job better than code you could write. What makes it better?
Hint: The book's example computes a checksum.
Answer:
Checksums, image conversion, compression, searching a huge file — all cases where a specialised tool has had years of work put into being correct and fast.
What makes it better is not cleverness but maturity: the edge cases have been found, and the performance work has been done, by people who only did that.
So a pipe is a way of borrowing that. The cost is a dependency on the tool being installed and a command string that has to be right — which is why checking the exit status matters, since a missing program looks exactly like a failing one.
Section
Section 4
Concept
Any file that contains Python code can be imported as a module. Suppose you have a file named wc.py that defines a function and then calls it.
def linecount(filename):
count = 0
for line in open(filename):
count += 1
return count
print(linecount('wc.py'))| Usage | What happens | Note |
|---|---|---|
| run as a script | prints 7 | it reads itself |
| import wc | prints 7 as well | the test code runs |
| wc.linecount('wc.py') | 7 | the function is available |
If you run this program, it reads itself and prints the number of lines in the file, which is 7. You can also import it — and then you have a module object which provides linecount.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 143-143
Picture it
One as a program to run, one as a library to import.
Figure (svg): Two columns showing the same file run as a script and imported as a module
The middle row is the surprise: normally when you import a module, it defines new functions but it doesn't run them — and this file runs its test code either way.
Worked example
The one line that makes a file work as both.
def linecount(filename):
count = 0
for line in open(filename):
count += 1
return count
if __name__ == '__main__':
print(linecount('wc.py'))| Usage | The value of __name__ | What happens |
|---|---|---|
| run as a script | __name__ is '__main__' | the test code runs |
| imported | __name__ is 'wc' | the test code is skipped |
| either way | linecount is defined | which is the point |
Note the problem.
Why: The only problem with the earlier example is that when you import the module it runs the test code at the bottom.
Note the built-in variable.
Why: __name__ is a built-in variable that is set when the program starts. If the program is running as a script, __name__ has the value '__main__'.
Note the effect.
Why: In that case the test code runs; otherwise, if the module is being imported, the test code is skipped.
Figure (svg): A flowchart showing the __name__ check routing between script and module behaviour
A file that is a usable program and a usable library at once. Programs that will be imported as modules often use this idiom, and it costs one line.
Verify: Check the value of __name__ when imported.
Why: It is the module's name — 'wc' for wc.py — which is what the book sets as an exercise. Knowing that the variable holds something rather than nothing makes the condition read as a comparison rather than a magic incantation.
Prediction
The file is being imported rather than run.
# in wc.py:
print(__name__)
# and elsewhere:
# >>> import wc| Usage | The value | Note |
|---|---|---|
| run as a script | '__main__' | the special value |
| imported | the module's name | 'wc' |
| the check | distinguishes the two |
Predict first
What does it print when wc is imported?
Correct: wc — when a file is imported, __name__ holds the module's name rather than '__main__'.
Why: __name__ is a built-in variable set when the program starts: '__main__' if the file is running as a script, and the module's own name if it is being imported. Nothing is skipped on import — the whole file runs, which is exactly why the idiom is needed.
Worked example
Python does nothing the second time.
>>> import wc
7
>>> import wc # nothing happens
>>>
# even if wc.py has been edited since| Action | What happens | Note |
|---|---|---|
| the first import | reads the file | and runs it |
| the second | does nothing | not even a re-read |
| after an edit | still the old version | until you restart |
Note the behaviour.
Why: If you import a module that has already been imported, Python does nothing — it does not re-read the file, even if it has changed.
Note when it bites.
Why: Editing a module while an interpreter session is open, and finding your changes have no effect.
Note the safe response.
Why: You can use the built-in function reload, but it can be tricky, so the safest thing to do is restart the interpreter and then import the module again.
Figure (svg): The state of the program after each line of Worked example the re-import warning, drawn as a ladder with one rung per traced line
A silent staleness. The symptom is that a fix appears not to work, and the cause is that the fixed file was never read.
Verify: Recognise the symptom before investigating the code.
Why: If a change has no effect at all — not a wrong effect, but none — and the module was imported earlier in the same session, the file has not been re-read. Restarting first costs seconds and rules out a whole category of imaginary bug.
Trap
A file defines useful functions and calls one at the bottom to check it works.
Test the function where it is defined
Why: Which is convenient while writing it.
Importing the file now runs the test — printing output, reading files, and doing work the importer never asked for. Normally when you import a module, it defines new functions but it doesn't run them.
Guard it with the __name__ check.
if __name__ == '__main__':
Why: The test code then runs only when the file is run as a script.
And the definitions still happen either way
Why: Which is what an importer wants.
This is one line, it costs nothing, and it makes every script you write importable. It is worth adopting as a default habit rather than adding when the need arises.
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C describes what people expect and not what happens. Importing runs the whole file top to bottom, including any calls at the bottom — which is precisely the problem the __name__ idiom solves. What is true is the weaker statement that importing does not CALL the functions it defines.
Faded example
Run it as a script, skip it on import.
Fill in the blanks
if __name__ == '__main__':
print(linecount('wc.py'))
Why: __name__ is a built-in variable set when the program starts, holding '__main__' when the file is run as a script and the module's name when it is imported. The guard costs one line and is what makes a file usable as both a program and a library.
Explain it
Not a wrong result — no change whatsoever.
Discussion prompt
A classmate edited a module, ran their code again in the same interpreter session, and the behaviour is byte-for-byte identical. Diagnose it.
Hint: Was the file read again?
Answer:
The module was already imported, so the second import did nothing — Python does not re-read the file, even if it has changed. Their edit is on disk and not in the session.
The signature is that nothing changed at all. A wrong fix produces different wrong behaviour; an unread fix produces exactly what it did before.
The safest response is to restart the interpreter and import again. There is a reload function, but it can be tricky, and restarting rules out the whole question in a couple of seconds.
Section
Section 5
Concept
When you are reading and writing files, you might run into problems with whitespace. These errors can be hard to debug because spaces, tabs and newlines are normally invisible.
>>> s = '1 2\t 3\n 4'
>>> print(s)
1 2 3
4
>>> print(repr(s))
'1 2\t 3\n 4'| Call | What you see | Note |
|---|---|---|
| print(s) | the characters are rendered | the tab and newline act |
| print(repr(s)) | backslash sequences | you can see what is there |
| the difference | rendering against describing |
The built-in function repr takes any object as an argument and returns a string representation of it. For strings, it represents whitespace characters with backslash sequences — which is exactly what you need when the problem is a character you cannot see.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 144-144
Picture it
The same value, rendered and described.
Figure (svg): Two columns showing a string printed normally and printed through repr
Which is why repr is the first thing to reach for when a file's contents look right and behave wrong.
Worked example
Two strings that print identically.
>>> a = 'total'
>>> b = 'total '
>>> print(a)
total
>>> print(b)
total
>>> a == b
False
>>> print(repr(b))
'total '| Step | What you see | Note |
|---|---|---|
| printing both | identical output | the space is invisible |
| comparing | False | and the reason is not visible |
| repr | the trailing space is shown | inside the quotes |
Notice the mystery.
Why: Two strings print the same and compare unequal, which looks like a bug in the comparison.
Reach for repr.
Why: It takes any object and returns a string representation, showing whitespace as it is.
Read the answer off.
Why: The quotes delimit the string, so a trailing space is plainly visible where printing hid it.
Figure (svg): The state of the program after each line of Worked example finding a stray character, drawn as a ladder with one rung per traced line
A trailing space, found in one line. This class of bug is common when reading files, where lines carry their newline and fields may be padded.
Verify: Apply it to a line read from a file.
Why: repr(line) shows the trailing '\n' that every line carries, which explains why a comparison against a plain word fails. That is the commonest whitespace bug in file handling, and strip is the usual fix — but only once you can see that the newline is there.
Prediction
The string contains a tab.
s = 'a\tb'
print(repr(s))| Call | What you see | Note |
|---|---|---|
| print(s) | a, a tab's worth of space, b | the tab acts |
| repr(s) | the escape sequence | shown literally |
| the quotes | delimit the string | so the ends are visible |
Predict first
What does this print?
Correct: 'a\tb' — repr represents whitespace characters with backslash sequences, and includes the quotes.
Why: print(s) would render the tab as horizontal space, hiding what character produced it. repr describes the string instead of rendering it, showing the escape sequence and delimiting the whole with quotes — which is also how a trailing space becomes visible.
Worked example
Different systems disagree about how a line ends.
# some systems use a newline: \n
# others use a return character: \r
# some use both: \r\n
>>> repr(line)
'total\r\n' # two characters, not one| Ending | What it is | Note |
|---|---|---|
| \n | a newline | one convention |
| \r | a return | another |
| \r\n | both | a third |
Note the disagreement.
Why: Different systems use different characters to indicate the end of a line. Some use a newline; others use a return character; some use both.
Note when it causes trouble.
Why: If you move files between different systems, these inconsistencies can cause problems.
Note how to see it.
Why: repr shows exactly which characters are there, which is the only reliable way to tell \n from \r\n.
Figure (svg): Three line-ending conventions shown as the characters they consist of
Three conventions, and a file from one system read on another carries the wrong ones. The symptom is usually a stray character at the end of every field.
Verify: Check what strip does about it.
Why: strip() with no argument removes whitespace including both \r and \n, so it handles all three conventions at once — which is why stripping lines read from a file is such a common habit. For most systems there are also applications to convert from one format to another, or you could write one yourself.
Trap
A program reads lines and tests if line == 'total':, which never matches.
Compare what you read to what you expect
Why: The line looks like the word when printed.
Every line carries its newline, so the value is 'total\n' and the comparison fails. Printing it shows total, which makes the comparison look correct and the bug invisible.
Strip the line, and use repr when unsure.
if line.strip() == 'total':
Why: Which removes the newline and any stray spaces.
And print(repr(line)) when a comparison fails inexplicably
Why: It shows exactly what is there.
The general rule for file work: whenever two things look the same and compare unequal, look at them through repr before looking at the comparison. The characters you cannot see are the usual explanation.
Error analysis
Mark each and say what repr would reveal.
Annotate
All four are invisible when printed, which is exactly why the book introduces repr in a chapter about files.
Faded example
Describe the string rather than rendering it.
Fill in the blanks
print(repr(line)) # shows \t and \n literally
Why: repr takes any object and returns a string representation of it, and for strings it represents whitespace characters with backslash sequences. Printing the string directly renders those characters instead, which is exactly what hides them — so repr is the tool when the problem is something you cannot see.
Real world
Two things that look identical and are not.
Discussion prompt
Think of a case outside programming where two entries looked the same and were treated as different. What was invisible?
Hint: A space you cannot see.
Answer:
A trailing space in a spreadsheet cell that splits one category into two; a name pasted with a non-breaking space; two apparently identical codes that will not match.
What is invisible is a character that takes up no distinguishable room — and every system involved renders it rather than describing it, so nobody can see the difference.
Which is what repr is for: a view that describes rather than renders. The general lesson is that when two things look identical and behave differently, the display is hiding something, and you need a different view rather than a closer look.
Comparison
Fill the blanks. Each keeps a different amount of the structure.
Comparison matrix
| Question | A text file | A dbm database | Pickle |
|---|---|---|---|
| Readable by a person? | yes | partly — the values are text | no |
| Keeps types? | no — everything becomes characters | no — strings or bytes only | yes, almost any object |
| Random access by key? | no | yes — like a dictionary | not by itself |
| When to use it | flat data someone will read | key-value data that must persist | program state, reloaded by itself |
And the combination of the last two is so common it has its own module, shelve.
Pattern
Five steps, and the fourth is the one worth making a habit.
Step 5 is the one that wastes the most time when forgotten. A fix that has no effect at all — rather than a different wrong effect — almost always means the file was never re-read.
Python documentation — pickle — Python object serialization pickle — Python object serialization
Check
Out to a string and back again.
t1 = [1, 2, 3]
t2 = pickle.loads(pickle.dumps(t1))
print(t1 is t2)| Step | What happens | Result |
|---|---|---|
| dumps | the contents become a string | the object does not travel |
| loads | builds something new | from the description |
| is | two objects | False |
Check your understanding
What does this print?
Answer: A
Why: Although the new object has the same value as the old, it is not in general the same object — so == is True and is is False. The book's summary is exact: pickling and then unpickling has the same effect as copying the object.
Check
The file is imported rather than run.
Check your understanding
What does if __name__ == '__main__': achieve?
Answer: A
Why: __name__ is a built-in variable set when the program starts: '__main__' when the file runs as a script, and the module's name when it is imported. So the guarded test code runs only in the first case — which is what stops an import doing work the importer never asked for.
Check
Two strings that print identically.
Check your understanding
Why does the book introduce repr in a chapter about files?
Answer: A
Why: repr takes any object and returns a string representation, showing whitespace characters as backslash sequences. That matters most for file work, where every line carries a line ending and different systems use different characters for it — differences that are invisible when printed.
Real world
Saving something in a form you can get back is a design decision with a cost.
Discussion prompt
Think of a file format you can open and read, and one you cannot. What does each buy, and what happens when the program that wrote the unreadable one is gone?
Hint: Plain text against a proprietary format.
Answer:
A readable format can be inspected, fixed by hand, read by other tools, and understood in ten years. It usually stores less structure and takes more space.
An opaque format keeps everything the program knew and is useless without that program. If it stops working or nobody has it any more, the data are effectively gone.
Which is exactly the text-file-against-pickle trade. Neither is right in general, and the question worth asking is who will need to read this, and when — because that is what decides which cost you can afford.
Commit first
Answer, then rate your confidence.
Predict first
You import a module, edit its file, and import it again in the same session. What happens?
Correct: Nothing — if you import a module that has already been imported, Python does nothing, and it does not re-read the file even if it has changed.
Why: The book gives this as an explicit warning, and it costs people real time because the symptom is so misleading: your fix appears to have no effect at all. That is the tell — a wrong fix produces different wrong behaviour, whereas an unread fix produces byte-for-byte what it did before. There is a built-in reload function, but the book notes it can be tricky, so the safest thing to do is restart the interpreter and import the module again. Making that the reflex whenever a change seems to do nothing rules out a whole category of imaginary bug in a couple of seconds.
Explain it
Three ways to save data, and how to choose.
Discussion prompt
A classmate asks whether to save their program's data as a text file, in a dbm database, or with pickle. Give them the question that decides it.
Hint: Who reads it back?
Answer:
Ask who will read it. If a person or another program must, it has to be a text file — which means the data have to be flat enough to write as lines, and the types will be lost.
If only their own program will read it back and the structure matters, pickle: it translates almost any type of object into a string and back, and the format is not meant for human readers.
And if they want to look things up by key without loading everything, a dbm database — which behaves like a dictionary on disk, with the limitation that keys and values must be strings or bytes. Combining it with pickle is so common that shelve exists to do exactly that.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The database is the easiest of the four, because almost everything transfers from dictionaries — the bytes and the missing methods are the only surprises. Pickling is one idea with a memorable summary: it has the same effect as copying. Pipes reuse the file interface entirely, and the one thing worth remembering is that None means success. And the __name__ idiom has the longest reach of anything in this lesson: it is one line, and it makes every script you write importable.
Connect it up
One page, from memory.
Draw it
Draw three boxes labelled text file, database and pickle, and beside each write what it preserves and what it loses. Underneath, write the round trip pickle.loads(pickle.dumps(t)) and note what == and is report about the result. Then write the __name__ guard from memory, with the value the variable holds in each of the two cases, and finish with one string shown both ways — printed, and through repr.
Recap
Four pages, and chapter 14 is finished: data that outlive the program, several ways.
| If you remember one thing | It is this |
|---|---|
| From dbm | A database is a dictionary on disk, with strings or bytes for keys and values. |
| From pickle | Pickling and unpickling has the same effect as copying the object. |
| From pipes | close returns the status, and None means it ended normally. |
| From modules | Importing runs the whole file. The __name__ guard is what stops that mattering. |
| From repr | print renders; repr describes. When something looks right and behaves wrong, describe it. |
The next chapter begins the object-oriented half of the book: programmer-defined types, with a class that represents a point in two-dimensional space — attributes, instances, and objects as return values.
Think Python, 2nd edition — Allen B. Downey §14.6-14.10, pp. 141-144 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.