This lesson introduces persistent programs, shows how to open a file for writing and what that destroys, and covers the format operator for putting non-string values into a file.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 14 — Files
§14.1-14.3, pp. 137-138
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138 — the pages these objectives are drawn from
Warm-up
You built one in chapter 13. Then the program ended.
Discussion prompt
Analysing Emma takes a noticeable time and produces a histogram of 7,214 words. Run the program twice and it does the work twice. What would you need to avoid that, and what is stopping you?
Hint: Where does the histogram live?
Answer:
You would need the histogram to survive after the program ends — written somewhere that is still there next time.
What stops you is that everything so far has lived in memory, which is gone the moment the program stops. Every program you have written starts with a clean slate.
This chapter is about the alternative: keeping data in permanent storage, so that a program can pick up where it left off.
Concept
Most of the programs we have seen so far are transient: they run for a short time and produce some output, but when they end, their data disappears. If you run the program again, it starts with a clean slate.
persistent — Pertaining to a program that runs indefinitely and keeps at least some of its data in permanent storage.
Other programs are persistent: they run for a long time or all the time, they keep at least some of their data in permanent storage, and if they shut down and restart, they pick up where they left off. Operating systems and web servers are the obvious examples.
Figure (svg): Two columns contrasting a transient program with a persistent one
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-137
Section
Section 1
Concept
To write a file, you have to open it with mode 'w' as a second parameter. If the file already exists, opening it in write mode clears out the old data and starts fresh — so be careful. If the file doesn't exist, a new one is created.
>>> fout = open('output.txt', 'w')| Situation | What happens | Note |
|---|---|---|
| the file exists | its contents are cleared | immediately, on open |
| the file does not exist | a new one is created | empty |
| either way | you get a file object | with methods for working with it |
This is the first operation in the book that can destroy something outside the program, and it does so at the moment of opening — before you have written a single character, and without asking.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138
Picture it
The second argument decides whether the old contents survive.
Figure (svg): Two columns comparing opening an existing file for reading and for writing
The clearing happens on open, not on write — so a program that opens a file for writing and then crashes has still emptied it.
Worked example
open, write, close — and the return value is worth reading.
>>> fout = open('output.txt', 'w')
>>> line1 = "This here's the wattle,\n"
>>> fout.write(line1)
24
>>> line2 = 'the emblem of our land.\n'
>>> fout.write(line2)
24
>>> fout.close()| Line | What happens | Note |
|---|---|---|
| open(..., 'w') | a file object | the file is now empty |
| write(line1) | returns 24 | the number of characters written |
| write(line2) | adds to the END | the file object tracks its position |
| close() | finishes | or it happens when the program ends |
Open in write mode.
Why: open returns a file object that provides methods for working with the file.
Write, and note the return value.
Why: The write method puts data into the file, and the return value is the number of characters that were written.
Write again.
Why: The file object keeps track of where it is, so if you call write again, it adds the new data to the end.
Figure (svg): The state of the program after each line of Worked example writing two lines, drawn as a ladder with one rung per traced line
Two lines in the file, and two counts of 24. When you are done writing, you should close the file — and if you don't, it gets closed for you when the program ends.
Verify: Count the characters yourself.
Why: "This here's the wattle," is 23 characters and the newline makes 24, which matches what write returned. Confirming the count includes the newline is worth doing once, because a missing newline is the commonest reason a written file looks wrong when it is read back.
Prediction
The file already contains a hundred lines.
fout = open('notes.txt', 'w')
# nothing is written yet
fout.close()| Step | What happens | Result |
|---|---|---|
| open with 'w' | clears the file | on opening |
| nothing written | the file stays empty | 0 characters |
| close | an empty file remains | the hundred lines are gone |
Predict first
What does notes.txt contain afterwards?
Correct: Nothing — if the file already exists, opening it in write mode clears out the old data and starts fresh.
Why: The clearing happens at open, before any write. That is what makes this trap so effective: a program can destroy a file without executing a single write, and without any warning. If you want to add to an existing file rather than replace it, mode 'a' appends instead.
Worked example
It happens automatically, and relying on that has a cost.
# explicit
fout = open('out.txt', 'w')
fout.write('data\n')
fout.close()
# or let it happen at the end of the program
# - but nothing else can rely on the file
# being complete until then| Approach | When the file is finished | Note |
|---|---|---|
| close() | finishes the file | at a point you choose |
| not closing | closed when the program ends | at a point you do not |
| in between | the file may be incomplete | for anything else reading it |
Note the rule.
Why: When you are done writing, you should close the file. If you don't close it, it gets closed for you when the program ends.
Note why should rather than must.
Why: The automatic close means a short script will work either way, which is why the omission so often goes unnoticed.
Note when it bites.
Why: A long-running program, or one that writes a file and then reads it back, cannot rely on the contents until the file is closed.
Figure (svg): A pipeline from opening a file through writing to closing it
Closing is not strictly required and is still worth doing, because it is the point at which the file is definitely complete — and that point is otherwise the end of the program.
Verify: Try reading a file back before closing it.
Why: The read may come back empty or short, because what was written is still buffered rather than on the disk. That failure looks like a bug in the writing code and is really a missing close, which is a good reason to make the habit automatic.
Trap
A program opens a file in write mode to see whether it is there, planning to read it if so.
Open it and find out
Why: Which works for reading, where a missing file raises.
Write mode creates the file if it is absent and empties it if it is present — so the check destroys exactly the data it was looking for, and reports success either way.
Test before opening, or open for reading.
os.path.exists tells you without touching anything
Why: Which the next lesson covers.
Or open for reading and handle the failure
Why: A missing file raises, and the file is unharmed.
The rule to internalise: mode 'w' is destructive at the moment of opening. Nothing warns you, nothing asks, and the old contents are not recoverable.
Prediction
The return value is easy to overlook.
fout = open('out.txt', 'w')
n = fout.write('hello\n')
print(n)| Part | What is true | Result |
|---|---|---|
| 'hello' | five characters | plus the newline |
| write | returns the count | 6 |
| the file | now holds the line | and is not yet closed |
Predict first
What does this print?
Correct: 6 — the return value is the number of characters that were written, and the newline counts.
Why: write returns a character count rather than None, which is unusual among the methods you have met — most modifying methods return nothing. The newline is a single character, so 'hello\n' is six. Forgetting that the newline counts is the commonest reason people expect 5.
Discrimination
Ask what the mode is.
Sort into buckets
For each call, is an existing file's content lost?
Socratic
It would be safer to clear on the first write.
Discussion prompt
Opening in write mode empties the file immediately, even if nothing is ever written. Why might the language do that rather than waiting?
Hint: What is the mode a statement about?
Answer:
Because the mode is a declaration of intent: you have said this file is going to be replaced, so the old contents are not wanted.
Waiting would mean the file's state depended on whether a write happened to occur, which would make the behaviour harder to predict rather than easier — sometimes cleared, sometimes not.
The practical consequence is the one to remember: opening is the destructive step. If you might not want to replace the file, do not open it in write mode at all — check first, or use append mode.
Section
Section 2
Concept
The argument of write has to be a string, so if we want to put other values in a file, we have to convert them. When applied to integers, % is the modulus operator — but when the first operand is a string, % is the format operator.
>>> 7 % 3
1
>>> camels = 42
>>> '%d' % camels
'42'
>>> 'I have spotted %d camels.' % camels
'I have spotted 42 camels.'| Expression | Which operator | Result |
|---|---|---|
| 7 % 3 | both integers | modulus: 1 |
| '%d' % camels | first operand a string | formatting: '42' |
| embedded in a sentence | the sequence can be anywhere | a longer string |
The first operand is the format string, which contains one or more format sequences specifying how the second operand is formatted. The result is a string — so '42' is not to be confused with the integer 42.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-138
Picture it
Which one applies is decided by the type on the left.
Figure (svg): Two columns showing the percent symbol acting as modulus and as formatting
That is unusual — most operators mean one thing — and it is why the book introduces the distinction explicitly before using it.
Worked example
Several values need a tuple.
>>> 'In %d years I have spotted %g %s.' % (3, 0.1, 'camels')
'In 3 years I have spotted 0.1 camels.'| Sequence | What it formats | From the tuple |
|---|---|---|
| %d | a decimal integer | 3 |
| %g | a floating-point number | 0.1 |
| %s | a string | 'camels' |
| the tuple | matched in order | one element per sequence |
Note when a tuple is needed.
Why: If there is more than one format sequence in the string, the second argument has to be a tuple.
Note how they are matched.
Why: Each format sequence is matched with an element of the tuple, in order — positionally, exactly like tuple assignment.
Note the three common sequences.
Why: %d for a decimal integer, %g for a floating-point number, and %s for a string.
Figure (svg): The state of the program after each line of Worked example three format sequences, drawn as a ladder with one rung per traced line
One string built from three values of three different types. The tuple's order must match the sequences' order, since the matching is positional.
Verify: Swap two values in the tuple and see which fails.
Why: Swapping the 3 and the 0.1 gives '%d' a float — which Python accepts by truncating — while swapping in the string raises. So some mismatches are loud and some are quiet, which is an argument for reading the sequence order carefully rather than relying on an error.
Prediction
The first operand is a string.
print('%d' % 7)| Part | What is true | Result |
|---|---|---|
| the first operand | a string | so % formats |
| %d | a decimal integer | 7 |
| the result | a string | '7' |
Predict first
What does this print?
Correct: 7 — the first operand is a string, so % is the format operator and the result is the string '7'.
Why: When applied to integers % is the modulus operator, but when the first operand is a string it formats. Option B would be the modulus of some pair of numbers; nothing here is arithmetic. Note that the result is the string '7', not the integer — which matters if it is used in a calculation afterwards.
Worked example
Too few elements, or the wrong type.
>>> '%d %d %d' % (1, 2)
TypeError: not enough arguments for format string
>>> '%d' % 'dollars'
TypeError: %d format: a number is required, not str| Mistake | What is wrong | Message |
|---|---|---|
| three sequences, two values | a count mismatch | not enough arguments |
| %d given a string | a type mismatch | a number is required |
| both | TypeError | and both messages are specific |
Check the count.
Why: The number of elements in the tuple has to match the number of format sequences in the string.
Check the types.
Why: The types of the elements have to match the format sequences — %d wants a number.
Read the two messages.
Why: In the first example there aren't enough elements; in the second, the element is the wrong type. Each message says which.
Figure (svg): A panel showing the three distinct TypeError messages the format operator produces
Two TypeErrors with two distinct messages. Both failures are loud, which is a relief after this chapter's silent file destruction.
Verify: Try too many elements rather than too few.
Why: '%d' % (1, 2) raises not all arguments converted during string formatting — a third distinct message. So all three count-and-type mismatches are caught, and the message tells you which one you have without further investigation.
Trap
A program writes '%s and %s' % 'a', 'b' — with the two values as separate arguments.
Pass the values the way you would to a function
Why: Comma-separated arguments are how everything else takes several values.
The format operator takes exactly one right-hand operand, so this formats 'a' alone and then makes a tuple of the result with 'b'. It raises about missing arguments — or, in a print call, silently produces something odd.
Put the values in a tuple.
'%s and %s' % ('a', 'b')
Why: One operand, which happens to hold two values.
A single value needs no tuple
Why: '%d' % 42 works, though '%d' % (42,) also does.
The asymmetry is worth noting: one value may be passed bare, and several must be a tuple. The single-element tuple form works in both cases, which is why some people always write it.
Sorting
Look at the type on the left.
Sort into buckets
For each expression, which operation does % perform?
Faded example
Several sequences need one tuple.
Fill in the blanks
print('%d camels and %s' % (3, 'a llama'))
Why: If there is more than one format sequence in the string, the second argument has to be a tuple, and each sequence is matched with an element in order. Passing them as separate comma-separated arguments would not work: the operator takes exactly one right-hand operand.
Error analysis
Mark each and name the failure.
Annotate
Every mismatch is caught, which makes the format operator one of the safer things in this chapter.
Section
Section 3
Concept
The argument of write has to be a string, so if we want to put other values in a file, we have to convert them to strings. The easiest way to do that is with str.
>>> x = 52
>>> fout.write(str(x))
# or with the format operator, which can do more:
>>> fout.write('%d\n' % x)| Call | What is passed | Result |
|---|---|---|
| fout.write(x) | an integer | TypeError |
| fout.write(str(x)) | converted | works |
| fout.write('%d\n' % x) | converted and formatted | with a newline |
str is the simplest conversion; the format operator does the same job and can add surrounding text, control the appearance of a number, and append the newline in one expression.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-138
Picture it
Both end in a string, which is the only thing write accepts.
Figure (svg): Two paths converting a number into a string suitable for writing
That is worth stating plainly: a text file has no notion of an integer. Everything in it is text, and reading it back means converting in the other direction.
Worked example
Two values per line, formatted and terminated.
fout = open('counts.txt', 'w')
for word, freq in hist.items():
fout.write('%s %d\n' % (word, freq))
fout.close()| Part | What it produces | Note |
|---|---|---|
| '%s %d\n' | a word, a space, a number, a newline | one line's worth |
| % (word, freq) | the two values, in order | matched positionally |
| write | one line per item | the position advances |
Build the line as a string.
Why: One format expression produces the whole line, including the separator and the newline.
Match the tuple to the sequences.
Why: %s takes the word and %d the count, in the order they appear.
Write and close.
Why: Each write appends at the current position, so the lines accumulate in order.
Figure (svg): A ladder showing successive writes appending lines to a file
A file with one word and count per line — the histogram made persistent, so a later run need not recompute it.
Verify: Check the newline is present.
Why: Without the \n every line would run into the next and the file would be one enormous line — recoverable, but only by knowing where the boundaries were. Since write does not add a newline of its own, supplying it is the writer's job, and forgetting it is the commonest defect in a written file.
Prediction
The argument is an integer.
fout = open('out.txt', 'w')
fout.write(52)| Part | What is true | Result |
|---|---|---|
| write | requires a string | strictly |
| 52 | an integer | not converted automatically |
| the result | TypeError | must be str |
Predict first
What happens?
Correct: A TypeError — the argument of write has to be a string, so the integer must be converted first.
Why: str(52) or '%d' % 52 both produce the string write wants. print would have accepted the integer and converted it, which is why the inconsistency catches people — print converts and terminates lines for you, and write does neither.
Worked example
Everything written becomes text, and the types are gone.
fout.write('%d\n' % 42) # writes the characters '4', '2'
# reading it back:
fin = open('out.txt')
line = fin.readline() # '42\n' - a STRING
n = int(line) # convert back explicitly| Direction | What happens | Note |
|---|---|---|
| writing | 42 becomes '42' | the type is discarded |
| reading | '42\n' | a string, with the newline |
| int(line) | back to a number | the conversion is yours to do |
Note what is stored.
Why: A text file is a sequence of characters, so the integer becomes the characters '4' and '2' and nothing records that it was a number.
Note what comes back.
Why: Reading gives a string, including the newline — the file cannot tell you it was an integer.
Convert explicitly.
Why: int() turns it back, and it tolerates the trailing newline, though strip makes the intent clearer.
Figure (svg): The state of the program after each line of Worked example what the file does not remember, drawn as a ladder with one rung per traced line
Round-tripping a number through a text file loses its type, and restoring it is the reader's responsibility. That is the limitation the pickle module addresses later in the chapter.
Verify: Try it with a list.
Why: str([1, 2, 3]) writes '[1, 2, 3]' and reading it back gives that string, with no easy way to recover the list — int and float have inverses, and a general value does not. That gap is exactly why the chapter goes on to databases and pickling.
Trap
A program writes fout.write(count) where count is an integer.
Pass the value you want in the file
Why: print accepts anything, so write looks like it should too.
It raises TypeError: write() argument must be str. print converts for you and write does not, which is an inconsistency between two things that both put text somewhere.
Convert first.
str(count), or a format expression
Why: The argument of write has to be a string.
And add the newline yourself
Why: write does not add one, whereas print does by default.
Both differences run the same way: print is the convenient one that converts and terminates, and write is the literal one that does exactly what it is told. Remembering that print is the exception makes both predictable.
Comparison
Fill the blanks. Both put text somewhere and they differ in three ways.
Comparison matrix
| Question | write | |
|---|---|---|
| Where does it go? | the screen | a file |
| Does it convert values? | yes — anything can be printed | no — the argument must be a string |
| Does it add a newline? | yes, by default | no — you supply it |
| What does it return? | None | the number of characters written |
print is the convenient one and write is the literal one. Every difference runs that way, which makes them easy to keep straight.
Faded example
Two conversions and a line ending, in one expression.
Fill in the blanks
fout.write('%d\n' % count)
Why: write does not add a line ending, so it has to be part of the string. Without it every value would run into the next and the file would be a single line — readable only by someone who knew where the boundaries had been.
Explain it to yourself
An integer goes in and a string comes out.
Discussion prompt
Explain why writing 42 to a text file and reading it back gives you a string rather than a number.
Hint: What is a text file made of?
Answer:
A text file is a sequence of characters, and nothing else. Writing 42 stores the characters '4' and '2'; there is nowhere in the file to record that they were meant as a number.
So reading gives characters back, and turning them into an integer is a decision the reading program makes — the file cannot make it, because the information was never stored.
Which is why the chapter goes on to pickling: a format that does record types can restore the original value, where a text file can only give you text. For numbers the conversion is easy either way; for a list or a dictionary it is not.
Section
Section 4
Concept
Three format sequences cover almost everything: a decimal integer, a floating-point number, and a string.
>>> '%d' % 42
'42'
>>> '%g' % 0.1
'0.1'
>>> '%s' % 'camels'
'camels'
>>> '%s' % 42
'42' # %s accepts anything| Sequence | What it formats | Note |
|---|---|---|
| %d | a decimal integer | requires a number |
| %g | a floating-point number | chooses a readable form |
| %s | a string | accepts any value |
%s is the forgiving one, because any value can be converted to a string. The other two are specific, and giving them the wrong type raises rather than guessing.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-139
Picture it
Two are strict and one takes anything.
Figure (svg): Three format sequences with the types each accepts
Using %s everywhere always works and gives up the checking — which is a fair trade when you are only writing text, and a loss when a type mistake would be worth catching.
Worked example
The strict sequence catches a mistake the forgiving one hides.
>>> count = 'twelve' # a bug: should be a number
>>> '%d items' % count
TypeError: %d format: a number is required, not str
>>> '%s items' % count
'twelve items' # legal, and probably wrong| Sequence | What happens | Note |
|---|---|---|
| %d with a string | raises immediately | the bug is found |
| %s with a string | formats happily | the bug survives |
| the difference | checking against convenience | a real trade |
Notice what %d refuses.
Why: The types of the elements have to match the format sequences, so a string where a number belongs raises.
Notice what %s accepts.
Why: Anything, because any value can be converted to a string — which means it can never tell you something is wrong.
Weigh the two.
Why: The strict sequence turns a type confusion into an immediate error; the forgiving one lets it through into the output.
Figure (svg): The state of the program after each line of Worked example why d is not just s, drawn as a ladder with one rung per traced line
%d is a small type check written into the format string. Using %s everywhere is simpler and gives that up.
Verify: Ask when %s is the right choice anyway.
Why: When the value genuinely may be of several types, or when it is already a string. Choosing %s deliberately is fine; choosing it to avoid thinking about types is what loses the check. That distinction — a decision against a default — is the same one behind get versus bracket lookup in lesson 11a.
Prediction
The format operator produces something specific.
x = '%d' % 42
print(type(x))| Part | What is true | Result |
|---|---|---|
| the format operator | the result is a string | always |
| '42' | characters, not a value | not the integer |
| type | <class 'str'> |
Predict first
What does this print?
Correct: <class 'str'> — the result of the format operator is a string.
Why: The book flags this directly: the result is the string '42', which is not to be confused with the integer value 42. Formatting is for producing output, so anything arithmetic should happen before it — otherwise the plus operator concatenates instead of adding.
Worked example
It picks a readable form for a floating-point number.
>>> '%g' % 0.1
'0.1'
>>> '%g' % 100000000.0
'1e+08'
>>> '%f' % 0.1
'0.100000' # always six decimal places| Sequence | What it chooses | Result |
|---|---|---|
| %g on a small number | plain form | 0.1 |
| %g on a large one | scientific notation | 1e+08 |
| %f | a fixed six places | regardless of the value |
Use %g for a general float.
Why: It formats a floating-point number in whichever of the two forms is more compact, which is usually what you want for output.
Note it switches to scientific notation.
Why: For very large or very small numbers, which keeps the string short and can surprise you if you expected digits.
Compare with %f.
Why: Always a fixed number of decimal places, which is better when columns must line up and worse when the magnitudes vary.
Figure (svg): Two columns comparing the general float format with the fixed-decimal one
%g adapts and %f does not. Neither is right in general; the choice depends on whether you want compactness or alignment.
Verify: Format a whole number as a float.
Why: '%g' % 3.0 gives '3' rather than '3.0', which is more compact and loses the visible indication that the value was a float. If that distinction matters in the output, %f or an explicit precision is the better choice — which is the kind of thing worth checking once rather than discovering in a report.
Trap
A program computes total = '%d' % count + 1, expecting to add one to the count.
Treat the formatted value as the number
Why: It looks like a number and prints like one.
The result of the format operator is a string, so this concatenates rather than adds — or raises, if the 1 is an integer. The book warns explicitly that '42' is not to be confused with the integer 42.
Format at the last moment, for output only.
Do the arithmetic on numbers
Why: total = count + 1, and format when writing.
Then '%d' % total for the file
Why: One conversion, at the boundary.
The general shape is worth keeping: values stay in their own types inside the program and become text only where they leave it. Converting early means every later operation deals with strings that are pretending to be numbers.
Discrimination
Two are strict about type and one is not.
Sort into buckets
For each value, which sequence fits best?
Faded example
A whole number, and a sequence that checks it is one.
Fill in the blanks
print('I have spotted %d camels.' % 42)
Why: %d formats a decimal integer and refuses a non-number, so it catches a value that should have been a count and is not. %s would also work here and would accept anything at all, giving up that check — which is fine when chosen deliberately and a loss when chosen by default.
Explain it
A very common early bug.
Discussion prompt
A classmate formatted a number, added one, and got '421' instead of 43. Explain what happened.
Hint: What did the format operator return?
Answer:
The format operator returns a string, so '%d' % 42 is the two characters '4' and '2', not the number forty-two.
Adding 1 to a string with the plus operator concatenates rather than adds — or raises, if the 1 was an integer. Either way the value stopped being a number the moment it was formatted.
The fix is to keep values in their own types inside the program and format only where they leave it, on the way to a file or the screen. Formatting early means everything after it is text pretending to be a number.
Section
Section 5
Concept
One of the simplest ways for programs to maintain their data is by reading and writing text files. An alternative is to store the state of the program in a database.
We have already seen programs that read text files; in this chapter we will see programs that write them. The choice between the routes is another data structure decision, of exactly the kind chapter 13 was about.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-137
Picture it
Persistence is about which part of the state survives the ending.
Figure (svg): A flowchart contrasting a transient program's ending with a persistent one's restart
Which is why the chapter is short on new concepts and long on mechanics: the idea is simple and the details of doing it safely are not.
Worked example
Compute once, save, and load thereafter.
import os
if os.path.exists('counts.txt'):
hist = load_counts('counts.txt') # fast
else:
hist = process_file('emma.txt') # slow
save_counts(hist, 'counts.txt')| Situation | What happens | Note |
|---|---|---|
| the file exists | load it | seconds saved |
| it does not | compute and save | once |
| subsequent runs | always the fast path | persistence |
Check whether the saved data exist.
Why: The existence check comes before any open, so nothing is destroyed by asking.
Load if they do, compute if they do not.
Why: The expensive work happens once, and every later run reads the result.
Save after computing.
Why: So that the next run takes the fast path.
Figure (svg): Two columns comparing recomputing every run with loading saved results
A program that pays the analysis cost once. This is the memo pattern from chapter 11, with a file in place of a dictionary — and the file survives the program's ending, which the dictionary could not.
Verify: Ask what makes the saved data stale.
Why: A change to emma.txt, or to the cleaning rules. The saved histogram records an answer without recording the question, so nothing detects that it no longer matches — which is the general hazard of caching that lesson 11c raised, and Fibonacci avoided by having answers that never change.
Prediction
The program writes nothing to storage.
hist = process_file('emma.txt')
print(len(hist))
# the program ends| Stage | What happens | Note |
|---|---|---|
| the histogram | in memory | while it runs |
| the program ends | memory is released | the data are gone |
| running again | starts with a clean slate | recomputes everything |
Predict first
Is this program transient or persistent?
Correct: Transient — it runs for a short time and produces some output, but when it ends, its data disappear.
Why: Reading a file does not make a program persistent; keeping its own data in permanent storage does. This one starts with a clean slate every run and recomputes the whole histogram, which is exactly the situation writing a file would fix.
Worked example
A text file is not always the right container.
# fine in a text file: one word and count per line
fout.write('%s %d\n' % (word, freq))
# awkward: a dictionary of lists of tuples
# - what separator? what escaping?
# - this is what pickle is for| Data | Route | Note |
|---|---|---|
| flat, simple values | a text file is fine | one line each |
| nested structures | awkward to encode | separators collide |
| arbitrary program data | pickle | types preserved |
Consider what the data look like.
Why: A flat mapping from words to counts is one line each and reads back easily.
Consider what a nested structure needs.
Why: A dictionary of lists of tuples has to be encoded with separators, and any separator can appear in the data — lesson 12c's compound-key hazard again.
Note where each route ends.
Why: A text file is simple, readable and lossy; pickle makes it easy to store program data as it is.
Figure (svg): The state of the program after each line of Worked example which route for which data, drawn as a ladder with one rung per traced line
The shape of the data decides the route. A text file is right for something flat and human-readable, and wrong for an arbitrary structure.
Verify: Ask what a text file buys that pickle does not.
Why: Readability by anything at all — you can open it, another program can read it, and it survives a change of language. That is a real advantage, which is why the format persists for data simple enough to fit it.
Trap
A program opens its output file at the top, then reads its input, and the two happen to be the same file.
Set everything up first
Why: Opening all the files at the start looks tidy.
Opening for writing clears the file, so the input is destroyed before it is read — and the program reports an empty input rather than an error.
Open the output only when you are ready to write it.
Read the input fully first
Why: Then nothing that follows can destroy it.
And never let the two names be the same
Why: Write to a different name, and rename afterwards if you must replace it.
This is the most destructive mistake in the chapter and it is entirely silent. The file is gone, the program succeeds, and the output is empty — which looks like a bug in the analysis rather than in the file handling.
Sorting
Ask what happens when it restarts.
Sort into buckets
For each program, which kind is it?
Faded example
Ask whether the saved data are there.
Fill in the blanks
import os
if os.path.exists('counts.txt'):
hist = load_counts('counts.txt')
else:
hist = process_file('emma.txt')
Why: os.path.exists checks whether a file or directory exists without opening it, so the check itself cannot destroy anything. Trying to find out by opening in write mode would create the file if it were absent and empty it if it were present — the destructive answer to a harmless question.
Real world
Most software you use is in the persistent column.
Discussion prompt
Think of something you use that would be useless if it forgot everything when it closed. What does it store, and where would it hurt most to lose it?
Hint: Anything with your work in it.
Answer:
A document editor, a messaging app, a game with saved progress, a browser with its history — all of them keep state on disk precisely so that closing is not losing.
What they store is the part you would mind recreating: the text, the conversation, the position, the settings. Anything cheap to recompute usually is not stored.
Which is the design question this chapter poses: what is expensive enough to save, and what is safer to recompute? Chapter 13's histogram is a good candidate because it takes real time; the two totals derived from it are not, because they are one line each.
Comparison
Fill the blanks. One mode is safe and the other is not.
Comparison matrix
| Question | open(name) | open(name, 'w') |
|---|---|---|
| If the file exists | read it, unchanged | clear it immediately |
| If it does not | FileNotFoundError | create a new one |
| What the methods do | read and iterate lines | write, returning a character count |
| Can it destroy data? | no | yes, on opening, silently |
The bottom row is the one to keep. Opening for writing is destructive before anything is written.
Pattern
Six steps, and the second exists because the first is destructive.
Step 1 is the one that prevents the chapter's worst mistake. Opening the input file for writing destroys it before a single line has been read, and the program then reports an empty input rather than an error.
Python documentation — Input and Output Input and Output
Check
The file already has contents.
Check your understanding
What happens when you open an existing file with mode 'w'?
Answer: A
Why: If the file already exists, opening it in write mode clears out the old data and starts fresh — at the moment of opening, before anything is written. That is why a program that opens a file for writing and then crashes has still emptied it.
Check
Two format sequences.
print('%s has %d legs' % ('a cat', 4))| Sequence | What it formats | Value |
|---|---|---|
| %s | a string | 'a cat' |
| %d | a decimal integer | 4 |
| the tuple | matched in order | positionally |
Check your understanding
What does this print?
Answer: A
Why: Each format sequence is matched with an element of the tuple, in order. Two sequences and two elements, with the types matching, so the result is the completed sentence. Passing the values as separate arguments rather than a tuple would fail, since the operator takes exactly one right-hand operand.
Check
One of these fails.
Check your understanding
Which call raises a TypeError?
Answer: A
Why: The argument of write has to be a string, so an integer must be converted first — with str, or with a format expression. print would have accepted the integer and converted it for you, which is the inconsistency that catches people.
Real world
Replacing a file rather than adding to it is a distinction with consequences.
Discussion prompt
Think of a time something was overwritten when it should have been added to, or the reverse. What made the difference invisible until afterwards?
Hint: Save as, against save.
Answer:
Saving over the original instead of a copy; a backup that replaced the previous one rather than joining it; an export that silently overwrote last month's.
What makes it invisible is that both operations succeed and look identical from outside — a file exists afterwards either way, and only its contents differ.
Which is exactly mode 'w' against mode 'a'. One character in the call decides whether the previous contents survive, nothing warns you, and the result is not recoverable. It is worth treating the write-mode open as the dangerous line it is.
Commit first
Answer, then rate your confidence. This one can cost real data.
Predict first
You call open('notes.txt', 'w') on a file containing a hundred lines, then close it without writing. What does the file contain?
Correct: Nothing — if the file already exists, opening it in write mode clears out the old data and starts fresh.
Why: The clearing happens at the moment of opening, not at the first write, which is what makes this the most dangerous line in the chapter. A program can destroy a file without executing a single write and without any warning — and the destruction is not recoverable. The book's own comment is simply so be careful. The practical consequences are worth stating: never open a file in write mode to find out whether it exists (os.path.exists answers that without touching it), never open the output file before the input has been fully read, and use mode 'a' when you mean to add rather than replace.
Explain it
print and write both put text somewhere, and they differ in three ways.
Discussion prompt
A classmate's file is one enormous line and their write call raised a TypeError before that. Explain both, in terms of how write differs from print.
Hint: print does two favours that write does not.
Answer:
The TypeError is because write takes only a string — print converts any value for you and write does not, so a number has to be passed through str or a format expression.
The single enormous line is because print adds a newline by default and write does not. Every line ending in a written file has to be supplied as part of the string.
Both differences run the same way: print is the convenient one and write is the literal one that does exactly what it is told. Remembering that print is the exception makes both predictable rather than surprising.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The persistence distinction is simple and worth being precise about, since reading a file does not make a program persistent. The write-mode warning is the one thing in this lesson that can cost you data, and it is worth over-learning: the clearing happens on open. The two meanings of % are unusual in the language and settle quickly once you know to look at the left operand. And the tuple matching is ordinary positional matching with three clear error messages, which makes it the most forgiving part of the lesson.
Connect it up
One page, from memory.
Draw it
Draw two columns headed transient and persistent, and put four programs in them. Beneath, draw the sequence open, write, write, close, and mark on it where the existing file is destroyed and where it becomes complete. Then write one format expression using all three of %d, %g and %s, with its tuple, and note beside it the two ways the matching can fail and the message each gives.
Recap
Two pages, and a program's data can outlive the program.
| If you remember one thing | It is this |
|---|---|
| From persistence | Reading a file does not make a program persistent. Writing one does. |
| From write mode | The file is cleared on open, before anything is written. |
| From write | It takes only strings and adds no newline. print does both favours; write does neither. |
| From the format operator | The type of the left operand decides whether % means modulus or formatting. |
| From the result | '42' is a string. Do the arithmetic first and format last. |
The next lesson deals with finding the file in the first place — the os module, current directories, absolute and relative paths, walking a directory tree — and with the try statement, which handles the many things that can go wrong when a program touches the file system.
Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.