14a Persistence, Reading and Writing Files, the Format Operator

This lesson introduces persistent programs, shows how to open a file for writing and what that destroys, and covers the format operator for putting non-string values into a file.

Subject: Python · 65 slides · code lesson

Open the interactive version of this deck

What this lesson covers

The lesson, slide by slide

1. Lesson 14a Persistence, Reading and Writing Files, the Format Operator

Title

Python · Chapter 14 — Files

§14.1-14.3, pp. 137-138

2. By the end of this lesson you can

Objectives

Five things, each one you can check yourself at an interpreter prompt.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138 — the pages these objectives are drawn from

3. Before we start: where did the histogram go?

Warm-up

You built one in chapter 13. Then the program ended.

Discussion prompt

Analysing Emma takes a noticeable time and produces a histogram of 7,214 words. Run the program twice and it does the work twice. What would you need to avoid that, and what is stopping you?

Hint: Where does the histogram live?

Answer:

You would need the histogram to survive after the program ends — written somewhere that is still there next time.

What stops you is that everything so far has lived in memory, which is gone the moment the program stops. Every program you have written starts with a clean slate.

This chapter is about the alternative: keeping data in permanent storage, so that a program can pick up where it left off.

4. The one idea behind this chapter: data that outlive the program

Concept

Most of the programs we have seen so far are transient: they run for a short time and produce some output, but when they end, their data disappears. If you run the program again, it starts with a clean slate.

persistent — Pertaining to a program that runs indefinitely and keeps at least some of its data in permanent storage.

Other programs are persistent: they run for a long time or all the time, they keep at least some of their data in permanent storage, and if they shut down and restart, they pick up where they left off. Operating systems and web servers are the obvious examples.

Figure (svg): Two columns contrasting a transient program with a persistent one

Everything you have written so far is in the left column.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-137

5. Opening a file for writing

Section

Section 1

6. Mode 'w', and what it destroys

Concept

To write a file, you have to open it with mode 'w' as a second parameter. If the file already exists, opening it in write mode clears out the old data and starts fresh — so be careful. If the file doesn't exist, a new one is created.

>>> fout = open('output.txt', 'w')
SituationWhat happensNote
the file existsits contents are clearedimmediately, on open
the file does not exista new one is createdempty
either wayyou get a file objectwith methods for working with it

This is the first operation in the book that can destroy something outside the program, and it does so at the moment of opening — before you have written a single character, and without asking.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138

7. Picture it: what each mode does to an existing file

Picture it

The second argument decides whether the old contents survive.

Figure (svg): Two columns comparing opening an existing file for reading and for writing

One extra argument, and the existing file's contents are gone.

The clearing happens on open, not on write — so a program that opens a file for writing and then crashes has still emptied it.

8. Worked example: writing two lines

Worked example

open, write, close — and the return value is worth reading.

>>> fout = open('output.txt', 'w')
>>> line1 = "This here's the wattle,\n"
>>> fout.write(line1)
24
>>> line2 = 'the emblem of our land.\n'
>>> fout.write(line2)
24
>>> fout.close()
LineWhat happensNote
open(..., 'w')a file objectthe file is now empty
write(line1)returns 24the number of characters written
write(line2)adds to the ENDthe file object tracks its position
close()finishesor it happens when the program ends

Open in write mode.

Why: open returns a file object that provides methods for working with the file.

Write, and note the return value.

Why: The write method puts data into the file, and the return value is the number of characters that were written.

Write again.

Why: The file object keeps track of where it is, so if you call write again, it adds the new data to the end.

Figure (svg): The state of the program after each line of Worked example writing two lines, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Two lines in the file, and two counts of 24. When you are done writing, you should close the file — and if you don't, it gets closed for you when the program ends.

Verify: Count the characters yourself.

Why: "This here's the wattle," is 23 characters and the newline makes 24, which matches what write returned. Confirming the count includes the newline is worth doing once, because a missing newline is the commonest reason a written file looks wrong when it is read back.

9. Predict: what happens to the existing file?

Prediction

The file already contains a hundred lines.

fout = open('notes.txt', 'w')
# nothing is written yet
fout.close()
StepWhat happensResult
open with 'w'clears the fileon opening
nothing writtenthe file stays empty0 characters
closean empty file remainsthe hundred lines are gone

Predict first

What does notes.txt contain afterwards?

  • Nothing — opening in write mode cleared it
  • The original hundred lines, since nothing was written
  • A copy of the original with a blank line added
  • It raises an error, since the file already existed

Correct: Nothing — if the file already exists, opening it in write mode clears out the old data and starts fresh.

Why: The clearing happens at open, before any write. That is what makes this trap so effective: a program can destroy a file without executing a single write, and without any warning. If you want to add to an existing file rather than replace it, mode 'a' appends instead.

10. Worked example: why closing matters

Worked example

It happens automatically, and relying on that has a cost.

# explicit
fout = open('out.txt', 'w')
fout.write('data\n')
fout.close()

# or let it happen at the end of the program
# - but nothing else can rely on the file
# being complete until then
ApproachWhen the file is finishedNote
close()finishes the fileat a point you choose
not closingclosed when the program endsat a point you do not
in betweenthe file may be incompletefor anything else reading it

Note the rule.

Why: When you are done writing, you should close the file. If you don't close it, it gets closed for you when the program ends.

Note why should rather than must.

Why: The automatic close means a short script will work either way, which is why the omission so often goes unnoticed.

Note when it bites.

Why: A long-running program, or one that writes a file and then reads it back, cannot rely on the contents until the file is closed.

Figure (svg): A pipeline from opening a file through writing to closing it

The clearing happens at the first stage and the completing at the last.

Closing is not strictly required and is still worth doing, because it is the point at which the file is definitely complete — and that point is otherwise the end of the program.

Verify: Try reading a file back before closing it.

Why: The read may come back empty or short, because what was written is still buffered rather than on the disk. That failure looks like a bug in the writing code and is really a missing close, which is a good reason to make the habit automatic.

11. Trap: opening a file for writing to check it exists

Trap

The trap

A program opens a file in write mode to see whether it is there, planning to read it if so.

Open it and find out

Why: Which works for reading, where a missing file raises.

Write mode creates the file if it is absent and empties it if it is present — so the check destroys exactly the data it was looking for, and reports success either way.

The fix

Test before opening, or open for reading.

os.path.exists tells you without touching anything

Why: Which the next lesson covers.

Or open for reading and handle the failure

Why: A missing file raises, and the file is unharmed.

The rule to internalise: mode 'w' is destructive at the moment of opening. Nothing warns you, nothing asks, and the old contents are not recoverable.

12. Predict: what does write return?

Prediction

The return value is easy to overlook.

fout = open('out.txt', 'w')
n = fout.write('hello\n')
print(n)
PartWhat is trueResult
'hello'five charactersplus the newline
writereturns the count6
the filenow holds the lineand is not yet closed

Predict first

What does this print?

  • 6
  • 5
  • None
  • The string 'hello'

Correct: 6 — the return value is the number of characters that were written, and the newline counts.

Why: write returns a character count rather than None, which is unusual among the methods you have met — most modifying methods return nothing. The newline is a single character, so 'hello\n' is six. Forgetting that the newline counts is the commonest reason people expect 5.

13. Discriminate: does this destroy the file?

Discrimination

Ask what the mode is.

Sort into buckets

For each call, is an existing file's content lost?

destroys the contents
open('f.txt', 'w'); open('f.txt', 'w') then immediately close
leaves the file alone
open('f.txt'); open('f.txt', 'r'); os.path.exists('f.txt'); reading every line of f.txt
yes
Both open in write mode, which clears the file at the moment of opening. Closing immediately does not help — the damage is done by open, not by write.
no
Reading, whether the mode is given explicitly or defaulted, never modifies the file; and checking existence does not open it at all.

14. Think it through: why clear the file on open?

Socratic

It would be safer to clear on the first write.

Discussion prompt

Opening in write mode empties the file immediately, even if nothing is ever written. Why might the language do that rather than waiting?

Hint: What is the mode a statement about?

Answer:

Because the mode is a declaration of intent: you have said this file is going to be replaced, so the old contents are not wanted.

Waiting would mean the file's state depended on whether a write happened to occur, which would make the behaviour harder to predict rather than easier — sometimes cleared, sometimes not.

The practical consequence is the one to remember: opening is the destructive step. If you might not want to replace the file, do not open it in write mode at all — check first, or use append mode.

15. The format operator

Section

Section 2

16. Two meanings for one symbol

Concept

The argument of write has to be a string, so if we want to put other values in a file, we have to convert them. When applied to integers, % is the modulus operator — but when the first operand is a string, % is the format operator.

>>> 7 % 3
1
>>> camels = 42
>>> '%d' % camels
'42'
>>> 'I have spotted %d camels.' % camels
'I have spotted 42 camels.'
ExpressionWhich operatorResult
7 % 3both integersmodulus: 1
'%d' % camelsfirst operand a stringformatting: '42'
embedded in a sentencethe sequence can be anywherea longer string

The first operand is the format string, which contains one or more format sequences specifying how the second operand is formatted. The result is a string — so '42' is not to be confused with the integer 42.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-138

17. Picture it: the same symbol, two operators

Picture it

Which one applies is decided by the type on the left.

Figure (svg): Two columns showing the percent symbol acting as modulus and as formatting

Same symbol. The type of the first operand decides which operation happens.

That is unusual — most operators mean one thing — and it is why the book introduces the distinction explicitly before using it.

18. Worked example: three format sequences

Worked example

Several values need a tuple.

>>> 'In %d years I have spotted %g %s.' % (3, 0.1, 'camels')
'In 3 years I have spotted 0.1 camels.'
SequenceWhat it formatsFrom the tuple
%da decimal integer3
%ga floating-point number0.1
%sa string'camels'
the tuplematched in orderone element per sequence

Note when a tuple is needed.

Why: If there is more than one format sequence in the string, the second argument has to be a tuple.

Note how they are matched.

Why: Each format sequence is matched with an element of the tuple, in order — positionally, exactly like tuple assignment.

Note the three common sequences.

Why: %d for a decimal integer, %g for a floating-point number, and %s for a string.

Figure (svg): The state of the program after each line of Worked example three format sequences, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

One string built from three values of three different types. The tuple's order must match the sequences' order, since the matching is positional.

Verify: Swap two values in the tuple and see which fails.

Why: Swapping the 3 and the 0.1 gives '%d' a float — which Python accepts by truncating — while swapping in the string raises. So some mismatches are loud and some are quiet, which is an argument for reading the sequence order carefully rather than relying on an error.

19. Predict: which operator applies?

Prediction

The first operand is a string.

print('%d' % 7)
PartWhat is trueResult
the first operanda stringso % formats
%da decimal integer7
the resulta string'7'

Predict first

What does this print?

  • 7
  • 1
  • 0
  • A TypeError

Correct: 7 — the first operand is a string, so % is the format operator and the result is the string '7'.

Why: When applied to integers % is the modulus operator, but when the first operand is a string it formats. Option B would be the modulus of some pair of numbers; nothing here is arithmetic. Note that the result is the string '7', not the integer — which matters if it is used in a calculation afterwards.

20. Worked example: the two ways it fails

Worked example

Too few elements, or the wrong type.

>>> '%d %d %d' % (1, 2)
TypeError: not enough arguments for format string
>>> '%d' % 'dollars'
TypeError: %d format: a number is required, not str
MistakeWhat is wrongMessage
three sequences, two valuesa count mismatchnot enough arguments
%d given a stringa type mismatcha number is required
bothTypeErrorand both messages are specific

Check the count.

Why: The number of elements in the tuple has to match the number of format sequences in the string.

Check the types.

Why: The types of the elements have to match the format sequences — %d wants a number.

Read the two messages.

Why: In the first example there aren't enough elements; in the second, the element is the wrong type. Each message says which.

Figure (svg): A panel showing the three distinct TypeError messages the format operator produces

Two TypeErrors with two distinct messages. Both failures are loud, which is a relief after this chapter's silent file destruction.

Verify: Try too many elements rather than too few.

Why: '%d' % (1, 2) raises not all arguments converted during string formatting — a third distinct message. So all three count-and-type mismatches are caught, and the message tells you which one you have without further investigation.

21. Trap: forgetting the tuple for a single value

Trap

The trap

A program writes '%s and %s' % 'a', 'b' — with the two values as separate arguments.

Pass the values the way you would to a function

Why: Comma-separated arguments are how everything else takes several values.

The format operator takes exactly one right-hand operand, so this formats 'a' alone and then makes a tuple of the result with 'b'. It raises about missing arguments — or, in a print call, silently produces something odd.

The fix

Put the values in a tuple.

'%s and %s' % ('a', 'b')

Why: One operand, which happens to hold two values.

A single value needs no tuple

Why: '%d' % 42 works, though '%d' % (42,) also does.

The asymmetry is worth noting: one value may be passed bare, and several must be a tuple. The single-element tuple form works in both cases, which is why some people always write it.

22. Sort: modulus or formatting?

Sorting

Look at the type on the left.

Sort into buckets

For each expression, which operation does % perform?

modulus
17 % 5; n % 2; total % len(t)
formatting
'%d' % 5; '%s items' % count; '%g' % 0.1
mod
In each the first operand is a number, so % is the arithmetic remainder operator — the one from chapter 5's divisibility tests.
fmt
In each the first operand is a string containing format sequences, so % builds a new string from the value on the right.

23. Complete it: format two values

Faded example

Several sequences need one tuple.

Fill in the blanks

print('%d camels and %s' % (3, 'a llama'))

Why: If there is more than one format sequence in the string, the second argument has to be a tuple, and each sequence is matched with an element in order. Passing them as separate comma-separated arguments would not work: the operator takes exactly one right-hand operand.

24. Error analysis: four format expressions

Error analysis

Mark each and name the failure.

Annotate

  • Line 1 has three sequences and two elements, so it raises TypeError: not enough arguments for format string.
  • Line 2 gives %d a string, so it raises TypeError: %d format: a number is required, not str.
  • Line 3 has one sequence and two elements, which raises TypeError: not all arguments converted during string formatting.
  • Line 4 is fine: %s accepts anything, since any value can be converted to a string — so an integer formats without complaint.
  • So three of the four raise, each with a message naming the specific problem, and the fourth is the reason %s is the forgiving one.
  • The two rules are that the counts must match and the types must match, and %s relaxes the second.

Every mismatch is caught, which makes the format operator one of the safer things in this chapter.

25. Converting values for a file

Section

Section 3

26. write takes a string, and nothing else

Concept

The argument of write has to be a string, so if we want to put other values in a file, we have to convert them to strings. The easiest way to do that is with str.

>>> x = 52
>>> fout.write(str(x))

# or with the format operator, which can do more:
>>> fout.write('%d\n' % x)
CallWhat is passedResult
fout.write(x)an integerTypeError
fout.write(str(x))convertedworks
fout.write('%d\n' % x)converted and formattedwith a newline

str is the simplest conversion; the format operator does the same job and can add surrounding text, control the appearance of a number, and append the newline in one expression.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-138

27. Picture it: two routes from a value to a file

Picture it

Both end in a string, which is the only thing write accepts.

Figure (svg): Two paths converting a number into a string suitable for writing

A file is a sequence of characters, so everything written has to become characters first.

That is worth stating plainly: a text file has no notion of an integer. Everything in it is text, and reading it back means converting in the other direction.

28. Worked example: writing a histogram to a file

Worked example

Two values per line, formatted and terminated.

fout = open('counts.txt', 'w')
for word, freq in hist.items():
    fout.write('%s %d\n' % (word, freq))
fout.close()
PartWhat it producesNote
'%s %d\n'a word, a space, a number, a newlineone line's worth
% (word, freq)the two values, in ordermatched positionally
writeone line per itemthe position advances

Build the line as a string.

Why: One format expression produces the whole line, including the separator and the newline.

Match the tuple to the sequences.

Why: %s takes the word and %d the count, in the order they appear.

Write and close.

Why: Each write appends at the current position, so the lines accumulate in order.

Figure (svg): A ladder showing successive writes appending lines to a file

The file object keeps track of where it is, so writes accumulate rather than overwrite.

A file with one word and count per line — the histogram made persistent, so a later run need not recompute it.

Verify: Check the newline is present.

Why: Without the \n every line would run into the next and the file would be one enormous line — recoverable, but only by knowing where the boundaries were. Since write does not add a newline of its own, supplying it is the writer's job, and forgetting it is the commonest defect in a written file.

29. Predict: does this write succeed?

Prediction

The argument is an integer.

fout = open('out.txt', 'w')
fout.write(52)
PartWhat is trueResult
writerequires a stringstrictly
52an integernot converted automatically
the resultTypeErrormust be str

Predict first

What happens?

  • A TypeError — the argument of write has to be a string
  • It writes 52 to the file
  • It writes nothing and returns 0
  • It writes the character with code 52

Correct: A TypeError — the argument of write has to be a string, so the integer must be converted first.

Why: str(52) or '%d' % 52 both produce the string write wants. print would have accepted the integer and converted it, which is why the inconsistency catches people — print converts and terminates lines for you, and write does neither.

30. Worked example: what the file does not remember

Worked example

Everything written becomes text, and the types are gone.

fout.write('%d\n' % 42)      # writes the characters '4', '2'

# reading it back:
fin = open('out.txt')
line = fin.readline()        # '42\n' - a STRING
n = int(line)                # convert back explicitly
DirectionWhat happensNote
writing42 becomes '42'the type is discarded
reading'42\n'a string, with the newline
int(line)back to a numberthe conversion is yours to do

Note what is stored.

Why: A text file is a sequence of characters, so the integer becomes the characters '4' and '2' and nothing records that it was a number.

Note what comes back.

Why: Reading gives a string, including the newline — the file cannot tell you it was an integer.

Convert explicitly.

Why: int() turns it back, and it tolerates the trailing newline, though strip makes the intent clearer.

Figure (svg): The state of the program after each line of Worked example what the file does not remember, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

Round-tripping a number through a text file loses its type, and restoring it is the reader's responsibility. That is the limitation the pickle module addresses later in the chapter.

Verify: Try it with a list.

Why: str([1, 2, 3]) writes '[1, 2, 3]' and reading it back gives that string, with no easy way to recover the list — int and float have inverses, and a general value does not. That gap is exactly why the chapter goes on to databases and pickling.

31. Trap: passing a non-string to write

Trap

The trap

A program writes fout.write(count) where count is an integer.

Pass the value you want in the file

Why: print accepts anything, so write looks like it should too.

It raises TypeError: write() argument must be str. print converts for you and write does not, which is an inconsistency between two things that both put text somewhere.

The fix

Convert first.

str(count), or a format expression

Why: The argument of write has to be a string.

And add the newline yourself

Why: write does not add one, whereas print does by default.

Both differences run the same way: print is the convenient one that converts and terminates, and write is the literal one that does exactly what it is told. Remembering that print is the exception makes both predictable.

32. Compare: print and write

Comparison

Fill the blanks. Both put text somewhere and they differ in three ways.

Comparison matrix

Questionprintwrite
Where does it go?the screena file
Does it convert values?yes — anything can be printedno — the argument must be a string
Does it add a newline?yes, by defaultno — you supply it
What does it return?Nonethe number of characters written

print is the convenient one and write is the literal one. Every difference runs that way, which makes them easy to keep straight.

33. Complete it: write a number with a newline

Faded example

Two conversions and a line ending, in one expression.

Fill in the blanks

fout.write('%d\n' % count)

Why: write does not add a line ending, so it has to be part of the string. Without it every value would run into the next and the file would be a single line — readable only by someone who knew where the boundaries had been.

34. Explain it yourself: why does a text file lose types?

Explain it to yourself

An integer goes in and a string comes out.

Discussion prompt

Explain why writing 42 to a text file and reading it back gives you a string rather than a number.

Hint: What is a text file made of?

Answer:

A text file is a sequence of characters, and nothing else. Writing 42 stores the characters '4' and '2'; there is nowhere in the file to record that they were meant as a number.

So reading gives characters back, and turning them into an integer is a decision the reading program makes — the file cannot make it, because the information was never stored.

Which is why the chapter goes on to pickling: a format that does record types can restore the original value, where a text file can only give you text. For numbers the conversion is easy either way; for a list or a dictionary it is not.

35. Choosing a format sequence

Section

Section 4

36. %d, %g and %s

Concept

Three format sequences cover almost everything: a decimal integer, a floating-point number, and a string.

>>> '%d' % 42
'42'
>>> '%g' % 0.1
'0.1'
>>> '%s' % 'camels'
'camels'
>>> '%s' % 42
'42'          # %s accepts anything
SequenceWhat it formatsNote
%da decimal integerrequires a number
%ga floating-point numberchooses a readable form
%sa stringaccepts any value

%s is the forgiving one, because any value can be converted to a string. The other two are specific, and giving them the wrong type raises rather than guessing.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 138-139

37. Picture it: three sequences and what they accept

Picture it

Two are strict and one takes anything.

Figure (svg): Three format sequences with the types each accepts

%s is the safe default; the other two say something about the value.

Using %s everywhere always works and gives up the checking — which is a fair trade when you are only writing text, and a loss when a type mistake would be worth catching.

38. Worked example: why %d is not just %s

Worked example

The strict sequence catches a mistake the forgiving one hides.

>>> count = 'twelve'      # a bug: should be a number
>>> '%d items' % count
TypeError: %d format: a number is required, not str
>>> '%s items' % count
'twelve items'            # legal, and probably wrong
SequenceWhat happensNote
%d with a stringraises immediatelythe bug is found
%s with a stringformats happilythe bug survives
the differencechecking against conveniencea real trade

Notice what %d refuses.

Why: The types of the elements have to match the format sequences, so a string where a number belongs raises.

Notice what %s accepts.

Why: Anything, because any value can be converted to a string — which means it can never tell you something is wrong.

Weigh the two.

Why: The strict sequence turns a type confusion into an immediate error; the forgiving one lets it through into the output.

Figure (svg): The state of the program after each line of Worked example why d is not just s, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

%d is a small type check written into the format string. Using %s everywhere is simpler and gives that up.

Verify: Ask when %s is the right choice anyway.

Why: When the value genuinely may be of several types, or when it is already a string. Choosing %s deliberately is fine; choosing it to avoid thinking about types is what loses the check. That distinction — a decision against a default — is the same one behind get versus bracket lookup in lesson 11a.

39. Predict: what type is the result?

Prediction

The format operator produces something specific.

x = '%d' % 42
print(type(x))
PartWhat is trueResult
the format operatorthe result is a stringalways
'42'characters, not a valuenot the integer
type<class 'str'>

Predict first

What does this print?

  • <class 'str'>
  • <class 'int'>
  • <class 'float'>
  • <class 'tuple'>

Correct: <class 'str'> — the result of the format operator is a string.

Why: The book flags this directly: the result is the string '42', which is not to be confused with the integer value 42. Formatting is for producing output, so anything arithmetic should happen before it — otherwise the plus operator concatenates instead of adding.

40. Worked example: what %g does

Worked example

It picks a readable form for a floating-point number.

>>> '%g' % 0.1
'0.1'
>>> '%g' % 100000000.0
'1e+08'
>>> '%f' % 0.1
'0.100000'        # always six decimal places
SequenceWhat it choosesResult
%g on a small numberplain form0.1
%g on a large onescientific notation1e+08
%fa fixed six placesregardless of the value

Use %g for a general float.

Why: It formats a floating-point number in whichever of the two forms is more compact, which is usually what you want for output.

Note it switches to scientific notation.

Why: For very large or very small numbers, which keeps the string short and can surprise you if you expected digits.

Compare with %f.

Why: Always a fixed number of decimal places, which is better when columns must line up and worse when the magnitudes vary.

Figure (svg): Two columns comparing the general float format with the fixed-decimal one

One adapts to the value and one does not, which is the whole difference.

%g adapts and %f does not. Neither is right in general; the choice depends on whether you want compactness or alignment.

Verify: Format a whole number as a float.

Why: '%g' % 3.0 gives '3' rather than '3.0', which is more compact and loses the visible indication that the value was a float. If that distinction matters in the output, %f or an explicit precision is the better choice — which is the kind of thing worth checking once rather than discovering in a report.

41. Trap: using the formatted result as a number

Trap

The trap

A program computes total = '%d' % count + 1, expecting to add one to the count.

Treat the formatted value as the number

Why: It looks like a number and prints like one.

The result of the format operator is a string, so this concatenates rather than adds — or raises, if the 1 is an integer. The book warns explicitly that '42' is not to be confused with the integer 42.

The fix

Format at the last moment, for output only.

Do the arithmetic on numbers

Why: total = count + 1, and format when writing.

Then '%d' % total for the file

Why: One conversion, at the boundary.

The general shape is worth keeping: values stay in their own types inside the program and become text only where they leave it. Converting early means every later operation deals with strings that are pretending to be numbers.

42. Discriminate: which sequence for this value?

Discrimination

Two are strict about type and one is not.

Sort into buckets

For each value, which sequence fits best?

%d
a count of items; a year
%g
an average; a proportion between 0 and 1
%s
a word; a name
d
Both are whole numbers where a decimal point would be wrong, and %d also refuses a non-number, which is a small type check worth having.
g
Both are floating-point values, and %g chooses a compact readable form rather than a fixed number of decimal places.
s
Both are text already, and %s is the sequence that accepts any value — which here is exactly right rather than a way of avoiding the question.

43. Complete it: format a count

Faded example

A whole number, and a sequence that checks it is one.

Fill in the blanks

print('I have spotted %d camels.' % 42)

Why: %d formats a decimal integer and refuses a non-number, so it catches a value that should have been a count and is not. %s would also work here and would accept anything at all, giving up that check — which is fine when chosen deliberately and a loss when chosen by default.

44. Explain it: why does my total come out as text?

Explain it

A very common early bug.

Discussion prompt

A classmate formatted a number, added one, and got '421' instead of 43. Explain what happened.

Hint: What did the format operator return?

Answer:

The format operator returns a string, so '%d' % 42 is the two characters '4' and '2', not the number forty-two.

Adding 1 to a string with the plus operator concatenates rather than adds — or raises, if the 1 was an integer. Either way the value stopped being a number the moment it was formatted.

The fix is to keep values in their own types inside the program and format only where they leave it, on the way to a file or the screen. Formatting early means everything after it is text pretending to be a number.

45. Making a program persistent

Section

Section 5

46. Two routes, and what each is good for

Concept

One of the simplest ways for programs to maintain their data is by reading and writing text files. An alternative is to store the state of the program in a database.

We have already seen programs that read text files; in this chapter we will see programs that write them. The choice between the routes is another data structure decision, of exactly the kind chapter 13 was about.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-137

47. Picture it: what a program keeps and what it loses

Picture it

Persistence is about which part of the state survives the ending.

Figure (svg): A flowchart contrasting a transient program's ending with a persistent one's restart

The whole difference is whether anything reached permanent storage before the end.

Which is why the chapter is short on new concepts and long on mechanics: the idea is simple and the details of doing it safely are not.

48. Worked example: making chapter 13's analysis persistent

Worked example

Compute once, save, and load thereafter.

import os

if os.path.exists('counts.txt'):
    hist = load_counts('counts.txt')     # fast
else:
    hist = process_file('emma.txt')      # slow
    save_counts(hist, 'counts.txt')
SituationWhat happensNote
the file existsload itseconds saved
it does notcompute and saveonce
subsequent runsalways the fast pathpersistence

Check whether the saved data exist.

Why: The existence check comes before any open, so nothing is destroyed by asking.

Load if they do, compute if they do not.

Why: The expensive work happens once, and every later run reads the result.

Save after computing.

Why: So that the next run takes the fast path.

Figure (svg): Two columns comparing recomputing every run with loading saved results

The same trade as any cache: speed against the risk of an answer that no longer matches.

A program that pays the analysis cost once. This is the memo pattern from chapter 11, with a file in place of a dictionary — and the file survives the program's ending, which the dictionary could not.

Verify: Ask what makes the saved data stale.

Why: A change to emma.txt, or to the cleaning rules. The saved histogram records an answer without recording the question, so nothing detects that it no longer matches — which is the general hazard of caching that lesson 11c raised, and Fibonacci avoided by having answers that never change.

49. Predict: transient or persistent?

Prediction

The program writes nothing to storage.

hist = process_file('emma.txt')
print(len(hist))
# the program ends
StageWhat happensNote
the histogramin memorywhile it runs
the program endsmemory is releasedthe data are gone
running againstarts with a clean slaterecomputes everything

Predict first

Is this program transient or persistent?

  • Transient — its data disappear when it ends
  • Persistent — it read a file, so its data are on disk
  • Persistent — the histogram survives in memory
  • Neither, since it produces output

Correct: Transient — it runs for a short time and produces some output, but when it ends, its data disappear.

Why: Reading a file does not make a program persistent; keeping its own data in permanent storage does. This one starts with a clean slate every run and recomputes the whole histogram, which is exactly the situation writing a file would fix.

50. Worked example: which route for which data

Worked example

A text file is not always the right container.

# fine in a text file: one word and count per line
fout.write('%s %d\n' % (word, freq))

# awkward: a dictionary of lists of tuples
# - what separator? what escaping?
# - this is what pickle is for
DataRouteNote
flat, simple valuesa text file is fineone line each
nested structuresawkward to encodeseparators collide
arbitrary program datapickletypes preserved

Consider what the data look like.

Why: A flat mapping from words to counts is one line each and reads back easily.

Consider what a nested structure needs.

Why: A dictionary of lists of tuples has to be encoded with separators, and any separator can appear in the data — lesson 12c's compound-key hazard again.

Note where each route ends.

Why: A text file is simple, readable and lossy; pickle makes it easy to store program data as it is.

Figure (svg): The state of the program after each line of Worked example which route for which data, drawn as a ladder with one rung per traced line

The whole run at once: each drop is one line of the program.

The shape of the data decides the route. A text file is right for something flat and human-readable, and wrong for an arbitrary structure.

Verify: Ask what a text file buys that pickle does not.

Why: Readability by anything at all — you can open it, another program can read it, and it survives a change of language. That is a real advantage, which is why the format persists for data simple enough to fit it.

51. Trap: writing the output file before reading the input

Trap

The trap

A program opens its output file at the top, then reads its input, and the two happen to be the same file.

Set everything up first

Why: Opening all the files at the start looks tidy.

Opening for writing clears the file, so the input is destroyed before it is read — and the program reports an empty input rather than an error.

The fix

Open the output only when you are ready to write it.

Read the input fully first

Why: Then nothing that follows can destroy it.

And never let the two names be the same

Why: Write to a different name, and rename afterwards if you must replace it.

This is the most destructive mistake in the chapter and it is entirely silent. The file is gone, the program succeeds, and the output is empty — which looks like a bug in the analysis rather than in the file handling.

52. Sort: transient or persistent?

Sorting

Ask what happens when it restarts.

Sort into buckets

For each program, which kind is it?

persistent
an operating system; a web server; a program that saves your progress
transient
a script that prints today's date; a calculator that exits after one sum; the word-count program from chapter 13
per
Each keeps at least some of its data in permanent storage and picks up where it left off — the book names operating systems and web servers as the standard examples.
tra
Each runs for a short time, produces output, and loses its data when it ends. Running it again starts with a clean slate.

53. Complete it: check before you open

Faded example

Ask whether the saved data are there.

Fill in the blanks

import os
if os.path.exists('counts.txt'):
hist = load_counts('counts.txt')
else:
hist = process_file('emma.txt')

Why: os.path.exists checks whether a file or directory exists without opening it, so the check itself cannot destroy anything. Trying to find out by opening in write mode would create the file if it were absent and empty it if it were present — the destructive answer to a harmless question.

54. Where persistence is the whole point

Real world

Most software you use is in the persistent column.

Discussion prompt

Think of something you use that would be useless if it forgot everything when it closed. What does it store, and where would it hurt most to lose it?

Hint: Anything with your work in it.

Answer:

A document editor, a messaging app, a game with saved progress, a browser with its history — all of them keep state on disk precisely so that closing is not losing.

What they store is the part you would mind recreating: the text, the conversation, the position, the settings. Anything cheap to recompute usually is not stored.

Which is the design question this chapter poses: what is expensive enough to save, and what is safer to recompute? Chapter 13's histogram is a good candidate because it takes real time; the two totals derived from it are not, because they are one line each.

55. Compare: reading and writing a file

Comparison

Fill the blanks. One mode is safe and the other is not.

Comparison matrix

Questionopen(name)open(name, 'w')
If the file existsread it, unchangedclear it immediately
If it does notFileNotFoundErrorcreate a new one
What the methods doread and iterate lineswrite, returning a character count
Can it destroy data?noyes, on opening, silently

The bottom row is the one to keep. Opening for writing is destructive before anything is written.

56. The procedure: writing data to a text file

Pattern

Six steps, and the second exists because the first is destructive.

  1. Decide the output filename, and make sure it is not the input's.
  2. Check whether the file exists, if losing it would matter.
  3. Open it with mode 'w', which clears it immediately.
  4. Convert each value to a string, with str or a format expression.
  5. Include the newline yourself, since write does not add one.
  6. Close the file when you are done writing.

Step 1 is the one that prevents the chapter's worst mistake. Opening the input file for writing destroys it before a single line has been read, and the program then reports an empty input rather than an error.

Python documentation — Input and Output Input and Output

57. Check yourself 1 of 3: write mode

Check

The file already has contents.

Check your understanding

What happens when you open an existing file with mode 'w'?

  • A. Its contents are cleared immediately (correct)
  • B. New data are added to the end
  • C. An error is raised, since the file exists
  • D. Nothing until the first write

Answer: A

Why: If the file already exists, opening it in write mode clears out the old data and starts fresh — at the moment of opening, before anything is written. That is why a program that opens a file for writing and then crashes has still emptied it.

Why B tempts people
That is mode 'a', for appending. Mode 'w' replaces rather than adds.
Why C tempts people
No error is raised. The silence is what makes this dangerous.
Why D tempts people
The clearing happens on open. Waiting for a write would make the file's state depend on whether one occurred.

58. Check yourself 2 of 3: the format operator

Check

Two format sequences.

print('%s has %d legs' % ('a cat', 4))
SequenceWhat it formatsValue
%sa string'a cat'
%da decimal integer4
the tuplematched in orderpositionally

Check your understanding

What does this print?

  • A. a cat has 4 legs (correct)
  • B. %s has %d legs
  • C. A TypeError about argument count
  • D. ('a cat', 4)

Answer: A

Why: Each format sequence is matched with an element of the tuple, in order. Two sequences and two elements, with the types matching, so the result is the completed sentence. Passing the values as separate arguments rather than a tuple would fail, since the operator takes exactly one right-hand operand.

Why B tempts people
The sequences are replaced rather than printed literally — that is what the operator does.
Why C tempts people
The counts match: two sequences, two elements. A mismatch would give this message.
Why D tempts people
The tuple is consumed by the formatting rather than printed.

59. Check yourself 3 of 3: what write accepts

Check

One of these fails.

Check your understanding

Which call raises a TypeError?

  • A. fout.write(52) (correct)
  • B. fout.write(str(52))
  • C. fout.write('%d' % 52)
  • D. fout.write('52')

Answer: A

Why: The argument of write has to be a string, so an integer must be converted first — with str, or with a format expression. print would have accepted the integer and converted it for you, which is the inconsistency that catches people.

Why B tempts people
str converts the integer to a string, which is exactly what the book recommends as the easiest way.
Why C tempts people
The format operator produces a string, so this is fine and can add surrounding text as well.
Why D tempts people
This is already a string, so nothing needs converting.

60. Where this shows up outside this course

Real world

Replacing a file rather than adding to it is a distinction with consequences.

Discussion prompt

Think of a time something was overwritten when it should have been added to, or the reverse. What made the difference invisible until afterwards?

Hint: Save as, against save.

Answer:

Saving over the original instead of a copy; a backup that replaced the previous one rather than joining it; an export that silently overwrote last month's.

What makes it invisible is that both operations succeed and look identical from outside — a file exists afterwards either way, and only its contents differ.

Which is exactly mode 'w' against mode 'a'. One character in the call decides whether the previous contents survive, nothing warns you, and the result is not recoverable. It is worth treating the write-mode open as the dangerous line it is.

61. Confidence wager: commit before you check

Commit first

Answer, then rate your confidence. This one can cost real data.

Predict first

You call open('notes.txt', 'w') on a file containing a hundred lines, then close it without writing. What does the file contain?

  • Nothing — write mode clears the file on opening
  • The original hundred lines, since nothing was written
  • The hundred lines plus a blank line
  • An error is raised because the file already existed

Correct: Nothing — if the file already exists, opening it in write mode clears out the old data and starts fresh.

Why: The clearing happens at the moment of opening, not at the first write, which is what makes this the most dangerous line in the chapter. A program can destroy a file without executing a single write and without any warning — and the destruction is not recoverable. The book's own comment is simply so be careful. The practical consequences are worth stating: never open a file in write mode to find out whether it exists (os.path.exists answers that without touching it), never open the output file before the input has been fully read, and use mode 'a' when you mean to add rather than replace.

62. Explain it to someone else

Explain it

print and write both put text somewhere, and they differ in three ways.

Discussion prompt

A classmate's file is one enormous line and their write call raised a TypeError before that. Explain both, in terms of how write differs from print.

Hint: print does two favours that write does not.

Answer:

The TypeError is because write takes only a string — print converts any value for you and write does not, so a number has to be passed through str or a format expression.

The single enormous line is because print adds a newline by default and write does not. Every line ending in a written file has to be supplied as part of the string.

Both differences run the same way: print is the convenient one and write is the literal one that does exactly what it is told. Remembering that print is the exception makes both predictable rather than surprising.

63. Exit ticket

Exit ticket

One honest answer. It decides what the next lesson opens with.

Predict first

Which of these is still least solid for you?

  • Transient against persistent, and what makes a program one or the other
  • Opening for writing, and what it destroys
  • The format operator, and its two meanings
  • Format sequences, and matching a tuple to them

Correct: Whichever you picked is the right answer — this one is for you, not for a mark.

Why: The persistence distinction is simple and worth being precise about, since reading a file does not make a program persistent. The write-mode warning is the one thing in this lesson that can cost you data, and it is worth over-learning: the clearing happens on open. The two meanings of % are unusual in the language and settle quickly once you know to look at the left operand. And the tuple matching is ordinary positional matching with three clear error messages, which makes it the most forgiving part of the lesson.

64. Synthesis: draw the map of this lesson

Connect it up

One page, from memory.

Draw it

Draw two columns headed transient and persistent, and put four programs in them. Beneath, draw the sequence open, write, write, close, and mark on it where the existing file is destroyed and where it becomes complete. Then write one format expression using all three of %d, %g and %s, with its tuple, and note beside it the two ways the matching can fail and the message each gives.

65. What you can do now

Recap

Two pages, and a program's data can outlive the program.

If you remember one thingIt is this
From persistenceReading a file does not make a program persistent. Writing one does.
From write modeThe file is cleared on open, before anything is written.
From writeIt takes only strings and adds no newline. print does both favours; write does neither.
From the format operatorThe type of the left operand decides whether % means modulus or formatting.
From the result'42' is a string. Do the arithmetic first and format last.

The next lesson deals with finding the file in the first place — the os module, current directories, absolute and relative paths, walking a directory tree — and with the try statement, which handles the many things that can go wrong when a program touches the file system.

Think Python, 2nd edition — Allen B. Downey §14.1-14.3, pp. 137-138 — everything on these slides traces back here

Sources

  1. Think Python, 2nd edition — Allen B. Downey — Allen B. Downey, Think Python: How to Think Like a Computer Scientist, 2nd edition (Green Tea Press, 2015), §14.1-14.3, pp. 137-138
  2. Python documentation — Input and Output
  3. Python documentation — Built-in Types

Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108