This lesson covers the current directory, relative and absolute paths, the os.path functions for inspecting files, a recursive directory walk, and the try statement for handling the many things that go wrong when a program touches the file system.
Subject: Python · 65 slides · code lesson
Open the interactive version of this deck
Title
Python · Chapter 14 — Files
§14.4-14.5, pp. 139-140
Objectives
Five things, each one you can check yourself at an interpreter prompt.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 139-140 — the pages these objectives are drawn from
Warm-up
You wrote open('emma.txt') and it worked. Then it did not.
Discussion prompt
The same program, run from a different folder, reports that emma.txt does not exist — though the file has not moved. What has changed, and what would make the program work from anywhere?
Hint: The filename does not say which folder.
Answer:
The current directory changed. A bare filename says nothing about which folder it is in, so Python looks in whatever directory the program is running from.
Giving the full location — starting from the root of the file system — would make it independent of where the program was started.
Those are the two kinds of path this lesson names, and the distinction between them explains a whole family of it worked yesterday failures.
Concept
Files are organised into directories. Every running program has a current directory, which is the default directory for most operations — so when you open a file for reading, Python looks for it in the current directory.
relative path — A path that starts from the current directory.
A simple filename like memo.txt is considered a path, but a relative one: it relates to the current directory. A path that begins with a separator does not depend on the current directory, and is called an absolute path.
Figure (svg): Two columns contrasting a relative path with an absolute one
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 139-139
Section
Section 1
Concept
The os module provides functions for working with files and directories — os stands for operating system. os.getcwd returns the name of the current directory, where cwd stands for current working directory.
>>> import os
>>> cwd = os.getcwd()
>>> cwd
'/home/dinsdale'
>>> os.path.abspath('memo.txt')
'/home/dinsdale/memo.txt'| Expression | What it gives | Note |
|---|---|---|
| os.getcwd() | the current directory | where relative paths start |
| 'memo.txt' | a relative path | means a file in that directory |
| os.path.abspath(...) | the absolute path | independent of the current directory |
A string like '/home/dinsdale' that identifies a file or directory is called a path. If the current directory is /home/dinsdale, the filename memo.txt would refer to /home/dinsdale/memo.txt.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 139-139
Picture it
The same short name, resolved from two different places.
Figure (svg): A pipeline showing a relative filename joined to the current directory to give an absolute path
Which is why a relative path can be correct in one run and wrong in the next, without the file or the program changing at all.
Worked example
Three questions you can ask without opening anything.
>>> os.path.exists('memo.txt')
True
>>> os.path.isdir('memo.txt')
False
>>> os.path.isdir('/home/dinsdale')
True
>>> os.path.isfile('memo.txt')
True| Function | The question | Note |
|---|---|---|
| exists | is there anything with this name? | file or directory |
| isdir | is it a directory? | False for a file |
| isfile | is it a file? | the complement |
Ask whether it exists.
Why: os.path.exists checks whether a file or directory exists, without opening anything.
Ask what kind of thing it is.
Why: If it exists, os.path.isdir checks whether it's a directory, and similarly os.path.isfile checks whether it's a file.
Note that none of these opens the file.
Why: So none of them can destroy anything, which the previous lesson's write-mode warning makes worth saying.
Figure (svg): The state of the program after each line of Worked example inspecting a path, drawn as a ladder with one rung per traced line
Three yes-or-no answers about a path, obtained without touching the file. These are the safe way to find out what is there.
Verify: Check the order in which to ask.
Why: isdir and isfile both return False for a name that does not exist, so exists distinguishes absent from present but the wrong kind — a distinction that matters when reporting the problem to a user, since the two need different advice.
Prediction
A bare filename, and a known current directory.
# current directory: /home/dinsdale
fin = open('memo.txt')| Part | What it contributes | Result |
|---|---|---|
| 'memo.txt' | a relative path | no directory part |
| the current directory | /home/dinsdale | supplies the rest |
| the file opened | /home/dinsdale/memo.txt |
Predict first
Which file does Python open?
Correct: /home/dinsdale/memo.txt — a relative path is completed by the current directory.
Why: When you open a file for reading, Python looks for it in the current directory, so the bare name refers to a file in whatever directory the program is running from. It does not search the disk, and an incomplete path is not an error — which is why the same program can open different files on different runs.
Worked example
listdir gives names; join makes them usable.
>>> os.listdir(cwd)
['music', 'photos', 'memo.txt']
>>> os.path.join(cwd, 'memo.txt')
'/home/dinsdale/memo.txt'| Call | What it gives | Note |
|---|---|---|
| listdir | files AND directories | bare names, no directory |
| the names alone | relative to the listed directory | not to the current one |
| os.path.join | a directory and a name | a complete path |
List the directory.
Why: os.listdir returns a list of the files — and other directories — in the given directory.
Notice the names are bare.
Why: They carry no directory part, so a name from listing /home/dinsdale is not usable as a path unless the current directory happens to be that one.
Join them.
Why: os.path.join takes a directory and a file name and joins them into a complete path.
Figure (svg): A directory listing with each bare name joined to its directory to give a full path
A list of names and a way to turn each into a usable path. The join step is what makes the recursive walk in the next idea work at any depth.
Verify: Ask why join rather than adding strings.
Why: Because the separator differs between systems, and join uses the right one — so the same code works on Windows and Unix. Concatenating with a hard-coded slash works on one and produces a path that does not exist on the other, which is a portability bug that is invisible until the program moves.
Trap
A program lists another directory and tries to open each name it finds.
Use the names the listing gave you
Why: They are the files in that directory, so they look ready to use.
The names carry no directory part, so they are interpreted relative to the current directory — which is somewhere else. The program reports that files it has just listed do not exist.
Join each name to the directory it came from.
path = os.path.join(dirname, name)
Why: Which is the first line of the book's own walk function.
Then every later operation uses the full path
Why: isfile, open, and the recursive call all take the joined path.
The symptom is distinctive: a file that the program just listed is reported as missing. That combination always means a bare name is being used where a path is needed.
Sorting
Does it depend on where the program is running?
Sort into buckets
For each path, which kind is it?
Faded example
A bare name from a listing is not enough.
Fill in the blanks
for name in os.listdir(dirname):
path = os.path.join(dirname, name)
Why: os.path.join takes a directory and a file name and joins them into a complete path, using the separator the system expects. Concatenating with a hard-coded slash would work on one platform and fail on another, and using the bare name would look for the file in the current directory rather than the one being listed.
Socratic
Absolute paths always work. Relative ones sometimes do not.
Discussion prompt
If absolute paths are unambiguous, why does anyone use relative ones?
Hint: What would an absolute path tie you to?
Answer:
An absolute path names one place on one machine. A program written with absolute paths only works where those exact directories exist, which is usually the author's own computer.
Relative paths describe a structure rather than a location — the data file next to this script — so the whole thing can be moved, copied, or shared and still work.
So the ambiguity is the price of portability. The practical rule is to use relative paths for things that travel with the program and absolute ones for fixed locations on a particular system, and to know which you are relying on.
Section
Section 2
Concept
The following example walks through a directory, prints the names of all the files, and calls itself recursively on all the directories.
def walk(dirname):
for name in os.listdir(dirname):
path = os.path.join(dirname, name)
if os.path.isfile(path):
print(path)
else:
walk(path)| Line | What it does | Note |
|---|---|---|
| listdir | names in this directory | files and directories mixed |
| join | a complete path | usable for the next test |
| isfile | a file: print it | the base case |
| else | a directory: walk it | the recursive case |
The os module provides a function called walk that is similar to this one but more versatile — the exercise is to read the documentation and use it to print the names of the files in a given directory and its subdirectories.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 140-140
Picture it
One call per directory, and the tree's shape drives it.
Figure (svg): A call diagram showing walk recursing into subdirectories
This is chapter 5's recursion with a real tree underneath it, and the base case is supplied by the file system rather than by an arithmetic condition.
Worked example
The order of the three lines matters.
for name in os.listdir(dirname):
path = os.path.join(dirname, name) # first
if os.path.isfile(path): # then test the PATH
print(path)
else:
walk(path) # and recurse on it| What is tested | What it means | Note |
|---|---|---|
| testing name | relative to the current directory | wrong |
| testing path | the actual location | right |
| recursing on name | would lose the directory | and fail one level down |
Join before testing.
Why: os.path.isfile(name) would ask about a file of that name in the current directory, which is not where the listing came from.
Test the joined path.
Why: Only the full path identifies the thing that was actually listed.
Recurse on the joined path.
Why: Otherwise the next call would receive a bare name and could not find the directory at all.
Figure (svg): The state of the program after each line of Worked example why the join must come first, drawn as a ladder with one rung per traced line
Three lines that must be in this order. Every operation after listdir needs the joined path, and the bare name is useful only as an ingredient.
Verify: Trace what happens two levels down.
Why: The recursive call receives an absolute path, so its own listdir and join produce longer absolute paths — the depth grows and the paths stay correct. If the bare name were passed instead, the second level would look in the current directory and find nothing, so the walk would silently stop after one level.
Prediction
The directory contains two folders and a file.
>>> os.listdir('/home/dinsdale')
['music', 'photos', 'memo.txt']| Entry | What it is | Note |
|---|---|---|
| music, photos | directories | included |
| memo.txt | a file | also included |
| the names | bare, no directory part | join is needed |
Predict first
What does the returned list contain?
Correct: Both the files and the subdirectories, as bare names — os.listdir returns a list of the files and other directories in the given directory.
Why: The entries are bare names with no directory part, which is why walk joins each to dirname before doing anything else. And because directories are included, walk has to test isfile rather than assuming — an assumption that would raise IsADirectoryError on the first folder.
Worked example
The recursion stops because directories run out.
# base case: a file - print it and return
if os.path.isfile(path):
print(path)
# recursive case: a directory - go deeper
else:
walk(path)
# and an empty directory: the loop body never runs| What is found | What happens | Note |
|---|---|---|
| a file | printed, no recursion | the base case |
| a directory | one more call | the recursive case |
| an empty directory | listdir gives [] | the loop ends immediately |
Identify the base case.
Why: A file is printed and nothing recurses — which is where a branch of the walk ends.
Identify the recursive case.
Why: A directory produces one more call, on a strictly deeper path.
Notice the empty directory.
Why: listdir returns an empty list, so the for loop body never runs and the call returns — which is a base case the code never mentions.
Figure (svg): A flowchart showing the two branches of the walk function
The recursion terminates because the tree is finite and every call goes strictly deeper. No counter or condition is needed; the structure supplies the ending.
Verify: Ask what would make it not terminate.
Why: A cycle in the tree — which a symbolic link pointing at an ancestor directory can create. The file system is normally a tree and can be made into a graph, so a walk that follows links needs to remember where it has been. That is one of the reasons the book mentions os.walk as more versatile.
Trap
A program lists a directory and opens each name, expecting text files.
Assume a directory contains files
Why: It usually does, and the ones you were thinking about are files.
os.listdir returns a list of the files and other directories, so opening a subdirectory raises IsADirectoryError — one of the three errors the next idea names.
Test what each entry is.
os.path.isfile(path) before opening
Why: Which is exactly what walk does.
Or catch the error and skip it
Why: The try statement, which is often simpler than testing everything.
Both are legitimate, and the next idea argues for the second: checking every possibility in advance takes a lot of time and code, and there are more possibilities than you will think of.
Ranking
Four steps, in order.
Put in order
Why: The join must come before the test, because isfile on a bare name asks about the current directory rather than the one being listed. Everything after the join uses the full path — the test, the printing, and the recursive call.
Faded example
The recursive call needs a usable path.
Fill in the blanks
if os.path.isfile(path):
print(path)
else:
walk(path)
Why: The recursive call takes the joined path, not the bare name — otherwise the next level would look in the current directory and find nothing, and the walk would stop silently after one level. Passing dirname instead would recurse forever on the same directory.
Explain it
A common and quiet failure.
Discussion prompt
A classmate's walk prints the files in the top directory and nothing from the subdirectories, with no error. Diagnose it.
Hint: What is being passed to the recursive call?
Answer:
They are almost certainly recursing on the bare name rather than the joined path — walk(name) instead of walk(path).
The recursive call then looks for a directory of that name in the current directory. It usually is not there, so listdir raises, or finds nothing, and the branch ends quietly.
The fix is one word, and the tell is worth remembering: a recursive traversal that works at the top level and does nothing below it is nearly always losing the path prefix at the recursive call.
Section
Section 3
Concept
A lot of things can go wrong when you try to read and write files.
>>> fin = open('bad_file')
FileNotFoundError: [Errno 2] No such file or directory: 'bad_file'
>>> fout = open('/etc/passwd', 'w')
PermissionError: [Errno 13] Permission denied: '/etc/passwd'
>>> fin = open('/home')
IsADirectoryError: [Errno 21] Is a directory: '/home'| Situation | The error | Number |
|---|---|---|
| the file is missing | FileNotFoundError | Errno 2 |
| you lack permission | PermissionError | Errno 13 |
| it is a directory | IsADirectoryError | Errno 21 |
To avoid these errors you could use functions like os.path.exists and os.path.isfile — but it would take a lot of time and code to check all the possibilities. If Errno 21 is any indication, there are at least twenty-one things that can go wrong.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 140-140
Picture it
Each has its own name, and the numbering hints at the rest.
Figure (svg): A panel listing three file errors with their error numbers
That last line is the argument. Enumerating the failures in advance is not a strategy when you do not know how many there are.
Worked example
Three checks, and the fourth thing still goes wrong.
if os.path.exists(name):
if os.path.isfile(name):
fin = open(name) # can STILL raise
else:
print('not a file')
else:
print('no such file')| Check | What it covers | What remains |
|---|---|---|
| exists | handles one failure | FileNotFoundError |
| isfile | handles another | IsADirectoryError |
| permission | not checked | PermissionError anyway |
Add the obvious checks.
Why: Existence and kind cover the two commonest failures and take five lines.
Notice what is left.
Why: Permission is not checked, so the open can still raise — and so can a dozen things neither of us has thought of.
Notice the second problem.
Why: Even a complete set of checks could go stale: the file can be deleted between the check and the open.
Figure (svg): Two columns comparing checking every possibility with trying and handling failure
Five lines of checking that still cannot guarantee the open succeeds. It would take a lot of time and code to check all the possibilities, and the result would still not be complete.
Verify: Ask what the checks are still good for.
Why: Giving a specific message. exists distinguishes no such file from it is a directory, which is more useful to a user than a generic failure — so the checks earn their place as diagnosis rather than as prevention. That is a different job from stopping the error, and it is worth doing separately.
Prediction
The path names a directory.
>>> fin = open('/home')| Fact | What it rules out | Result |
|---|---|---|
| the path exists | so not FileNotFoundError | |
| it is a directory | not a file | IsADirectoryError |
| the number | Errno 21 |
Predict first
Which error does this raise?
Correct: IsADirectoryError — the path exists and is the wrong kind of thing.
Why: open works on files, so a directory raises a specific error naming that. The distinction matters: FileNotFoundError would mean the path was wrong, whereas this one confirms the path is right and the target is a folder — usually because a name from os.listdir was opened without checking isfile.
Worked example
Each names a different situation, and the right response differs.
# FileNotFoundError: check the name and the current directory
# PermissionError: the file exists; you may not open it that way
# IsADirectoryError: the name is right and points at a folder| Error | What it tells you | Where to look |
|---|---|---|
| not found | a naming or location problem | os.getcwd often explains it |
| permission denied | an access problem | the path is correct |
| is a directory | a kind problem | the path is correct too |
Read FileNotFoundError as a location problem.
Why: The file is not where you looked — which for a relative path usually means the current directory is not what you assumed.
Read PermissionError as the opposite.
Why: The path is right, and the file exists; what failed is access, which no amount of checking the name will fix.
Read IsADirectoryError as a kind problem.
Why: The path is right and points at a folder, which usually means a listing was used where a file was expected.
Figure (svg): The state of the program after each line of Worked example reading the three messages, drawn as a ladder with one rung per traced line
Three errors that look similar and mean quite different things. Two of them confirm the path is correct, which narrows the investigation considerably.
Verify: Use os.getcwd when a file is not found.
Why: Printing the current working directory alongside the failing name shows immediately whether the program is looking where you think — and it is the single most useful thing to print for a FileNotFoundError on a relative path, because that error is much more often about the directory than about the name.
Trap
A program catches any failure to open and reports that the file does not exist.
Assume the commonest cause
Why: Most of the time the file really is missing.
A permission problem or a directory now reports a missing file, so the user checks the name — which is correct — and gets nowhere. The message actively points away from the cause.
Report what actually happened.
Let the specific error through, or name it in the message
Why: The three have three different remedies.
And print the path you tried
Why: Absolute, if you can, since a relative path is only half the story.
This is the same principle as lesson 11b's raise: report the failure where it happens, in terms that identify it. A message that mislabels the problem is worse than the raw traceback.
Discrimination
Three situations, three errors.
Sort into buckets
For each situation, which error does open raise?
Two truths and a lie
Two are true. Keep the lie.
Eliminate the wrong options
Rule out the two true statements.
Survives elimination: C
Why: C is false twice over. Permission is not covered by either check, so the open can still raise — and even a complete set of checks would leave a gap, since the file can be deleted or its permissions changed between the check and the open. That gap is one reason to try rather than to check.
Real world
The gap between checking and acting is a general problem.
Discussion prompt
Think of a situation where you confirmed something was available and it was gone by the time you acted. What made the check useless?
Hint: Anything other people can also change.
Answer:
A seat shown as free that is taken by the time you book; an item in stock that sells out during checkout; a slot that two people claim at once.
What makes the check useless is that the world can change between checking and acting, so the answer describes a moment that has already passed.
Which is why systems that must be right attempt the action and handle the failure, rather than asking first — exactly the argument for try over a sequence of checks. The check is still useful for telling someone what is likely; it just cannot be a guarantee.
Section
Section 4
Concept
It is better to go ahead and try — and deal with problems if they happen — which is exactly what the try statement does. The syntax is similar to an if...else statement.
catch — To prevent an exception from terminating a program, using the try and except statements.
try:
fin = open('bad_file')
except:
print('Something went wrong.')| Part | What happens | Note |
|---|---|---|
| the try clause | runs first | the thing that might fail |
| all goes well | the except clause is skipped | and execution proceeds |
| an exception occurs | jump out, run except | the failure is handled |
Python starts by executing the try clause. If all goes well, it skips the except clause and proceeds. If an exception occurs, it jumps out of the try clause and runs the except clause.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 140-141
Picture it
Exactly one of the two clauses runs to completion.
Figure (svg): A flowchart showing the two paths through a try and except statement
The important difference from if...else is that the try clause may be abandoned partway through — so anything after the failing line does not run.
Worked example
The book criticises its own example, and the criticism is the lesson.
# the book's example, with its own verdict:
try:
fin = open('bad_file')
except:
print('Something went wrong.') # not very helpful| Aspect | What is true | Note |
|---|---|---|
| catching | the program does not stop | which is the mechanism |
| this message | says nothing useful | the book says so |
| what it should do | fix, retry, or exit gracefully | the three good responses |
Note what catching achieves mechanically.
Why: Handling an exception with a try statement is called catching an exception, and it stops the exception terminating the program.
Note the book's own criticism.
Why: In this example, the except clause prints an error message that is not very helpful.
Note what a good handler does.
Why: In general, catching an exception gives you a chance to fix the problem, or try again, or at least end the program gracefully.
Figure (svg): The state of the program after each line of Worked example what catching is for, drawn as a ladder with one rung per traced line
Three legitimate purposes, and printing a vague message is none of them. Catching without doing one of the three converts a clear failure into a confusing one.
Verify: Compare with not catching at all.
Why: The uncaught traceback names the exception, the file, and the line — considerably more than Something went wrong. So a handler that discards that information and continues is worse than no handler, which is why catching should be a decision about what to do rather than a reflex.
Prediction
The file does not exist.
try:
fin = open('bad_file')
print('opened')
except:
print('Something went wrong.')| Line | What happens | Output |
|---|---|---|
| open raises | the try clause is abandoned | immediately |
| print('opened') | never runs | it is after the failing line |
| the except clause | runs | 'Something went wrong.' |
Predict first
What does this print?
Correct: Something went wrong. — the exception jumps out of the try clause, so the print after it never runs.
Why: If an exception occurs, Python jumps out of the try clause and runs the except clause — abandoning the rest of the try, not merely skipping the failing line. That is the key difference from an if statement, and it is why wrapping a lot of code in one try leaves half-finished work behind.
Worked example
Fix it, retry, or end gracefully.
# fix the problem: fall back to a default
try:
fin = open(name)
except FileNotFoundError:
fin = open('defaults.txt')
# or end gracefully, saying what failed
except FileNotFoundError:
print('Cannot find', os.path.abspath(name))| Response | When it fits | Note |
|---|---|---|
| fix | supply an alternative | the program continues correctly |
| retry | ask again, or wait | for transient failures |
| end gracefully | say what happened, then stop | better than a traceback for a user |
Fix the problem where there is a sensible alternative.
Why: A missing configuration file can fall back to defaults, and the program carries on doing the right thing.
Retry where the failure may be temporary.
Why: Which is common for networks and rare for files.
End gracefully otherwise.
Why: Report what failed, in terms the user can act on, and stop — rather than continuing with data you do not have.
Figure (svg): Two columns contrasting a vague catch-all handler with one that names the exception and responds
Three responses, chosen by what the failure means. Naming the exception in the except clause is what lets you respond to one kind and let the others through.
Verify: Ask what a bare except costs.
Why: It catches everything, including a typo in the try clause that raises NameError — so a genuine bug in your own code is silently reported as a file problem. Naming the exception you expect keeps unrelated failures visible, which is why the book's bare except is worth improving on even though it is what the section shows.
Trap
A whole function body is placed inside a try, with one except at the end.
Protect everything at once
Why: One handler is simpler than several, and covers the whole operation.
Any failure anywhere in the body now reaches the same handler, so a bug in the processing is reported as a file error — and the try clause is abandoned wherever it failed, leaving half-done work behind.
Wrap the smallest thing that can fail.
Just the open, or just the operation you expect to fail
Why: So the handler knows what went wrong.
And name the exception
Why: So unrelated failures are not swallowed.
The narrower the try clause, the more the except clause knows. A handler covering twenty lines can only say that one of twenty things failed, which is not enough to fix, retry, or report usefully.
Invariant
Two runs of the same statement.
Step through it
In which frames does the code after the try statement run?
In both — the second and the fourth. That is the point of catching: the program continues either way, and only the route through the statement differs.
Faded example
Naming the exception keeps other failures visible.
Fill in the blanks
try:
fin = open(name)
except FileNotFoundError:
print('Cannot find', os.path.abspath(name))
Why: Naming the exception means this handler responds to a missing file and lets everything else through — a permission problem, or a typo in your own code raising NameError, stays visible instead of being reported as a missing file. A bare except catches all of them, which is what makes it a blunt instrument.
Explain it to yourself
Both avoid a crash. Only one of them can be complete.
Discussion prompt
Explain, using the book's own reasoning, why it is better to go ahead and try than to check every possibility first.
Hint: How many things can go wrong?
Answer:
Because you cannot enumerate the failures. The book's remark about Errno 21 is the argument: there are at least twenty-one things that can go wrong, and checking each would take a lot of time and code.
And because a check describes a moment that has passed. A file can be deleted or its permissions changed between the check and the open, so even complete checking would leave a gap.
The try statement covers every failure mode with one construct and has no gap, because the handling is attached to the attempt itself. Catching then gives you a chance to fix the problem, or try again, or at least end the program gracefully.
Section
Section 5
Concept
The two techniques are not alternatives so much as different jobs. Checking tells a user what is wrong; trying keeps the program running whatever happens.
def read_or_default(name, fallback):
try:
return open(name).read()
except FileNotFoundError:
print('Not found:', os.path.abspath(name))
return fallback| Part | What it does | Note |
|---|---|---|
| the try clause | the smallest thing that can fail | one operation |
| the named exception | only a missing file | others propagate |
| the message | the absolute path | which is the useful part |
| the return | a sensible fallback | the program continues |
The message uses the absolute path because a relative one is only half the story — and for a FileNotFoundError, the missing half is usually the problem.
Python documentation — os.path — Common pathname manipulations os.path — Common pathname manipulations
Picture it
One prevents nothing and explains well; the other prevents everything and explains nothing.
Figure (svg): Two columns separating the diagnostic role of checks from the protective role of try
So the checks belong inside the handler as often as before the attempt — diagnosing the failure you have caught rather than trying to prevent it.
Worked example
The checks are more useful after the failure than before it.
try:
fin = open(name)
except OSError:
if not os.path.exists(name):
print('No such file:', os.path.abspath(name))
elif os.path.isdir(name):
print('That is a directory:', name)
else:
print('Cannot open:', name)| Part | What it contributes | Note |
|---|---|---|
| the try | covers every failure | nothing slips past |
| the checks | run only after a failure | no cost in the normal case |
| the messages | specific to the cause | useful to a user |
Attempt first.
Why: The try covers every possible failure, including the ones you have not thought of.
Diagnose after.
Why: The os.path checks now explain a failure that has definitely happened, rather than trying to predict one.
Note the cost.
Why: None in the normal case — the checks run only when something has already gone wrong.
Figure (svg): A flowchart showing an attempt followed by diagnostic checks inside the handler
Complete coverage and a specific message, with the checks doing the job they are good at. This is the shape worth reaching for.
Verify: Ask why the checks cannot be stale here.
Why: They can be, in principle — the file could appear between the failure and the check — but it does not matter, because they are producing a message rather than deciding what to do. A slightly wrong message is a much smaller problem than a wrong decision, which is why the same imprecise checks are fine in this position and not in the other.
Prediction
The handler prints and does nothing else.
try:
fin = open('missing.txt')
except FileNotFoundError:
print('not found')
for line in fin:
print(line)| Stage | What happens | Result |
|---|---|---|
| open raises | fin is never assigned | the try is abandoned |
| the handler | prints and returns | fin still unassigned |
| for line in fin | fin does not exist | NameError |
Predict first
What happens when the loop runs?
Correct: A NameError — fin was never assigned, because the try clause was abandoned before the assignment completed.
Why: Catching an exception keeps the program running, and that is only useful if the handler leaves things in a usable state. Here it prints and falls through with the work undone, so the failure reappears further from its cause — which is worse than the original traceback. The handler should have supplied a fallback or returned.
Worked example
The commonest path bug, and its fix.
import os
# fragile: depends on where the program is run from
fin = open('data/emma.txt')
# robust: relative to the script, not the current directory
here = os.path.dirname(os.path.abspath(__file__))
fin = open(os.path.join(here, 'data', 'emma.txt'))| Approach | What it depends on | Note |
|---|---|---|
| the relative path | completed by the current directory | wherever that is |
| __file__ | the script's own path | known regardless |
| join | the data file beside the script | portable |
Identify what the fragile version assumes.
Why: That the program is run from the directory containing data — which is true when you test it and often false later.
Anchor to the script instead.
Why: os.path.dirname of the script's absolute path gives the directory the program lives in, which does not move.
Join from there.
Why: The data file is described by its position relative to the script rather than to the current directory.
Figure (svg): The state of the program after each line of Worked example making a program work from anywhere, drawn as a ladder with one rung per traced line
A program that finds its own data whatever directory it is started from — using join for portability and an anchor that does not depend on the caller.
Verify: Test it by running from a different directory.
Why: The fragile version raises FileNotFoundError and the anchored one works, which is the whole difference. Running from elsewhere at least once is worth doing deliberately, because this bug is invisible while you develop in the project directory and appears the moment anyone else runs the program.
Trap
A handler prints a message and lets the program continue with the file object unassigned.
Keep the program running
Why: Which is what catching is for.
The next line uses a variable that was never assigned, so the program fails again with a NameError — further from the cause and harder to read than the original error would have been.
Make sure the handler leaves the program in a usable state.
Supply a fallback, or return, or exit
Why: Fix the problem, try again, or end gracefully — the book's own three.
Never fall through with the work undone
Why: If there is nothing sensible to continue with, stopping is the graceful option.
Catching is a promise that the program can proceed. If it cannot, the honest thing is to report clearly and stop — which is still much better than an uncaught traceback for a user, and much worse than pretending nothing happened.
Sorting
Ask whether you want a decision or a message.
Sort into buckets
For each job, which technique fits?
Faded example
A relative path is half the story.
Fill in the blanks
except FileNotFoundError:
print('Cannot find', os.path.abspath(name))
Why: os.path.abspath shows where the program actually looked, which for a relative path is the missing half of the diagnosis. Printing the bare name tells the user something they already knew; printing the absolute path usually reveals that the current directory was not what anyone assumed.
Explain it
The classic path failure, and it has a standard cause.
Discussion prompt
A classmate's program finds its data file when they run it and not when anyone else does. Diagnose it and give them the fix.
Hint: Where are they running it from?
Answer:
They are opening a relative path, which is completed by the current directory — and they always run from the project folder, where it happens to be right.
Anyone starting the program from elsewhere gets a different current directory, so the same relative path names a file that is not there. The program has not changed and neither has the file.
The fix is to anchor to something that does not move: os.path.dirname(os.path.abspath(__file__)) gives the script's own directory, and joining from there finds the data wherever the program is started. Printing os.getcwd() in the handler would have shown them the problem immediately.
Comparison
Fill the blanks. They do different jobs and are often used together.
Comparison matrix
| Question | os.path checks | try / except |
|---|---|---|
| What does it cover? | the cases you thought of | every failure, including ones you did not |
| Can it go stale? | yes — the file can change after the check | no — the handling is attached to the attempt |
| How specific is it? | very — exists, isdir, isfile | as specific as the exception you name |
| What is it best for? | producing a useful message | keeping the program running |
Which is why the checks often belong inside the handler: attempt first for coverage, then diagnose what actually failed.
Pattern
Five steps, and the second is what most programs skip.
Step 5's one of three things is the test of a handler. Printing a vague message and falling through is none of them, and it produces a second failure further from the cause.
Python documentation — os.path — Common pathname manipulations os.path — Common pathname manipulations
Check
The current directory is /home/dinsdale.
Check your understanding
Which file does open('memo.txt') refer to?
Answer: A
Why: A simple filename is a relative path: it relates to the current directory, so if the current directory is /home/dinsdale, the filename memo.txt refers to /home/dinsdale/memo.txt. Python does not search elsewhere, which is why the same program can open different files depending on where it is run.
Check
An exception occurs partway through the try clause.
try:
a = open('missing.txt')
b = 1
except:
print('failed')| Line | What happens | Note |
|---|---|---|
| open raises | the try clause is abandoned | immediately |
| b = 1 | never runs | it is after the failing line |
| except | runs | 'failed' |
Check your understanding
Is b assigned?
Answer: A
Why: If an exception occurs, Python jumps out of the try clause and runs the except clause — so everything after the failing line is abandoned, not merely the failing line itself. That is why wrapping a lot of code in one try can leave half-finished work behind.
Check
os.path.exists and isfile are available.
Check your understanding
Why does the book prefer a try statement to checking every possibility first?
Answer: A
Why: It would take a lot of time and code to check all the possibilities — and if Errno 21 is any indication, there are at least twenty-one things that can go wrong. It is better to go ahead and try, and deal with problems if they happen, which also closes the gap where a file changes between the check and the open.
Real world
Ask permission or ask forgiveness — a choice that comes up everywhere.
Discussion prompt
Think of a process that checks eligibility before acting, and one that acts and handles rejection. What does each get wrong?
Hint: What happens between the check and the action?
Answer:
Checking first gives a clear answer up front and can be wrong by the time you act — the seat was free when you looked and is taken when you book.
Acting first and handling the failure is always accurate about the moment that matters, and it tells you less in advance: you find out by trying.
Which is why the two are usually combined: check to tell someone what is likely, and handle the failure to be correct. That is exactly the shape this lesson arrives at — try for coverage, and os.path checks inside the handler for the message.
Commit first
Answer, then rate your confidence.
Predict first
Why is checking with os.path.exists before opening not enough to prevent errors?
Correct: Because many other things can go wrong, and the file can change between the check and the open.
Why: The book makes the first half of this argument directly: you could use functions like os.path.exists and os.path.isfile, but it would take a lot of time and code to check all the possibilities — and if Errno 21 is any indication, there are at least twenty-one things that can go wrong. Existence does not cover permission, and neither covers the failures you have not thought of. The second half is the gap: a check describes a moment that has already passed, so a file can be deleted or its permissions changed between the check and the open. It is better to go ahead and try, and deal with problems if they happen — which is exactly what the try statement does, and it leaves no gap because the handling is attached to the attempt itself.
Explain it
Two techniques that look like alternatives and are not.
Discussion prompt
A classmate asks whether they should check a file exists or use try. Give them the answer and the reason each is good at its own job.
Hint: One covers everything and one explains well.
Answer:
Use try, because it is the only one that covers every failure — including permission problems and the twenty or so possibilities neither of you has enumerated — and because a check can go stale before the open.
Then use the os.path checks inside the handler, where they explain what actually went wrong: absent, a directory, or something else. That is a specific message rather than a prediction.
And whatever the handler does, it must do one of three things — fix the problem, try again, or end gracefully. Printing a message and carrying on with the work undone just moves the failure somewhere less informative.
Exit ticket
One honest answer. It decides what the next lesson opens with.
Predict first
Which of these is still least solid for you?
Correct: Whichever you picked is the right answer — this one is for you, not for a mark.
Why: The path distinction explains a whole family of it works on my machine failures, and printing os.getcwd() in a handler is the fastest diagnosis there is. The walk is chapter 5's recursion over a real tree, with the join step as the thing that must come first. The argument against exhaustive checking is the section's real content and is easy to skim past. And the try statement is mechanically simple, with the difficulty entirely in what the handler should do — fix, retry, or end gracefully, and nothing else counts.
Connect it up
One page, from memory.
Draw it
Draw a small directory tree and write beside two of its files their relative and absolute paths, marking which one depends on the current directory. Underneath, write the walk function from memory and circle the line that must come first. Then draw the two paths through a try statement, and beside it list the three errors open can raise and the three things a handler may usefully do.
Recap
Two pages, and a program can deal with a file system it did not create.
| If you remember one thing | It is this |
|---|---|
| From paths | A filename is not a location. The current directory supplies the rest. |
| From join | Never build a path by concatenating a separator you typed. |
| From walk | Join before you test, or you are asking about the wrong directory. |
| From the errors | There are more than you can enumerate, which is the argument for try. |
| From try | A handler must fix, retry, or end gracefully. A message alone is none of those. |
The next lesson finishes the chapter with the other routes to persistence: dbm databases that behave like dictionaries on disk, pickling to store arbitrary objects, pipes for running other programs, writing your own modules with the __name__ idiom, and repr for debugging invisible whitespace.
Think Python, 2nd edition — Allen B. Downey §14.4-14.5, pp. 139-140 — everything on these slides traces back here
Want this taught 1-on-1? Alexander tutors Python — $55/session, free consultation.