A numbered spec turned into a working program: reading it as a data flow, choosing a dictionary over parallel lists, seeding and guarding the counting loop, validating input without swallowing your own bugs, handling ties correctly, and splitting the code so the logic can be tested without typing votes by hand.
Subject: IT Support & Networking · 64 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
IT · Programming
Eight spec steps, one program, and the grader that runs it on inputs you have not seen
Objectives
The assignment is short but unforgiving: the grader runs your program on inputs you never see, so the program has to be right for reasons you can explain, not because it happened to work once.
The Python Tutorial — python.org — the reference this deck follows
Section
Section 1
Warm-up
Two minutes with the spec and a pen.
Discussion prompt
For any program like this one, what are the three questions you should answer before typing a single line?
Hint: Input, state, output.
Answer:
What comes in, what shape does it have, and where does it come from.
What has to be remembered while the program runs, and in what container.
What goes out, in exactly what format, since a grader usually compares your output character by character.
Answer those three and the code is mostly transcription.
Concept
A spec written as eight steps is usually one data flow with the stages spelled out. Grouping them makes the program's shape obvious.
Read votes, reject the invalid ones, count the rest, find the highest count, print the result. Everything in the spec is one of those five.
CS50x — Harvard's introduction to computer science, the plurality and runoff problem sets — the plurality problem set uses exactly this shape
Picture it
Five boxes. Each becomes one function.
Figure (svg): A pipeline of five stages: read votes, validate, tally, find max, and report
The single biggest cause of a tangled program is starting to type before this picture exists.
Sorting
The spec is deliberately written out of order.
Sort into buckets
Sort each requirement into its stage.
Notice how little state there is. One dictionary is the entire memory of this program.
Section
Section 2
Concept
You need to look up a candidate by name and change their number. That is precisely what a dictionary does, in one step, no searching.
dictionary — A container that maps keys to values. Looking up a key takes the same time whether there are three entries or three million.
Python Standard Library — Mapping Types (dict) — the mapping type reference
Picture it
Three entries, each one a key pointing at a value.
Figure (svg): A dictionary drawn as three key boxes with arrows to three value boxes, mapping Alice to three, Bob to two and Charlie to one
Concept
The alternative is one list of names and one list of counts, kept in the same order. It works, and it is how you would do it in C.
names = ['Alice', 'Bob', 'Charlie']
counts = [0, 0, 0]
# to add a vote you must first find the position
i = names.index('Bob')
counts[i] += 1| operation | with two lists | with a dictionary |
|---|---|---|
| add a vote | search for the index, then increment | one increment |
| risk | the lists can drift out of step | impossible; one structure |
| unknown name | index raises ValueError | handled with a membership test |
| lines of code | more | fewer |
The parallel-list version is not wrong. It is just two things to keep synchronised where one would do, and that is where bugs come from.
The Python Tutorial — python.org — Data Structures
Prediction
A dictionary with three entries, and a lookup for a name that is not one of them.
Predict first
tally = {'Alice': 3}; print(tally['Bob'])
Correct: It raises KeyError.
Why: Square-bracket lookup on a missing key is an error, not a default. This matters here because a vote for an unknown candidate would crash the program rather than being rejected politely, which is exactly the case the grader will test.
Concept
Each is correct in a different situation, and the assignment needs the first one.
# 1. membership test: reject anything not on the ballot
if name in tally:
tally[name] += 1
else:
print('Invalid vote.')
# 2. get with a default: read without inserting
count = tally.get(name, 0)
# 3. setdefault: insert on first sight, for an open ballot
tally.setdefault(name, 0)
tally[name] += 1| form | missing key becomes | use it when |
|---|---|---|
| name in tally | rejected | the candidate list is fixed, as here |
| tally.get(name, 0) | reads as 0, not stored | you are only reading a count |
| tally.setdefault(name, 0) | created with 0 | any name is allowed, like a write-in |
This assignment has a fixed ballot, so the membership test is the right one. The others would silently accept a misspelling as a new candidate.
Python Standard Library — Mapping Types (dict) — dict methods
Discrimination
Choosing the container is most of the design.
Sort into buckets
List, or dictionary?
Section
Section 3
Worked example
Start from a list, not from input. A function that takes a list can be tested; a function that reads from the keyboard cannot.
Build the tally with every candidate at zero
Why: Starting from zero means a candidate with no votes still appears in the output, which the spec requires.
Walk the votes once
Why: One pass, one increment. There is never a reason to loop twice here.
Reject anything not on the ballot
Why: The membership test both validates and protects the increment from a KeyError.
def tally_votes(votes, candidates):
tally = {name: 0 for name in candidates}
for vote in votes:
if vote in tally:
tally[vote] += 1
else:
print('Invalid vote.')
return tally| vote | valid? | Alice | Bob | Charlie |
|---|---|---|---|---|
| start | - | 0 | 0 | 0 |
| Alice | yes | 1 | 0 | 0 |
| Bob | yes | 1 | 1 | 0 |
| Alice | yes | 2 | 1 | 0 |
| Dave | no | 2 | 1 | 0 |
| Charlie | yes | 2 | 1 | 1 |
| Alice | yes | 3 | 1 | 1 |
Verify: that the totals sum to the number of valid votes
Why: Three plus one plus one is five, and six votes were offered with one rejected. If the sum does not match, a vote was counted twice or lost.
Picture it
Six votes, six increments, one pass.
Figure (svg): Six rows showing each incoming vote on the left and the state of the tally after it on the right
Invariant
Step through the loop and watch two quantities.
Step through it
What does the gap between the two numbers tell you, and what would it mean if the gap were negative?
A negative gap is impossible, so if you ever see one, the loop is counting something twice. This is a good assertion to add while debugging.
Trap
A version that adds candidates the first time it sees them.
Annotate
Pass your own test, fail the grader's
Why: Your test data has no typos in it. The grader's does.
Seed the tally from the ballot, then only ever increment.
tally = {name: 0 for name in candidates}
for vote in votes:
if vote in tally:
tally[vote] += 1
else:
print('Invalid vote.')| input | seeded version | grow-as-you-go version |
|---|---|---|
| 'Alice' | Alice becomes 1 | Alice becomes 1 |
| 'alice' | Invalid vote. | a second candidate appears |
| candidate with no votes | reported as 0 | missing entirely |
The ballot is fixed, so the keys should be too
Why: Deciding the keys up front is what makes validation possible at all.
Error analysis
It runs, raises nothing, and gives the wrong answer on every input with a repeat in it.
Annotate
| votes | buggy result | correct result |
|---|---|---|
| Alice | Alice 1 | Alice 1 |
| Alice, Alice | Alice 1 | Alice 2 |
| Alice, Bob, Alice | Alice 1, Bob 1 | Alice 2, Bob 1 |
A single-vote test would pass. This is why the test list in section six has a repeat in it.
Fill the middle
The guard is written; the body is missing.
Fill in the blanks
for vote in votes:
if vote in tally:
tally[vote] += 1
else:
print('Invalid vote.')
Why: The augmented assignment reads the current value, adds one, and stores it back, all in one expression. Writing it the long way is equally correct and four characters longer.
Pattern
Four lines, and every one of them is load-bearing.
Any counting problem you meet later, in any language, is this loop with a different key.
Python Standard Library — collections.Counter — the standard library ships this loop as Counter, once you are allowed to use it
Section
Section 4
Concept
Whatever the user types, input hands you text. A number typed at the prompt is the characters of a number, not the number.
n = input('Number of voters: ')
print(n + 1) # TypeError: can only concatenate str
n = int(input('Number of voters: '))
print(n + 1) # works| expression | type | value | result |
|---|---|---|---|
| input() | str | '6' | text |
| input() + 1 | - | - | TypeError |
| int(input()) | int | 6 | a number |
| int('six') | - | - | ValueError |
The last row is the one graders test: a non-numeric answer to a numeric prompt.
The Python Tutorial — python.org — Input and Output
Prediction
The program calls int(input()) and the user types 'six'.
Predict first
What does the program do?
Correct: Raises ValueError and stops.
Why: int refuses text it cannot parse and raises immediately. If the spec says to keep asking, the fix is a try and except inside a while loop. If it does not, an unhandled crash may be acceptable, but knowing which is which is part of reading the spec.
Worked example
Loop until the input parses, and only then leave the loop.
Loop forever, and break out on success
Why: A while-true loop with a break is the idiomatic shape here; a flag variable is more code doing the same thing.
Attempt the conversion inside a try block
Why: Only the risky line goes inside. Putting the whole body in there hides real bugs behind the same handler.
Handle exactly the exception you expect
Why: Catching a bare exception would also swallow a keyboard interrupt and every typo in your own code.
def ask_int(prompt):
while True:
try:
return int(input(prompt))
except ValueError:
print('Please enter a whole number.')| user types | int() does | loop |
|---|---|---|
| 'six' | raises ValueError | prints the message, asks again |
| '' | raises ValueError | prints the message, asks again |
| '6' | returns 6 | returns, leaving the loop |
| ' 6 ' | returns 6 | returns; int strips whitespace |
Verify: that the last row surprises you in the right direction
Why: int tolerates surrounding whitespace but not an empty string. Knowing which conversions are forgiving is worth more than guessing.
Picture it
This is the output shape the grader compares against.
Figure (svg): A terminal session prompting for six votes, printing Invalid vote for one of them, and reporting Alice as the winner
Graders usually compare output character by character. A trailing space or a missing full stop fails a program that is otherwise perfect.
Trap
A defensive-looking handler.
Annotate
Spend an afternoon on a bug the traceback would have named
Why: This is the single most expensive habit a beginner can pick up.
Catch the specific exception, around the specific line.
try:
n = int(input('Number of voters: '))
except ValueError:
print('Please enter a whole number.')
tally = tally_votes(read_votes(n), candidates)| error | bare except | specific except |
|---|---|---|
| user typed 'six' | friendly message | friendly message |
| typo in a function name | friendly message, bug hidden | full traceback, bug named |
| Ctrl+C | swallowed | program exits as expected |
Keep the risky line alone inside the try
Why: Everything else moves out, so nothing else can be caught by accident.
Section
Section 5
Concept
Ask for the highest count first, then ask who reached it. Doing it in that order handles ties for free.
best = max(tally.values())
winners = [name for name, count in tally.items() if count == best]| step | value with a clear win | value with a tie |
|---|---|---|
| tally | Alice 3, Bob 2, Charlie 1 | Alice 3, Bob 3, Charlie 1 |
| max of the values | 3 | 3 |
| winners list | ['Alice'] | ['Alice', 'Bob'] |
| what to print | Winner: Alice | Winner: Alice and Bob |
Python Standard Library — Mapping Types (dict) — dict.items and dict.values
Picture it
Two candidates on three votes each.
Figure (svg): Three bars showing Alice and Bob tied on three votes with Charlie on one, with both leaders marked as reaching the maximum
Picture it
Scan the tally once, keeping everyone whose count equals the maximum.
Figure (svg): Three candidate rows with the two on the maximum count highlighted and the third marked as below it
A program that returns a single name is not wrong on your test data. It is wrong on the grader's.
Worked example
Print every candidate with their total, then the winner or winners.
Sort by count descending, then by name
Why: Sorting by two keys at once keeps the output stable, so two runs on the same data always print the same order.
Print the table
Why: An f-string with a width specifier lines the columns up without any manual padding.
Print the winners, joined
Why: Joining a list handles one name and three names with the same line of code.
for name, count in sorted(tally.items(), key=lambda kv: (-kv[1], kv[0])):
print(f'{name:<10}{count}')
best = max(tally.values())
winners = [n for n, c in tally.items() if c == best]
print('Winner: ' + ' and '.join(sorted(winners)))| candidate | count | printed line |
|---|---|---|
| Alice | 3 | Alice 3 |
| Bob | 2 | Bob 2 |
| Charlie | 1 | Charlie 1 |
Verify: by running it twice on the same input
Why: The order must be identical both times. If it is not, the sort key is incomplete and the grader may see a different order than you did.
Picture it
This is what the grader compares against.
Figure (svg): A terminal showing the tally printed in aligned columns followed by the winner line
If the spec shows an example of the output, copy its spacing exactly. That example is the specification.
Notation
One line does a lot here, and copying it without understanding it means you cannot change it later.
Annotate
| key returned for | tuple | sorts |
|---|---|---|
| Alice, 3 | (-3, 'Alice') | first |
| Bob, 2 | (-2, 'Bob') | second |
| Charlie, 1 | (-1, 'Charlie') | third |
Using reverse=True instead would put the counts in the right order and the tied names in the wrong one, which is the subtle version of this bug.
Python HOWTO — Sorting Techniques — sorting by multiple keys
Elimination
The tally is Alice 3, Bob 3, Charlie 1.
Eliminate the wrong options
Which one reports both leaders?
Survives elimination: A
Why: Compute the maximum first, then collect everyone equal to it. The other three all collapse the answer to a single winner at the moment they are evaluated, which is the exact case the spec asks you to handle.
Check
Solve it on paper before you click.
Check your understanding
tally = {'Alice': 3, 'Bob': 3, 'Charlie': 1}. What does max(tally, key=tally.get) return?
Answer: A
Why: max returns one item. With the key function it compares counts, and on a tie it keeps the first one it encountered. The tie is not detected and not reported, which is why this shape is wrong for this assignment.
Section
Section 6
Concept
Split the program so the logic takes arguments and returns values, and only a thin outer layer touches input and print.
Then the interesting part can be called from a test with a known list, and you never type six votes by hand again.
The Python Tutorial — python.org — Defining Functions
Picture it
Logic first, input last. Reversing this is why testing feels impossible.
Figure (svg): Four steps: write a function taking a list, call it from a test, compare against an expected dictionary, and only then wire it to input
Concept
Five functions matching the five stages, and a main that wires them together.
CANDIDATES = ['Alice', 'Bob', 'Charlie']
def tally_votes(votes, candidates):
tally = {name: 0 for name in candidates}
for vote in votes:
if vote in tally:
tally[vote] += 1
else:
print('Invalid vote.')
return tally
def winners_of(tally):
best = max(tally.values())
return sorted(n for n, c in tally.items() if c == best)
def report(tally):
for name, count in sorted(tally.items(), key=lambda kv: (-kv[1], kv[0])):
print(f'{name:<10}{count}')
print('Winner: ' + ' and '.join(winners_of(tally)))
def main():
n = int(input('Number of voters: '))
votes = [input('Vote: ') for _ in range(n)]
report(tally_votes(votes, CANDIDATES))
if __name__ == '__main__':
main()| function | takes | returns | touches input or print? |
|---|---|---|---|
| tally_votes | a list and the ballot | a dict | prints rejections only |
| winners_of | a dict | a sorted list | no |
| report | a dict | nothing | prints |
| main | nothing | nothing | reads input |
Two of the four are pure: same input, same output, no side effects. Those two are the ones worth testing, and they are where the marks are.
PEP 8 — Style Guide for Python Code — naming and layout conventions used here
Worked example
No typing votes in. Just call the function.
Write the input and the expected output side by side
Why: If you cannot write the expected output, you do not yet understand the requirement.
Assert equality between dictionaries
Why: Two dictionaries compare equal when they have the same keys and values, regardless of insertion order.
Add the edge cases the grader will use
Why: An empty vote list, an all-invalid list, and a tie. Those three catch nearly everything.
def test_tally():
votes = ['Alice', 'Bob', 'Alice', 'Dave', 'Charlie', 'Alice']
got = tally_votes(votes, CANDIDATES)
assert got == {'Alice': 3, 'Bob': 1, 'Charlie': 1}
assert tally_votes([], CANDIDATES) == {'Alice': 0, 'Bob': 0, 'Charlie': 0}
assert winners_of({'Alice': 3, 'Bob': 3, 'Charlie': 1}) == ['Alice', 'Bob']
print('all tests passed')| case | input | expected |
|---|---|---|
| normal | six votes, one invalid | Alice 3, Bob 1, Charlie 1 |
| empty | no votes | every candidate on 0 |
| tie | a tally with two on 3 | both names, sorted |
Verify: by breaking the code on purpose
Why: Change the increment to add two and re-run. If the test still passes, the test is not testing what you think it is.
Picture it
The traceback names the line, the value, and the call that got there. Read it from the bottom up.
Figure (svg): A Python traceback ending in a KeyError for the name Dave, with the failing line inside tally_votes highlighted
Beginners read tracebacks from the top and give up. The last two lines are the answer, and everything above them is the route taken to get there.
Explain it
Two sentences, out loud.
Discussion prompt
Why is it worth restructuring the program so the logic never calls input?
Hint: How many times will you retype six votes before you stop bothering?
Answer:
Because a function that reads the keyboard can only be exercised by a human typing, so it gets tested once and then never again.
A function that takes a list can be exercised a hundred times a second, which means you can afford to check every edge case every time you change anything.
The Python Tutorial — python.org — Defining Functions
Comparison
The four functions, and what makes each testable or not.
Comparison matrix
| function | pure? | how you would test it |
|---|---|---|
| tally_votes | nearly; it prints rejections | call it with a list, compare the dict |
| winners_of | yes | call it with a tally, compare the list |
| report | no; it prints | capture stdout, or test the pieces it calls |
| main | no; it reads input | run the program by hand, once |
The rule this table is teaching: push the impure parts to the edges, and keep the middle pure.
Section
Section 7
Concept
The last line names the error and the offending value. The line just above it is the code that raised it. Everything higher up is how execution got there.
Two lines answer nearly every question you have. The rest is context you only need when the two lines are not enough.
The Python Tutorial — python.org — Errors and Exceptions
Matching
These five cover almost every failure in a program of this shape.
Match the pairs
Why: Every one of these names its own cause once you know the vocabulary. KeyError is always a missing key; ValueError is always a value of the right type but the wrong content. Learning the five saves more time than any debugging technique.
Trap
Something is wrong, so prints go in.
Annotate
Trade one confusing problem for a wall of text
Why: Print debugging works, but scattering it is what makes it painful.
One print, at the one place the state changes, showing only what changed.
for vote in votes:
if vote in tally:
tally[vote] += 1
else:
print('rejected:', repr(vote))| technique | output volume | finds the bug |
|---|---|---|
| print everything | eighteen lines | eventually |
| print only the anomaly | one line | immediately |
| assert an invariant | silent until it breaks | at the exact moment |
Use repr rather than str when printing input
Why: It shows the quotes and any stray whitespace, which is exactly the class of bug that makes a valid-looking vote get rejected.
Socratic
A vote that looks correct is being rejected.
Discussion prompt
The user typed Alice and the program says invalid. What would repr show that print would not?
Hint: What is invisible when you print a string without its quotes?
Answer:
The quotes and the whitespace. print shows Alice; repr shows 'Alice ' with a trailing space, or 'alice' in lower case.
A trailing space from a copy-paste is the single most common cause of a valid-looking value failing a membership test.
Which suggests the fix the spec probably wants anyway: strip the input, and decide explicitly whether the comparison is case sensitive.
Section
Section 8
Concept
Counting first preferences and taking the highest is plurality. It is what this assignment asks for, and it is the simplest voting rule there is.
The follow-up assignment is usually runoff, where each ballot is a ranked list and the lowest candidate is eliminated and their ballots redistributed until someone passes half.
CS50x — Harvard's introduction to computer science, the plurality and runoff problem sets — plurality then runoff, in that order
Picture it
The counting loop is identical. What changes is what happens after it.
Figure (svg): A five step flow contrasting plurality, which counts once, with runoff, which counts, eliminates the lowest, redistributes and repeats
Which is the real reason to write tally_votes as a function that takes a list: runoff calls it repeatedly, on a shrinking candidate set.
Trade off
Fill in what each one costs.
Comparison matrix
| plurality | runoff | |
|---|---|---|
| ballot shape | one name | a ranked list |
| passes of the tally | one | one per elimination round |
| can elect someone most voters dislike | yes, with a split field | much less easily, since it needs a majority |
| extra code you need | none | elimination and redistribution |
Note the third row: it is the actual reason runoff exists, and it is worth understanding before writing the code that implements it.
Two truths and a lie
Thinking ahead to the runoff version.
Eliminate the wrong options
Which is true?
Survives elimination: A
Why: Passing the candidate list in rather than reading a global is what makes the function reusable across rounds. That single design choice is the difference between runoff being a small extension and a rewrite.
Commit first
Answer, then rate your confidence honestly.
Predict first
Six votes: Alice, Alice, Bob, Bob, Charlie, Charlie. Under plurality, who wins?
Correct: Nobody; it is a three-way tie.
Why: All three candidates have two votes, so the maximum is 2 and all three reach it. A program that returns a single name reports Alice and is wrong, which is exactly the case the winners list exists to handle.
Pattern
Eight spec steps, five stages, four functions, three tests.
| symptom | cause, nearly always |
|---|---|
| KeyError on a vote | square-bracket lookup on a name not in the tally |
| a candidate missing from the output | the tally was grown from votes instead of seeded |
| a tie reports one name | max used on the dict instead of on the values |
| TypeError adding to input() | the string was never converted with int |
| a mysterious 'something went wrong' | a bare except swallowing your own bug |
CS50x — Harvard's introduction to computer science, the plurality and runoff problem sets — the same failure modes, in the plurality problem set
Ranking
Build it in this order and you will never be debugging two things at once.
Put in order
Why: The test comes third, before the rest of the program exists, because from that point on every later change is checked automatically. Writing main first is the natural instinct and it is what makes the last hour of the assignment miserable.
Missing information
The spec says: read the votes, count them, print each candidate's total, print the winner.
Discussion prompt
Name three things the grader will test that the spec does not mention.
Hint: Think about the inputs a malicious grader would choose.
Answer:
What happens on a tie. Almost every spec omits it and almost every grader tests it.
What happens with zero voters, where max of an empty sequence raises ValueError.
Whether the candidate list is case sensitive, and what an unknown name should print.
When a spec is silent, pick the behaviour that cannot crash, and say in a comment which choice you made and why.
Edge cases
The user enters 0 at the first prompt.
Discussion prompt
Trace the program. Where does it break, and what is the smallest fix?
Hint: Which sequence is actually empty: the votes, or the candidates?
Answer:
The tally is built with every candidate on zero, which is fine. Then max of the values returns 0, which is also fine, and every candidate ties for the win.
The genuine crash case is max on an empty sequence, which happens only if the candidate list itself is empty.
The smallest honest fix is a guard: if there are no candidates, say so and stop. Do not silently print a winner that does not exist.
Python Standard Library — Mapping Types (dict) — max raises ValueError on an empty iterable
Check
Solve it on paper before you click.
Check your understanding
With the tally Alice 3, Bob 3, Charlie 1, what does winners_of return?
Answer: A
Why: The maximum count is 3, and two candidates reach it, so the comprehension keeps both. Sorting makes the order predictable so the output is the same on every run.
Exit ticket
One honest answer, and it decides what the next session opens with.
Predict first
Which of these is still shakiest?
Correct: Whichever you named is where we start.
Why: All five are drillable in about fifteen minutes each, and the last one is the highest leverage: once the logic is testable, every other item on the list becomes something you can check instead of something you have to be careful about.
Connect it up
Blank editor, twenty minutes, no looking at these slides.
Draw it
Draw the five-stage flow from memory. Then write tally_votes and winners_of, and a test for each. Do not write main until both tests pass.
Anything you had to look up is the agenda for next session. Bring the file.
Recap
Six sections, and the fourth one is the one graders quietly test on every assignment of this shape.
| input | expected output |
|---|---|
| Alice, Bob, Alice, Dave, Charlie, Alice | Alice 3, Bob 1, Charlie 1, one rejection |
| no votes at all | every candidate 0, all tied |
| a tally with two on 3 | both names, sorted, joined with 'and' |
The Python Tutorial — python.org — every construct used in this deck
Want this taught 1-on-1? Alexander tutors IT Support & Networking — $55/session, free consultation.