CS 161, Lesson 9, in 50 slides and code mode. It covers three classes of memory-safety vulnerability: format-string bugs, through printf's walk up the stack, the %x leak, and the %n write; integer conversion bugs, through signed against unsigned, overflow that under-allocates, and truncation; and off-by-one bugs, through the fence-post error, a single-null overwrite of the saved frame pointer, and strncpy leaving a string unterminated. It includes several full trace tables and is anchored to textbook sections 3.3 to 3.5.
Subject: Computer Security · 91 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS 161 · Lesson 9 of 45
printf's stack walk · %x leak / %n write · signed vs unsigned · the single null byte
Objectives
printf(user_input) is a vulnerability and printf("%s", user_input) is safe, and why the compiler cannot catch the difference.% modifier reads on the stack — starting 8 bytes above printf's rip, walking up 4 bytes each — and tell %x (read) apart from %n (write).size_t parameter bypasses a bounds check, and how count * size overflow under-allocates.strncpy may leave a string not null-terminated.Warm-up
Discussion prompt
Before we open L09 · Format Strings, Integer Conversion & Off-by-One Bugs: without looking back, what was the main idea of L08 · Buffer Overflows & Stack Smashing, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
CS 161 Lesson 8 (50 slides, code mode): why C has no bounds checking, the vulnerable gets(buf) frame, the 8+4=12-byte offset to the rip, little-endian address typing, three shellcode-placement strategies, NOP sleds, the worm history, and the fgets fix. Toy/sandbox examples only.
Section
Part A · §3.3
Concept
Today's three classes look unrelated, but they share a theme from L8: attacker data crossing into a role it shouldn't have. A format string becomes a program; a length becomes a size; one extra byte becomes a frame pointer.
| Class | Data that escapes its role | Becomes |
|---|---|---|
| Format string (§3.3) | user text | printf's instructions |
| Integer conversion (§3.4) | a signed length | a huge unsigned size |
| Off-by-one (§3.5) | one extra byte | the saved frame pointer |
All three are memory-safety bugs in C, and all three end the same way: corrupted memory the attacker steers toward control of the program.
Comparison
Comparison matrix
From Three bugs, one root: data trusted as control: refill the Data that escapes its role column from what you know. The rest of the table is as it appeared.
| Class | Data that escapes its role | Becomes |
|---|---|---|
| Format string (§3.3) | user text | printf's instructions |
| Integer conversion (§3.4) | a signed length | a huge unsigned size |
| Off-by-one (§3.5) | one extra byte | the saved frame pointer |
Concept
printf takes a variable number of arguments. It learns how many — and of what type — only by reading the format string at runtime and counting the % modifiers it finds.
variadic function — A function that accepts a variable-length argument list (declared with ...). The callee, not the compiler, decides how many arguments to pull off the stack.
printf("x=%d y=%d\n", x, y); // 2 modifiers -> reads 2 args
printf("hello\n"); // 0 modifiers -> reads 0 args| Format string | % modifiers | Args printf will fetch |
|---|---|---|
| "x=%d y=%d\n" | 2 | 2 (x and y) |
| "hello\n" | 0 | 0 |
| "%d %d %d" | 3 | 3 |
Trade off
Comparison matrix
From printf is variadic — it trusts its first argument: every row here is a choice with a cost. Fill the % modifiers column, then say which row you would actually pick and what you give up for it.
| Format string | % modifiers | Args printf will fetch |
|---|---|---|
| "x=%d y=%d\n" | 2 | 2 (x and y) |
| "hello\n" | 0 | 0 |
| "%d %d %d" | 3 | 3 |
Intuition
Think of the format string as a little program that printf executes. Each % is an instruction: go grab the next argument off the stack and print it this way.
If the attacker writes that little program (because the format string came from user input), the attacker decides how many arguments printf grabs and what it does with them. The data on the stack is no longer printf's to interpret — it's the attacker's.
| Who controls the format string? | Who decides what printf reads? |
|---|---|
| The programmer (literal) | The programmer — safe |
| The attacker (user input) | The attacker — vulnerable |
Counterexample
Discussion prompt
Think of the format string as a little program that printf executes. Each % is an instruction: go grab the next argument off the stack and print it this way.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
The whole bug is which slot the user input lands in: as data behind a %s it is harmless; as the format string itself it is interpreted.
// SAFE: user_input is DATA, consumed by the %s
printf("%s", user_input);
// VULNERABLE: user_input IS the format string
printf(user_input);| Call | Role of user_input | Verdict |
|---|---|---|
| printf("%s", user_input) | argument printed by %s | safe |
| printf(user_input) | the format string | vulnerable (CWE-134) |
Analogy
Discussion prompt
Explain Safe vs vulnerable: one argument of difference by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The whole bug is which slot the user input lands in: as data behind a %s it is harmless; as the format string itself it is interpreted.
Concept
printf is legally variadic — taking 0, 1, or 20 arguments is all valid. So the compiler cannot flag printf(buf) as wrong: a format string with zero % modifiers and zero extra arguments is a perfectly legal call.
The mismatch only exists at runtime, once the actual string is known. (Modern compilers offer a heuristic warning — -Wformat-security — but it is opt-in and only fires on a non-literal format string; the language itself permits the call.)
| When is the format known? | Can it be type-checked? |
|---|---|
| Literal "%d" at compile time | yes (optional -Wformat) |
| Runtime user_input | no — string not known yet |
Concept
Any function that takes a format string has the same exposure if user data reaches that parameter: fprintf, sprintf, snprintf, vprintf, and especially syslog(priority, user_msg) — a very common real-world sink.
syslog(LOG_INFO, user_msg); // VULNERABLE if user_msg has %n
syslog(LOG_INFO, "%s", user_msg); // safe
sprintf(out, fmt_from_user, ...); // VULNERABLE format parameter| Function | Format parameter | Risk if user-controlled |
|---|---|---|
| printf | 1st | leak + write |
| fprintf / sprintf | 2nd | leak + write |
| snprintf | 3rd | leak + write (bounded output) |
| syslog | 2nd | leak + write (classic real bug) |
Discrimination
Sort into buckets
Sort these by Format parameter, from memory, without looking back at It is not only printf. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
When printf runs and meets its first % modifier, its internal argument pointer starts 8 bytes above printf's saved rip and walks upward 4 bytes for each modifier consumed.
high addr
... (further stack args)
arg #1 (first real argument) rip + 8 <- argptr STARTS here
&format-string rip + 4
saved rip of printf rip + 0
saved ebp of printf
low addr (printf's locals)| Stack slot | Offset from printf's rip | What lives there |
|---|---|---|
| printf's saved rip | rip + 0 | return address into the caller |
| format-string pointer | rip + 4 | address of the format string |
| first variadic arg | rip + 8 | where argptr begins |
Estimation
Predict first
Output: x=11 y=22 z=33. Every modifier had a matching argument, so every read landed on real data.
Commit before you compute: what does Worked: printf("x=%d y=%d z=%d", x, y, z) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Each %d consumes the next 4-byte slot, walking up
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. First %d reads rip+8 (x), second reads rip+12 (y), third reads rip+16 (z).
Worked example
int x = 11, y = 22, z = 33;
printf("x=%d y=%d z=%d", x, y, z); // three %d, three argsArgs are pushed in reverse (cdecl), so arg1 lands lowest
Why: The caller pushes z, then y, then x, then the format-string pointer last — so reading UP from rip+8 gives x, then y, then z, in source order.
Each %d consumes the next 4-byte slot, walking up
Why: First %d reads rip+8 (x), second reads rip+12 (y), third reads rip+16 (z). The argptr advances 4 bytes per modifier.
| Modifier | Reads slot | Value found | Prints |
|---|---|---|---|
| 1st %d | rip + 8 | x = 11 | 11 |
| 2nd %d | rip + 12 | y = 22 | 22 |
| 3rd %d | rip + 16 | z = 33 | 33 |
Output: x=11 y=22 z=33. Every modifier had a matching argument, so every read landed on real data.
Pattern
Step through it
Step through Worked: printf("x=%d y=%d z=%d", x, y, z) one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
The compiler never complained — the call is legal. The mismatch is a silent information leak (CWE-134).
Commit before you compute: what does Worked: the MISMATCH — three %d, only two args come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The third %d keeps walking — past the end of the arguments
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. There is no third argument, so printf reads rip+16: whatever 4 bytes happen to sit there (a saved register, a leftover local, another frame).
Worked example
int x = 11, y = 22;
printf("x=%d y=%d z=%d", x, y); // THREE %d, only TWO args!First two %d read the real args
Why: rip+8 gives x, rip+12 gives y — both supplied by the caller, both correct.
The third %d keeps walking — past the end of the arguments
Why: There is no third argument, so printf reads rip+16: whatever 4 bytes happen to sit there (a saved register, a leftover local, another frame). It leaks unrelated stack data.
| Modifier | Reads slot | Holds | Prints |
|---|---|---|---|
| 1st %d | rip + 8 | x = 11 (real arg) | 11 |
| 2nd %d | rip + 12 | y = 22 (real arg) | 22 |
| 3rd %d | rip + 16 | leftover stack word | garbage / leaked value |
The compiler never complained — the call is legal. The mismatch is a silent information leak (CWE-134).
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
The third %d keeps walking — past the end of the arguments
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The compiler never complained — the call is legal. The mismatch is a silent information leak (CWE-134).
Concept
%x prints the next stack word as hex. A string of %x %x %x simply walks up the stack, leaking one word per modifier. %08x zero-pads each to 8 hex digits so the dump lines up.
// attacker sends this AS the format string:
printf("%08x %08x %08x %08x"); // == printf(user_input)| Modifier # | Reads slot | Leaks |
|---|---|---|
| %08x #1 | rip + 8 | stack word 1 (hex) |
| %08x #2 | rip + 12 | stack word 2 (hex) |
| %08x #3 | rip + 16 | stack word 3 (hex) |
| %08x #4 | rip + 20 | stack word 4 (hex) |
No matching arguments exist, so every %x dumps live stack memory — saved registers, return addresses, canaries, pointers. This is how a format-string bug defeats ASLR (Lesson 15): it reads back a real runtime address.
Explain it
Discussion prompt
Explain %x — the read/leak primitive to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
%x prints the next stack word as hex. A string of %x %x %x simply walks up the stack, leaking one word per modifier. %08x zero-pads each to 8 hex digits so the dump lines up.
Concept
%n does not print anything. It writes the number of bytes printed so far into the address held by the next argument slot — turning a read bug into an arbitrary memory write.
%n — Stores the count of characters output so far into an int* taken from the next argument slot. Attacker-controlled width specifiers control the value written; attacker-placed pointers control where.
// CONCEPTUAL / TOY only — do not build a real exploit
int count;
printf("hello%n", &count); // count <- 5 ("hello" is 5 chars)| Modifier | Action | Effect |
|---|---|---|
| %x | read a stack word | leak (information disclosure) |
| %n | write the running count | arbitrary write (memory corruption) |
| %100d%n | pad to 100, then write | write the value 100 to a target |
Because the attacker controls both the address slot (by walking the stack) and the value (via field-width padding), %n is a write-what-where primitive — the seed of a control-flow hijack, kept strictly conceptual here.
Definition probe
Sort into buckets
Every line below is part of the definition of variadic function or of %n — one or the other, never both. Put each where it belongs.
...).; The callee, not the compiler, decides how many arguments to pull off the stack.int* taken from the next argument slot.; Attacker-controlled width specifiers control the value written; attacker-placed pointers control where....). The callee, not the compiler, decides how many arguments to pull off the stack.int* taken from the next argument slot. Attacker-controlled width specifiers control the value written; attacker-placed pointers control where.Intuition
Walking the stack %x by %x is tedious. The $ notation (%7$x) lets an attacker jump directly to the 7th argument slot without printing the six before it — like an array index into the stack.
This is why a format-string bug is so powerful: once the attacker knows which slot holds a target pointer, %N$n writes to exactly that slot. The stack walk is no longer linear and clumsy — it is random access into memory.
| Modifier | Reads / writes slot | Equivalent to |
|---|---|---|
| %x | next slot, then advance | linear walk |
| %7$x | the 7th arg slot directly | stack[7] read |
| %7$n | the 7th arg slot directly | stack[7] write |
Estimation
Predict first
The leaked return address is a real runtime pointer — exactly what an attacker needs to bypass ASLR before the L8-style hijack.
Commit before you compute: what does Worked: leaking 4 stack words with %08x come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Eventually the walk reaches the attacker's own buffer
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. If the buffer holding the format string is itself on the stack above rip+8, one of the %08x prints 41414141 — proving the attacker can see (and later, with %n, place) controlled bytes among the read slots.
Worked example
// attacker input becomes the format string:
// printf("AAAA %08x %08x %08x %08x")
// 'AAAA' = 0x41414141 sits in the buffer on the stackThe literal bytes print first
Why: AAAA is emitted verbatim; the four %08x then each read one stack word, walking up from rip+8.
Eventually the walk reaches the attacker's own buffer
Why: If the buffer holding the format string is itself on the stack above rip+8, one of the %08x prints 41414141 — proving the attacker can see (and later, with %n, place) controlled bytes among the read slots.
| Output field | Reads slot | Typical value |
|---|---|---|
| %08x #1 | rip + 8 | a saved register, e.g. 0xb7fc0000 |
| %08x #2 | rip + 12 | a stack/canary word |
| %08x #3 | rip + 16 | a return address (defeats ASLR) |
| %08x #4 | rip + 20 | 41414141 (the buffer itself) |
The leaked return address is a real runtime pointer — exactly what an attacker needs to bypass ASLR before the L8-style hijack.
Error analysis
Annotate
Walk the callouts on Worked: leaking 4 stack words with %08x. Each one is a place this is easy to get subtly wrong.
AAAA is emitted verbatim; the four %08x then each read one stack word, walking up from rip+8.%08x prints 41414141 — proving the attacker can see (and later, with %n, place) controlled bytes among the read slots.Concept
Combine the two halves: %x finds a live address (the rip slot, defeating ASLR), and %n writes to an address the attacker placed in the buffer. Read then write = a full hijack chain — all from one printf(buf).
| Phase | Modifier | Gains the attacker |
|---|---|---|
| recon | %x / %08x | a real stack/code address (defeats ASLR) |
| target | %N$x | confirms which slot holds a chosen pointer |
| write | %N$n | stores a value at that pointer |
Kept strictly conceptual — the takeaway is that a single uncontrolled format string yields BOTH disclosure and corruption, which is why CWE-134 is rated so severely.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Printing a user-supplied string buf.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: If buf contains %x or %n, printf interprets them as a tiny program and walks the stack — leaking or writing memory.
Printing a user-supplied string buf.
Why: If buf contains %x or %n, printf interprets them as a tiny program and walks the stack — leaking or writing memory. The string's content is now executable instructions to printf.
Trap
Printing a user-supplied string buf.
Write printf(buf) because "it's just text"
Why: If buf contains %x or %n, printf interprets them as a tiny program and walks the stack — leaking or writing memory. The string's content is now executable instructions to printf.
Printing a user-supplied string buf.
Write printf("%s", buf) — buf is DATA, never the format
Why: The literal "%s" is the only format string; buf is consumed as a plain argument. Any % inside buf is printed verbatim, not interpreted. (Also valid: fputs(buf, stdout) / puts(buf).)
Trap
Is printf("%d %d", x) a compile error?
Assume the compiler rejects too-few-args
Why: It does not. printf is variadic — taking one extra-or-fewer argument is a legal call. The second %d silently reads off the end of the arguments at runtime.
Is printf("%d %d", x) a compile error?
Treat it as a runtime bug the language permits
Why: The C standard allows variadic calls; the count mismatch is invisible at compile time. The extra %d reads uninitialized stack data — a leak. Only an opt-in heuristic (-Wformat) might warn.
Anomaly
Predict first
A student writes this, and it looks reasonable:
What does a chain of %x modifiers risk?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confuses a READ with a WRITE. %x only discloses memory; if you think they're equivalent you'll miss that %n CORRUPTS memory — a far more severe primitive.
What does a chain of %x modifiers risk?
Why: Confuses a READ with a WRITE. %x only discloses memory; if you think they're equivalent you'll miss that %n CORRUPTS memory — a far more severe primitive.
Trap
What does a chain of %x modifiers risk?
Lump %x and %n together as 'reading the stack'
Why: Confuses a READ with a WRITE. %x only discloses memory; if you think they're equivalent you'll miss that %n CORRUPTS memory — a far more severe primitive.
What does a chain of %x modifiers risk?
Separate %x (read/leak) from %n (write/corrupt)
Why: %x walks the stack and leaks words; %n writes the running byte count into a pointer slot. One is information disclosure, the other is an arbitrary memory write — different CWE consequences entirely.
Section
Part B · §3.4
Concept
A 32-bit word is just bits. Whether it means a signed value (two's complement) or an unsigned value depends only on the type the code uses to read it. The bug is when one piece of code disagrees with another about which interpretation applies.
Callback to L4: the all-ones word 0xFFFFFFFF is -1 read as int, but 4294967295 (~4.29e9) read as unsigned.
| Bits | as signed int | as unsigned (size_t) |
|---|---|---|
| 0x00000005 | 5 | 5 |
| 0xFFFFFFFF | -1 | 4294967295 |
| 0x80000000 | -2147483648 | 2147483648 |
Pattern
Step through it
Step through Same bits, two meanings one row at a time. What is driving the change, and what would the row after the last one be?
Concept
When C compares a signed and an unsigned operand of the same width, the signed one is converted to unsigned first. So (int)-1 < (unsigned)10 is evaluated as 4294967295u < 10u — which is false. The negative number wins by becoming gigantic.
int s = -1;
unsigned u = 10;
if (s < u) /* FALSE: -1 -> 4294967295u, not < 10 */
puts("reached?"); // never prints| Comparison | Promoted to | Result |
|---|---|---|
| (int)-1 < (unsigned)10 | 4294967295u < 10u | false (surprising) |
| (int)5 < (unsigned)10 | 5u < 10u | true |
| len < sizeof(buf) | if len<0: huge < small | false -> guard skipped |
This is exactly why a if (len < bufsize) guard fails for negative len: the comparison silently promotes len to a huge unsigned value.
Intuition
Lengths and sizes in C library calls are usually size_t, which is unsigned. If your own code carries a length as a signed int and an attacker makes it negative, the moment it crosses into a size_t it flips to a huge positive number.
So a check like if (len < MAX) happily passes for a 'small' negative len — but the very same len, handed to memcpy, is interpreted as billions of bytes. The guard and the copy disagree about what len means.
| len (signed view) | Passes if (len < MAX)? | len at memcpy (unsigned view) |
|---|---|---|
| -1 | yes (-1 < MAX) | 4294967295 bytes |
| -100 | yes | ~4.29 billion bytes |
Comparison
Comparison matrix
From Intuition — a negative length becomes 'enormous': refill the Passes if (len < MAX)? column from what you know. The rest of the table is as it appeared.
| len (signed view) | Passes if (len < MAX)? | len at memcpy (unsigned view) |
|---|---|---|
| -1 | yes (-1 < MAX) | 4294967295 bytes |
| -100 | yes | ~4.29 billion bytes |
Ranking
Put in order
Put the moves of Worked: signed length bypasses the bounds check into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. As a signed int, -1 is clearly not > 64, so the guard if (len > 64) return; does NOT trigger.
Worked example
void copy_into(char *src, int len) { // len is SIGNED
char buf[64];
if (len > 64) return; // guard: reject too-big
memcpy(buf, src, len); // memcpy wants size_t (unsigned)
}
// attacker calls copy_into(src, -1);Attacker passes len = -1
Why: As a signed int, -1 is clearly not > 64, so the guard if (len > 64) return; does NOT trigger. The check is satisfied.
memcpy reinterprets len as size_t
Why: memcpy's third parameter is size_t (unsigned). The bits 0xFFFFFFFF that meant -1 now mean 4294967295 — memcpy tries to copy ~4.29 GB into a 64-byte buffer.
| Step | len as int | len as size_t | Outcome |
|---|---|---|---|
| guard if (len > 64) | -1 | — | -1 > 64 is false -> passes |
| memcpy(buf, src, len) | — | 4294967295 | massive overflow of buf[64] |
Verify the fix: make len unsigned (size_t) end-to-end
Why: If len were declared size_t, then -1 could never be passed as a 'small' value — it would already be 4294967295 at the guard, and if (len > 64) return; would correctly reject it before memcpy.
Blank canvas
Draw it
Draw what Worked: signed length bypasses the bounds check just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Intuition
A 32-bit unsigned value is an odometer with 2^32 positions. Add past the top and it silently rolls back to 0 — there is no error, no exception, just a small number where a huge one belonged.
So count * size for big enough operands produces a tiny product. malloc happily returns a tiny buffer; the code 'knows' it asked for room for count items and writes all of them. The allocator and the loop disagree about how big the request was.
| True product | 2^32 wraps to | Gap the loop overflows |
|---|---|---|
| 0x100000010 | 0x10 (16) | everything past 16 bytes |
| 0x100000000 | 0x0 (0) | the entire write |
Concept
malloc(count * size) computes the product in fixed-width arithmetic. If count * size overflows past 2^32, it wraps to a small number — so malloc returns a tiny buffer, but the following loop still writes count full elements.
// 32-bit: count * size can wrap around 2^32
void *p = malloc(count * size);
for (unsigned i = 0; i < count; i++)
p[i] = read_element(); // writes count elements -> heap overflow| count | size | count * size (true) | wrapped (mod 2^32) | malloc gets |
|---|---|---|---|---|
| 0x10000001 | 16 | 0x100000010 | 0x10 | 16 bytes |
| 65536 | 65536 | 0x100000000 | 0 | 0 bytes |
malloc allocates the wrapped (tiny) size; the loop writes the true (huge) count of elements. The surplus spills past the allocation — a heap overflow (CWE-190) that feeds the heap exploits of Lesson 10.
Step zero
Discussion prompt
**Worked: count * size overflow, step by step** — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Compute the true product
Answer:
Worked example
unsigned count = 0x10000001; // attacker-chosen
unsigned size = 16;
int *p = malloc(count * size); // 0x10000001 * 16
for (unsigned i = 0; i < count; i++)
p[i] = 0; // writes count intsCompute the true product
Why: 0x10000001 * 16 = 0x100000010 — a 33-bit result that does not fit in 32 bits.
32-bit arithmetic keeps only the low 32 bits
Why: 0x100000010 mod 2^32 = 0x10 = 16. malloc allocates just 16 bytes — room for 4 ints, not 0x10000001 of them.
The loop writes count elements anyway
Why: i runs 0 .. 0x10000000, writing ~268 million ints into a 16-byte allocation — a massive heap overflow (CWE-190 feeding CWE-787).
| Quantity | Value | Note |
|---|---|---|
| count * size (true) | 0x100000010 | 33 bits |
| count * size (stored) | 0x10 = 16 | wrapped to 32 bits |
| malloc returns | 16 bytes | fits 4 ints |
| loop writes | 0x10000001 ints | billions past the buffer |
Verify the fix: check for overflow before multiplying
Why: Guard with if (count != 0 && size > SIZE_MAX / count) return NULL; (or use calloc, which performs this check internally) so a wrapping product can never reach malloc.
Concept
Attackers rarely type '-1'. Negative or huge lengths arise naturally: a header field read as a signed int, a subtraction like end - start that goes negative, or atoi on attacker text returning a negative number — all then flow into a size_t parameter.
int len = end - start; // if start > end, len < 0
if (len > MAX) return; // negative passes
memcpy(dst, src, len); // len -> huge size_t| Source of len | Attacker makes it… | At size_t becomes |
|---|---|---|
| int field from input | -1 | 4294967295 |
| end - start | negative (start > end) | huge positive |
| atoi(user_text) | -100 | ~4.29e9 |
Concept
Assigning a wider integer into a narrower type drops the high bytes. A size_t or int stored into a short (2 bytes) or char (1 byte) keeps only the low bits — the validated 'big' value becomes a small one the rest of the code trusts.
size_t big = 0x10000; // 65536
short small = big; // short is 2 bytes
// small now holds 0x0000 = 0 (the 0x1 high byte is dropped)| Source value | Source type | Target type | Stored value |
|---|---|---|---|
| 0x10000 (65536) | size_t (4B) | short (2B) | 0x0000 = 0 |
| 0x1FF (511) | int (4B) | char (1B) | 0xFF = -1 / 255 |
| 0x12345678 | int (4B) | short (2B) | 0x5678 |
A length checked as a large size_t can be truncated to a tiny short used for allocation, while the original large value still drives the copy — another size mismatch.
Explain it
Discussion prompt
Explain Truncation on assignment to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A length checked as a large size_t can be truncated to a tiny short used for allocation, while the original large value still drives the copy — another size mismatch.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Bounding an allocation by element count.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A count below MAX can still make count * size overflow when size is large — the product wraps, malloc under-allocates, and the loop overruns.
Bounding an allocation by element count.
Why: A count below MAX can still make count * size overflow when size is large — the product wraps, malloc under-allocates, and the loop overruns. Checking count is not checking the PRODUCT.
Trap
Bounding an allocation by element count.
Validate count alone, then malloc(count * size)
Why: A count below MAX can still make count * size overflow when size is large — the product wraps, malloc under-allocates, and the loop overruns. Checking count is not checking the PRODUCT.
Bounding an allocation by element count.
Check the multiplication itself: if (size && count > SIZE_MAX/size) fail;
Why: Verify the product cannot exceed SIZE_MAX before multiplying — or call calloc(count, size), which is contractually required to detect this overflow and return NULL.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Verify the product cannot exceed SIZE_MAX before multiplying — or call calloc(count, size), which is contractually required to detect this overflow and return NULL.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
A count below MAX can still make count * size overflow when size is large — the product wraps, malloc under-allocates, and the loop overruns. Checking count is not checking the PRODUCT.
Trap
Bounds-checking a length before a copy.
Carry the length as a signed int and test if (len > MAX) return;
Why: A negative len sails through the signed comparison, then becomes a huge unsigned value at memcpy/malloc. The guard and the library disagree about the type, so the check is meaningless.
Bounds-checking a length before a copy.
Use size_t for the length AND check both ends: if (len > MAX) return;
Why: With an unsigned size_t, there are no negatives to slip past, so the comparison agrees with what memcpy will do. Match the signedness of your length to the API that consumes it.
Section
Part C · §3.5
Concept
A fence-post (off-by-one) bug writes exactly one element too many — almost always a <= where it should be <. A buffer of N slots has valid indices 0…N−1; index N is one past the end.
char a[N];
for (int i = 0; i <= N; i++) // BUG: should be i < N
a[i] = src[i]; // a[N] writes one past the buffer| N | Valid indices | Loop with i <= N writes | Overflow? |
|---|---|---|---|
| 8 | 0..7 | 0..8 (nine writes) | a[8] is out of bounds |
| 64 | 0..63 | 0..64 | a[64] is out of bounds |
One extra byte sounds harmless — the next slide shows why, on the stack, that single byte can be exploitable (CWE-193).
Analogy
Discussion prompt
Explain The fence-post error by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A fence-post (off-by-one) bug writes exactly one element too many — almost always a <= where it should be <. A buffer of N slots has valid indices 0…N−1; index N is one past the end.
Intuition
A fence with N panels has N+1 posts. Mixing up 'panels' and 'posts' is the original off-by-one. An array a[N] has N slots but its highest valid index is N−1 — the index is a post, the count is a panel.
So the safe loop bound is i < N, never i <= N. The <= counts one post too many and writes a[N], the slot that does not exist.
| Array size N | Valid indices | Count of valid indices | First invalid |
|---|---|---|---|
| 4 | 0,1,2,3 | 4 | a[4] |
| 8 | 0..7 | 8 | a[8] |
Ranking
Put in order
Put the moves of Worked: the classic <= fence-post trace into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Four valid writes a[0]..a[3] copy 'W','X','Y','Z' — exactly the 4-byte buffer.
Worked example
char a[4];
char src[] = "WXYZ"; // 4 chars
for (int i = 0; i <= 4; i++) // BUG: <= instead of <
a[i] = src[i];Iterations i = 0..3 fill the buffer
Why: Four valid writes a[0]..a[3] copy 'W','X','Y','Z' — exactly the 4-byte buffer.
Iteration i = 4 runs because the test is <=
Why: 4 <= 4 is true, so the loop executes once more, writing a[4] — one slot beyond the buffer (and reading src[4], the '\0').
| i | i <= 4 ? | Writes a[i] | In bounds? |
|---|---|---|---|
| 0 | true | 'W' | yes |
| 1 | true | 'X' | yes |
| 2 | true | 'Y' | yes |
| 3 | true | 'Z' | yes |
| 4 | true | '\0' | NO — one past end |
| 5 | false | — | loop ends |
Verify the fix: use i < 4 (strictly less)
Why: With i < 4 the loop stops after a[3], writing exactly the 4 valid slots. The valid index range for a[N] is 0..N-1, so the loop bound must be < N.
Sorting
Sort into buckets
These are the pieces of L09 · Format Strings, Integer Conversion & Off-by-One Bugs, out of order. Put each one back under the part of the lesson it belongs to.
Intuition
Locals sit below ebp, and the saved frame pointer (sfp) sits at ebp+0 — right above the local buffer. A buffer that ends just below the sfp has its first out-of-bounds byte land exactly on the sfp's lowest address byte.
Because x86 is little-endian, that lowest-address byte is the least-significant byte of the sfp. Overwriting it shifts the saved frame pointer by up to 255 bytes — and when the caller restores ebp from that corrupted sfp, the next function's whole frame is relocated to attacker-influenced ground.
| Stack slot (high -> low) | Offset | One byte over buf hits… |
|---|---|---|
| saved rip | ebp + 4 | — |
| saved ebp (sfp) | ebp + 0 | its least-significant byte |
| char buf[...] (top) | ebp - 1 | last valid byte |
Step zero
Discussion prompt
Worked: the single-null-byte overflow — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: buf[0..7] fill the buffer correctly
Answer:
Worked example
void f(char *src) {
char buf[8]; // buf occupies ebp-8 .. ebp-1
for (int i = 0; i <= 8; i++) // off-by-one: writes buf[0..8]
buf[i] = src[i]; // buf[8] is at ebp+0 = sfp low byte
}buf[0..7] fill the buffer correctly
Why: Eight valid writes cover ebp-8 through ebp-1 — the whole 8-byte buffer.
buf[8] writes ONE byte past the buffer
Why: buf[8] lands at ebp+0, the saved frame pointer. If src[8] is a NUL (0x00), it zeroes the sfp's least-significant byte.
The corrupted sfp shifts the caller's frame
Why: On leave/ret, the caller pops the tampered sfp into ebp. Its low byte changed by up to 255, so the caller's frame pointer now points to a different stack location — one the attacker may control, making the single byte exploitable.
| Write | Address | Slot | Effect |
|---|---|---|---|
| buf[0..7] | ebp-8 .. ebp-1 | the buffer | normal, in bounds |
| buf[8] | ebp + 0 | sfp low byte | one-byte overwrite of saved ebp |
| caller leave/ret | — | ebp <- sfp | frame relocated by up to 255 bytes |
Blank canvas
Draw it
Draw what Worked: the single-null-byte overflow just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
Not every off-by-one is a <= loop. A classic is forgetting the +1 for the terminator: allocating strlen(s) bytes instead of strlen(s) + 1, then strcpy-ing — the '\0' lands one byte past the allocation.
char *d = malloc(strlen(s)); // BUG: missing + 1 for '\0'
strcpy(d, s); // writes strlen(s)+1 bytes| strlen(s) | malloc gets | strcpy writes | Overflow? |
|---|---|---|---|
| 5 | 5 bytes | 6 bytes (incl. '\0') | 1 byte over |
| 20 | 20 bytes | 21 bytes | 1 byte over |
Fix: malloc(strlen(s) + 1). The terminator is real storage that the size arithmetic must account for — a single forgotten +1 is a heap off-by-one (CWE-193).
Counterexample
Discussion prompt
Fix: malloc(strlen(s) + 1). The terminator is real storage that the size arithmetic must account for — a single forgotten +1 is a heap off-by-one (CWE-193).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
strncpy(dst, src, n) copies at most n bytes. If src is shorter than n, it pads the rest with NUL — terminated. But if src is exactly n bytes or longer, strncpy copies n bytes and stops — with no terminating NUL.
char buf[8];
strncpy(buf, src, sizeof(buf)); // n = 8
// if strlen(src) >= 8, buf has NO '\0' terminator
printf("%s", buf); // reads PAST buf until it finds a 0| strlen(src) | n | Bytes copied | Null-terminated? |
|---|---|---|---|
| 3 | 8 | 3 + 5 NUL padding | yes — padded |
| 7 | 8 | 7 + 1 NUL | yes — fits with terminator |
| 8 | 8 | 8 (fills buf) | NO — no room for '\0' |
| 20 | 8 | 8 (truncated) | NO — no '\0' |
An un-terminated string is a time bomb: the next strlen/%s/strcpy keeps reading past buf until it stumbles on a zero byte — leaking or corrupting adjacent memory (CWE-193 / out-of-bounds read).
Trade off
Comparison matrix
From strncpy does NOT always null-terminate: every row here is a choice with a cost. Fill the n column, then say which row you would actually pick and what you give up for it.
| strlen(src) | n | Bytes copied | Null-terminated? |
|---|---|---|---|
| 3 | 8 | 3 + 5 NUL padding | yes — padded |
| 7 | 8 | 7 + 1 NUL | yes — fits with terminator |
| 8 | 8 | 8 (fills buf) | NO — no room for '\0' |
| 20 | 8 | 8 (truncated) | NO — no '\0' |
Intuition
Read strncpy(dst, src, n) as 'spend a budget of n bytes'. If the source runs out early, the leftover budget is spent on NUL padding. If the budget runs out first, copying just stops — and a terminator was never in the budget.
The '\0' is a side effect of padding, not a guarantee of the function. People remember 'strncpy is the safe strcpy' and forget the safety only holds when the source is shorter than the buffer.
| Relationship | Budget spent on | Terminated? |
|---|---|---|
| strlen(src) < n | chars + NUL padding | yes |
| strlen(src) >= n | chars only, no room left | NO |
Comparison
Comparison matrix
From Intuition — strncpy is a copy budget, not a terminator: refill the Budget spent on column from what you know. The rest of the table is as it appeared.
| Relationship | Budget spent on | Terminated? |
|---|---|---|
| strlen(src) < n | chars + NUL padding | yes |
| strlen(src) >= n | chars only, no room left | NO |
Ranking
Put in order
Put the moves of Worked: strncpy when src shorter vs src == bufsize into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. strncpy copies 'h','i' then pads the remaining 2 bytes with NUL.
Worked example
char buf[4];
// case 1: src = "hi" (len 2)
strncpy(buf, "hi", 4); // -> 'h','i','\0','\0'
// case 2: src = "data" (len 4)
strncpy(buf, "data", 4); // -> 'd','a','t','a' (NO terminator!)Case 1: src (2) is shorter than n (4)
Why: strncpy copies 'h','i' then pads the remaining 2 bytes with NUL. buf is a proper C string.
Case 2: src (4) equals n (4)
Why: strncpy copies all 4 chars and stops — every byte of buf is used, leaving NO room for '\0'. buf is NOT a valid C string.
| src | buf[0] | buf[1] | buf[2] | buf[3] | Terminated? |
|---|---|---|---|---|---|
| "hi" | h | i | \0 | \0 | yes |
| "data" | d | a | t | a | NO |
Verify the safe pattern: reserve a byte and terminate by hand
Why: Use strncpy(buf, src, sizeof(buf)-1); buf[sizeof(buf)-1] = '\0'; — copy at most n-1 and force the last byte to NUL, so buf is always a valid string regardless of src's length.
Error analysis
Annotate
Walk the callouts on Worked: strncpy when src shorter vs src == bufsize. Each one is a place this is easy to get subtly wrong.
strncpy(buf, src, sizeof(buf)-1); buf[sizeof(buf)-1] = '\0'; — copy at most n-1 and force the last byte to NUL, so buf is always a valid string regardless of src's length.Anomaly
Predict first
A student writes this, and it looks reasonable:
Using strncpy(buf, src, sizeof(buf)) for safety.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: When strlen(src) >= sizeof(buf), strncpy fills the whole buffer and writes NO terminator.
Using strncpy(buf, src, sizeof(buf)) for safety.
Why: When strlen(src) >= sizeof(buf), strncpy fills the whole buffer and writes NO terminator. The next %s/strlen runs off the end — an out-of-bounds read.
Trap
Using strncpy(buf, src, sizeof(buf)) for safety.
Assume the result is always a valid C string
Why: When strlen(src) >= sizeof(buf), strncpy fills the whole buffer and writes NO terminator. The next %s/strlen runs off the end — an out-of-bounds read.
Using strncpy(buf, src, sizeof(buf)) for safety.
Copy sizeof(buf)-1 and set the last byte to '\0' yourself
Why: strncpy(buf, src, sizeof(buf)-1); buf[sizeof(buf)-1] = '\0'; guarantees termination for any src. strncpy's contract pads only when src is SHORTER than n — never rely on it otherwise.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sizing the risk of a single out-of-bounds byte.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Ignores WHERE that byte lands. One byte past a stack buffer overwrites the least-significant byte of the saved frame pointer — relocating the caller's whole frame by up to 255 bytes.
Sizing the risk of a single out-of-bounds byte.
Why: Ignores WHERE that byte lands. One byte past a stack buffer overwrites the least-significant byte of the saved frame pointer — relocating the caller's whole frame by up to 255 bytes. Quantity isn't the threat; position is.
Trap
Sizing the risk of a single out-of-bounds byte.
Dismiss a one-byte overflow as cosmetic
Why: Ignores WHERE that byte lands. One byte past a stack buffer overwrites the least-significant byte of the saved frame pointer — relocating the caller's whole frame by up to 255 bytes. Quantity isn't the threat; position is.
Sizing the risk of a single out-of-bounds byte.
Treat the location as the threat: it hits the sfp low byte
Why: A single null byte on the sfp shifts the next frame pointer, which the attacker can aim. Off-by-one is a recognized exploitable class (CWE-193), not a rounding error.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
% is an instruction: go grab the next argument off the stack and print it this way.; The whole bug is which slot the user input lands in: as data behind a %s it is harmless; as the format string itself it is interpreted.buf.; Is printf("%d %d", x) a compile error?Constraint
Discussion prompt
Run Spotting all three classes — a checklist with this step confiscated:
Multiplication: does malloc(count*size) or similar risk wrapping past 2^32? Check for overflow before allocating.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
printf(buf), fprintf, syslog(buf))? Force it into a %s argument.%x leaks the stack, %n writes it — argptr starts at rip+8 and walks up 4 bytes per modifier.int then meet an unsigned size_t API? Make it size_t end-to-end.malloc(count*size) or similar risk wrapping past 2^32? Check for overflow before allocating.short/char and then trusted? Watch the dropped high bytes.<= where it should be <? One byte past a stack buffer hits the sfp low byte.strncpy/snprintf actually leave a '\0'? Reserve a byte and set it by hand.Pattern
printf(buf), fprintf, syslog(buf))? Force it into a %s argument.%x leaks the stack, %n writes it — argptr starts at rip+8 and walks up 4 bytes per modifier.int then meet an unsigned size_t API? Make it size_t end-to-end.malloc(count*size) or similar risk wrapping past 2^32? Check for overflow before allocating.short/char and then trusted? Watch the dropped high bytes.<= where it should be <? One byte past a stack buffer hits the sfp low byte.strncpy/snprintf actually leave a '\0'? Reserve a byte and set it by hand.Edge cases
Discussion prompt
Spotting all three classes — a checklist works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
printf(buf), fprintf, syslog(buf))? Force it into a %s argument.%x leaks the stack, %n writes it — argptr starts at rip+8 and walks up 4 bytes per modifier.int then meet an unsigned size_t API? Make it size_t end-to-end.malloc(count*size) or similar risk wrapping past 2^32? Check for overflow before allocating.short/char and then trusted? Watch the dropped high bytes.<= where it should be <? One byte past a stack buffer hits the sfp low byte.strncpy/snprintf actually leave a '\0'? Reserve a byte and set it by hand.Elimination
Eliminate the wrong options
Where does the SECOND %d read from, and what does it find?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The argptr starts at rip+8 (the first arg, x) and advances 4 bytes per modifier. The first %d reads rip+8 = x; the second %d reads rip+12, which holds no supplied argument — it is whatever leftover stack word sits there, leaked as an integer. This is a CWE-134 information disclosure.
Check
printf("%d %d", x) is called with only ONE argument supplied. Trace where each modifier reads before you click.
Check your understanding
Where does the SECOND %d read from, and what does it find?
Answer: A
Why: The argptr starts at rip+8 (the first arg, x) and advances 4 bytes per modifier. The first %d reads rip+8 = x; the second %d reads rip+12, which holds no supplied argument — it is whatever leftover stack word sits there, leaked as an integer. This is a CWE-134 information disclosure.
Concept
%x/%n, printf interprets it; pass user data as a %s argument.%x reads/leaks a stack word; %n WRITES the running count to a pointer.memcpy/malloc.strlen(src) >= n; reserve a byte and terminate by hand.Concept
%x defeats ASLR, L15) and WRITE (%n is write-what-where).%n writes to an arbitrary address, skipping straight over the canary instead of overwriting through it.count*size under-allocates, then the loop writes the true count past the allocation.Concept
%x stack-read and %n write write-up.printf(buf) with -Wformat-security, feed it %08x %08x %08x, and watch live stack words appear (toy, authorized targets only).Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Format Strings · Integer Conversion · Off-by-One. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can explain why printf(user_input) is exploitable, trace printf's argptr from rip+8 upward, tell %x (leak) from %n (write), show how a signed length and a wrapped count*size corrupt memory, and locate the off-by-one byte on the saved frame pointer.
| Vuln class | Core mechanism | CWE |
|---|---|---|
| Format string | user data AS format; argptr walks rip+8 up; %x leaks, %n writes | CWE-134 |
| Integer conversion | signed len -> huge size_t; count*size wraps; truncation drops high bytes | CWE-190 |
| Off-by-one | <= writes one past; single byte hits sfp low byte; strncpy may not terminate | CWE-193 |
Want this taught 1-on-1? Alexander tutors Computer Security — $55/session, free consultation.