CS 161, Lesson 8, in 50 slides and code mode. It explains why C has no bounds checking, then works the vulnerable gets(buf) frame, the 8 + 4 = 12-byte offset to the saved rip, and how an address is typed in little-endian order. It covers three shellcode-placement strategies, NOP sleds, the history of the worm, and the fgets fix. The examples are toys running in a sandbox, and it is anchored to textbook sections 3.1 to 3.2.
Subject: Computer Security · 88 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
CS 161 · Lesson 8 of 45
buffer overflows · gets(buf) · the 12-byte offset · little-endian rip · shellcode · NOP sleds
Objectives
char buf[8] and compute the 8 + 4 = 12-byte offset from buf to the saved return address.gets with fgets(buf, sizeof(buf), stdin) and say why it stops the overflow.Warm-up
Discussion prompt
Before we open L08 · Buffer Overflows & Stack Smashing: without looking back, what was the main idea of L07 · Reading x86 Assembly + GDB (and the C→memory bridge), and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
CS 161 Lesson 7 (50 slides, code mode): read gcc -S -O0 output, recognize prologue/body/epilogue, decode MOV/LEA/ADD/SUB/CMP/JE-JNE-JMP and AT&T addressing modes, annotate assembly instruction-by-instruction, and step through it in gdb (stepi, info registers, x/8xw $esp) — closing on the unbounded C strcpy copy loop that sets up the Lesson 8 overflow. Multiple full trace tables.
Concept
Every example here is a toy in a sandbox, for authorized security education — exactly how the textbook teaches it. You learn the attack so you can recognize and fix the bug.
We never write real shellcode bytes; the payload's machine code stays abstract as [shellcode]. The skill being built is reading a frame and spotting the overflow.
Counterexample
Discussion prompt
Every example here is a toy in a sandbox, for authorized security education — exactly how the textbook teaches it. You learn the attack so you can recognize and fix the bug.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
We never write real shellcode bytes; the payload's machine code stays abstract as [shellcode]. The skill being built is reading a frame and spotting the overflow.
Section
Part 1 · §3.1
Concept
Declaring char buffer[4] reserves four bytes — but C does nothing to stop you reading or writing buffer[5]. The access just runs off the end into whatever sits next in memory.
buffer overflow — Writing past the end of a fixed-size buffer, corrupting adjacent memory the program never meant you to touch.
In a memory-safe language an out-of-range index throws; in C it is silent and the write simply lands somewhere it shouldn't.
Analogy
Discussion prompt
Explain C does not check array bounds by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Declaring char buffer[4] reserves four bytes — but C does nothing to stop you reading or writing buffer[5]. The access just runs off the end into whatever sits next in memory.
Concept
An out-of-bounds write is only dangerous because of what sits next to the buffer. On the stack, that neighbor is control-flow metadata — the sfp and the rip.
Overflow into another local is a bug; overflow into the rip is a takeover. The frame layout decides which one you get.
| Overflow reaches | Severity |
|---|---|
| another local variable | data corruption / logic bug |
| the saved sfp | frame confusion on return |
| the saved rip | control-flow hijack |
Comparison
Comparison matrix
From Where the overflow lands matters: refill the Severity column from what you know. The rest of the table is as it appeared.
| Overflow reaches | Severity |
|---|---|
| another local variable | data corruption / logic bug |
| the saved sfp | frame confusion on return |
| the saved rip | control-flow hijack |
Intuition
C is low-level: it exposes the bare machine, so a bounds check would be overhead the language refuses to add for you. Speed and control were the whole point.
C is also old — decades of legacy code (operating systems, network daemons, embedded firmware) are written in it and still running today.
So the programmer alone is responsible for every bound. Miss one, and the door is open.
Explain it
Discussion prompt
Explain Why C is like this to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
C is low-level: it exposes the bare machine, so a bounds check would be overhead the language refuses to add for you. Speed and control were the whole point.
Intuition
Think of char buf[8] as eight painted parking spots. C paints the lines but posts no attendant: nothing stops you from parking in spot 9, 10, 11 — they belong to the cars (variables) parked next door.
A safe language hires the attendant (a bounds check) who turns you away at spot 9. C hands you the keys and trusts you to count.
Concept
Buffer overflows are a property of manual memory management, not of C's syntax. C++ and Objective-C have the exact same exposure.
The cure is not a different C-family dialect — it is bounds-checked APIs and, ultimately, memory-safe languages (the Week 4 preview).
Pattern
Predict first
The table runs: buffer[0..3] | yes | writes inside the array · buffer[4] | no | writes 1 byte past — no error
In What 'no bounds check' looks like, given the rows so far: what is the next one — the row where Access is buffer[5]?
Correct: buffer[5] | no | writes 2 bytes past — no error
| Access | In bounds? | What C does |
|---|---|---|
| buffer[0..3] | yes | writes inside the array |
| buffer[4] | no | writes 1 byte past — no error |
| buffer[5] | no | writes 2 bytes past — no error |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. buffer[4] reserves exactly four bytes: buffer[0] through buffer[3].
Worked example
char buffer[4];
buffer[0] = 'h';
buffer[1] = 'i';
buffer[2] = '!';
buffer[3] = 0;
buffer[5] = 'X'; // legal C: no error, writes past the endIndices 0–3 are the only valid slots
Why: buffer[4] reserves exactly four bytes: buffer[0] through buffer[3].
buffer[5] writes two bytes past the array
Why: C computes the address buffer + 5 and stores 'X' there — into whatever variable or metadata happens to live at that address.
| Access | In bounds? | What C does |
|---|---|---|
| buffer[0..3] | yes | writes inside the array |
| buffer[4] | no | writes 1 byte past — no error |
| buffer[5] | no | writes 2 bytes past — no error |
Trade off
Comparison matrix
From What 'no bounds check' looks like: every row here is a choice with a cost. Fill the In bounds? column, then say which row you would actually pick and what you give up for it.
| Access | In bounds? | What C does |
|---|---|---|
| buffer[0..3] | yes | writes inside the array |
| buffer[4] | no | writes 1 byte past — no error |
| buffer[5] | no | writes 2 bytes past — no error |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Reading a line into a small buffer.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Sounds reasonable — but gets has no idea how big buf is.
Reading a line into a small buffer.
Why: Sounds reasonable — but gets has no idea how big buf is. It was never told the size.
Trap
Reading a line into a small buffer.
Assume gets(buf) stops once buf is full
Why: Sounds reasonable — but gets has no idea how big buf is. It was never told the size.
Reading a line into a small buffer.
gets writes until newline/EOF, however long the input
Why: §3.1: nothing — not the compiler, not the runtime, not gets — bounds-checks. A 100-byte input into char buf[8] overflows by 92 bytes. The programmer must pick a length-limited API.
Section
Part 2 · §3.2
Intuition
An attacker scans for two things: a fixed-size buffer and an unbounded write into it. vulnerable() hands over both — char buf[8] and gets(buf) — in three lines.
The defender's job is the mirror image: spot the same two ingredients and break the pairing before they ship.
Concept
void vulnerable() {
char buf[8];
gets(buf);
}Eight bytes of local buffer, filled by gets — which reads an unbounded line of user input. The whole vulnerability is in these three lines.
| Element | Role |
|---|---|
| char buf[8] | 8-byte local, attacker-fillable |
| gets(buf) | unbounded copy of user input into buf |
| (return) | pops the saved rip — the prize |
Picture it
Figure (svg): Stack frame from high to low: rip at ebp+4, sfp at ebp+0, then the 8-byte buf below. gets fills buf from low to high addresses, upward toward sfp and rip.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Once the prologue anchors ebp, the frame is fixed. From high to low addresses: the saved return address (rip), then the saved frame pointer (sfp), then buf.
Concept
Once the prologue anchors ebp, the frame is fixed. From high to low addresses: the saved return address (rip), then the saved frame pointer (sfp), then buf.
Figure (svg): Stack frame from high to low: rip at ebp+4, sfp at ebp+0, then the 8-byte buf below. gets fills buf from low to high addresses, upward toward sfp and rip.
| Slot (high→low) | Size | Offset from ebp |
|---|---|---|
| rip (saved return addr) | 4 B | ebp+4 |
| sfp (saved ebp) | 4 B | ebp+0 |
| buf[0..7] | 8 B | below ebp |
Ranking
Put in order
Put the moves of Tracing a too-long input byte by byte into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. buf[0..7] = the eight 'A's — exactly the declared size, still in bounds.
Worked example
Say the user types 16 characters into vulnerable(). Watch each group of bytes land as gets writes upward.
input = "AAAAAAAA" "BBBB" "CCCC" // 8 + 4 + 4 = 16 chars
// gets copies them buf[0], buf[1], ... upward, no limitFirst 8 bytes fill buf
Why: buf[0..7] = the eight 'A's — exactly the declared size, still in bounds.
Next 4 bytes overwrite the sfp
Why: The four 'B's land on the saved frame pointer at ebp+0 — already out of bounds, silently.
Last 4 bytes overwrite the rip
Why: The four 'C's overwrite the return address at ebp+4. On return the CPU jumps to 0x43434343 ('CCCC').
| Input bytes | Lands on | In bounds? |
|---|---|---|
| AAAAAAAA (0–7) | buf[0..7] | yes |
| BBBB (8–11) | sfp | no |
| CCCC (12–15) | rip | no — hijack |
Pattern
Step through it
Step through Tracing a too-long input byte by byte one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Two directions are in play and they are opposite. The stack grows down (each new frame sits at lower addresses). But gets fills buf upward: buf[0] is the lowest byte, buf[1] next, and so on toward higher addresses.
So a long input starts at buf[0] and marches up — through buf[7], then over the sfp, then over the rip sitting above it. The control-flow metadata is directly in the line of fire.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Which way does a long gets() input spread?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Confuses where new frames go (down) with the direction a single copy writes.
Which way does a long gets() input spread?
Why: Confuses where new frames go (down) with the direction a single copy writes. If writes went down, they'd never reach the rip — and the attack couldn't exist.
Trap
Which way does a long gets() input spread?
Conclude the overflow runs DOWN, away from rip
Why: Confuses where new frames go (down) with the direction a single copy writes. If writes went down, they'd never reach the rip — and the attack couldn't exist.
Which way does a long gets() input spread?
The copy writes UP: buf[0], buf[1], ... toward higher addresses
Why: §3.2: array indexing increases the address. buf[0] is lowest; extra bytes climb past buf[7] into sfp, then rip. Frame-growth direction and write direction are independent.
Section
Part 3 · §3.2
Concept
To reach the rip, the input must first fill all of buf (8 bytes) and then all of the sfp (4 bytes). That is 8 + 4 = 12 bytes of filler before the first byte of the rip.
The sfp is 4 bytes because this is 32-bit x86. (In 64-bit it would be 8 — a classic source of the offset-16 mistake.)
\[ \text{offset to rip} = \underbrace{8}_{\text{buf}} + \underbrace{4}_{\text{sfp}} = 12 \text{ bytes} \]
Intuition
It is tempting to stop at 8 — that's the buffer. But the sfp is parked between buf and the rip, and there is no way around it: input that climbs out of buf hits the sfp before it ever reaches the rip.
So you 'spend' 4 extra bytes overwriting the sfp just to get past it. 8 to fill, 4 to step over: 12 before the prize.
Step zero
Discussion prompt
Worked example: overwrite rip with 0xDEADBEEF — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: 12 garbage bytes first
Answer:
Worked example
Suppose (toy/sandbox) some attacker code already sits at address 0xDEADBEEF. We want vulnerable() to return straight into it.
payload = 'A' * 8 # fill buf[0..7]
+ 'A' * 4 # fill sfp (4 bytes, 32-bit)
+ '\xef\xbe\xad\xde' # rip = 0xDEADBEEF, little-endian12 garbage bytes first
Why: 8 to fill buf + 4 to fill the sfp. The value of these 12 bytes does not matter — they are just spacers to reach the rip.
Then the 4-byte address, LSB first
Why: x86 is little-endian, so 0xDEADBEEF is typed \xef\xbe\xad\xde — 0xEF first, 0xDE last.
On return, ret pops the overwritten rip into eip
Why: ret reads its target from the stack. We replaced that target with 0xDEADBEEF, so execution jumps there — the hijack.
| Input byte range | Overwrites | Resulting value |
|---|---|---|
| bytes 0–7 | buf[0..7] | 0x41 ('A') × 8 |
| bytes 8–11 | sfp | 0x41414141 |
| bytes 12–15 | rip | 0xDEADBEEF |
Error analysis
Annotate
Walk the callouts on Worked example: overwrite rip with 0xDEADBEEF. Each one is a place this is easy to get subtly wrong.
\xef\xbe\xad\xde — 0xEF first, 0xDE last.ret reads its target from the stack. We replaced that target with 0xDEADBEEF, so execution jumps there — the hijack.Trap
Encoding 0xDEADBEEF into the payload.
Write the bytes in reading order: \xde\xad\xbe\xef
Why: Looks like the number — but on a little-endian machine those bytes load as 0xEFBEADDE, a wrong address, and the jump lands nowhere useful.
Encoding 0xDEADBEEF into the payload.
Write the bytes LSB first: \xef\xbe\xad\xde
Why: §3.2 / Lesson 5: little-endian stores the least-significant byte at the lowest address. Reversed, the 4 bytes load back as exactly 0xDEADBEEF.
Anomaly
Predict first
A student writes this, and it looks reasonable:
What do the 12 filler bytes need to be?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Wastes effort — and worse, may avoid bytes like 0x00 that actually break gets early.
What do the 12 filler bytes need to be?
Why: Wastes effort — and worse, may avoid bytes like 0x00 that actually break gets early. The filler's value was never the point.
Trap
What do the 12 filler bytes need to be?
Carefully craft the 12 bytes to be 'correct' data
Why: Wastes effort — and worse, may avoid bytes like 0x00 that actually break gets early. The filler's value was never the point.
What do the 12 filler bytes need to be?
Any non-terminating bytes; only their COUNT matters
Why: The 12 bytes are pure spacing to position the rip overwrite. Use 'A' × 12. (One real constraint: avoid bytes that terminate the copy early, e.g. a null for some string functions.)
Concept
The rip is data on the stack, and ret trusts it blindly. Once attacker input can reach those 4 bytes, the attacker chooses where the function returns — i.e. what the CPU executes next.
The book's blunt rule: if your program has a buffer overflow bug, assume it is exploitable and an attacker can take control.
Section
Part 4 · §3.2
Intuition
Little-endian stores the least-significant byte at the lowest address. Since gets writes from low addresses upward, the first byte you type is the byte that ends up lowest — so you type the smallest-place byte first.
| Address (low→high) | Byte typed | Meaning |
|---|---|---|
| lowest | 0xEF | bits 0–7 |
| +1 | 0xBE | bits 8–15 |
| +2 | 0xAD | bits 16–23 |
| highest | 0xDE | bits 24–31 |
Read back as a 32-bit word, those four bytes are 0xDEADBEEF — the value we wanted.
Pattern
Step through it
Step through Little-endian, one more time one row at a time. What is driving the change, and what would the row after the last one be?
Concept
shellcode — Injected machine code the attacker wants the CPU to run — classically, code that spawns a shell, giving interactive control of the machine.
Hijacking the rip only redirects execution. Shellcode is what you redirect it to. The three strategies below differ only in where the shellcode lives and where rip points.
Reminder: in this deck the shellcode bytes are always abstract — written [shellcode] — never real machine code.
Estimation
Predict first
Easiest case: the code you want to run is already loaded somewhere at a known address. You don't inject anything — you just redirect.
Commit before you compute: what does Strategy 1: code already in memory at a known address come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: rip = address of the existing code
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. ret jumps straight there; nothing else is injected.
Worked example
Easiest case: the code you want to run is already loaded somewhere at a known address. You don't inject anything — you just redirect.
payload = [12 garbage bytes] + [address]
# fill buf+sfp rip = address of existing code12 garbage bytes to reach the rip
Why: 8 (buf) + 4 (sfp), same as before.
rip = address of the existing code
Why: ret jumps straight there; nothing else is injected. (This previews ret2libc, Lesson 13 — redirecting to code that's already present.)
| Payload region | Bytes | Overwrites |
|---|---|---|
| garbage | 0–11 | buf + sfp |
| address | 12–15 | rip → existing code |
Comparison
Comparison matrix
From Strategy 1: code already in memory at a known address: refill the Overwrites column from what you know. The rest of the table is as it appeared.
| Payload region | Bytes | Overwrites |
|---|---|---|
| garbage | 0–11 | buf + sfp |
| address | 12–15 | rip → existing code |
Ranking
Put in order
Put the moves of Strategy 2: inject 8-byte shellcode into the buffer into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. It must fit in the 8 bytes of buf in this case.
Worked example
If the shellcode is small enough to fit in buf itself (here, 8 bytes), put it in the buffer and point rip back down at buf.
payload = [shellcode (8B)] # fills buf[0..7]
+ [4 garbage bytes] # over the sfp
+ [address of buf] # rip -> start of bufShellcode occupies buf[0..7]
Why: It must fit in the 8 bytes of buf in this case.
4 garbage bytes overwrite the sfp
Why: Spacer to reach the rip; the sfp's value is irrelevant once we are leaving for good.
rip = address of buf, so ret runs the buffer
Why: ret jumps to buf, and the CPU executes the bytes we just wrote there as instructions.
| Payload region | Bytes | Overwrites / does |
|---|---|---|
| [shellcode] | 0–7 | buf — executed later |
| garbage | 8–11 | sfp (spacer) |
| &buf | 12–15 | rip → buf |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
rip = address of buf, so ret runs the buffer
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
If the shellcode is small enough to fit in buf itself (here, 8 bytes), put it in the buffer and point rip back down at buf.
Step zero
Discussion prompt
Strategy 3: shellcode too big — place it above the rip — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: 12 garbage bytes reach the rip
Answer:
Worked example
If the shellcode is large (say ~100 bytes), it won't fit in the 12 bytes below the rip. Put the shellcode above the rip and aim rip at it.
payload = [12 garbage bytes] # buf + sfp
+ [address of rip+4] # rip -> just past itself
+ [shellcode (~100B)] # sits above the rip12 garbage bytes reach the rip
Why: Same 8 + 4 offset as always.
rip = address of rip+4
Why: rip+4 is the next stack slot up — the first byte of the shellcode that follows the address in the payload.
Shellcode lives above the rip and runs on return
Why: We are no longer limited by buf's size; the payload simply continues upward past the return address into the shellcode.
| Payload region | Stack location | Role |
|---|---|---|
| garbage (12 B) | buf + sfp | spacer |
| &(rip+4) | rip | redirect target |
| [shellcode] | rip+4 and up | executed code |
Error analysis
Annotate
Walk the callouts on Strategy 3: shellcode too big — place it above the rip. Each one is a place this is easy to get subtly wrong.
Concept
The three strategies are a decision tree driven by two questions: do I need to inject code at all, and if so does it fit below the rip?
| Situation | Strategy | rip points to |
|---|---|---|
| target code already loaded | 1 — redirect | the existing code |
| small shellcode (≤ buf) | 2 — inject in buf | &buf |
| large shellcode | 3 — place above rip | &(rip+4) |
All three reuse the same 12-byte offset to reach the rip — only the destination and the code's location change.
Intuition
Stack addresses wobble between runs (environment size, argv length, the OS). Demanding the exact byte address of your shellcode is brittle — off by one and the jump misses.
The fix is to make a bigger target to hit: that is what the NOP sled is for.
Concept
NOP sled — A run of no-op bytes placed before the shellcode; landing anywhere in the run 'slides' the CPU forward into the shellcode.
In practice the exact stack address is hard to predict. Prepend many NOPs to the shellcode and aim rip near them: any landing inside the sled executes harmless no-ops until it reaches the real code.
It turns a pinpoint-accuracy problem into a hit-the-broad-side problem. (Modern defenses like ASLR exist to ruin this — Lesson 15.)
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of buffer overflow, shellcode, NOP sled as L08 · Buffer Overflows & Stack Smashing uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Anomaly
Predict first
A student writes this, and it looks reasonable:
How many filler bytes before the rip, for char buf[8]?
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Fills buf but stops at the sfp.
How many filler bytes before the rip, for char buf[8]?
Why: Fills buf but stops at the sfp. The next 4 bytes you write land on the saved frame pointer, not the rip — the address is off by one word.
Trap
How many filler bytes before the rip, for char buf[8]?
Use 8 — just the size of buf
Why: Fills buf but stops at the sfp. The next 4 bytes you write land on the saved frame pointer, not the rip — the address is off by one word.
How many filler bytes before the rip, for char buf[8]?
Use 12 — buf (8) + sfp (4)
Why: The sfp sits between buf and the rip. You must cross both to reach the return address: 8 + 4 = 12.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The sfp sits between buf and the rip. You must cross both to reach the return address: 8 + 4 = 12.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Fills buf but stops at the sfp. The next 4 bytes you write land on the saved frame pointer, not the rip — the address is off by one word.
Section
Part 5 · §3.2
Concept
The Morris Worm (late 1980s) used a buffer overflow to spread — one of the first internet-wide incidents and the reason CERT exists.
Aleph One's 1996 Phrack article 'Smashing the Stack for Fun and Profit' turned the technique into common knowledge.
| Event | When | Significance |
|---|---|---|
| Morris Worm | late 1980s | buffer overflow used to self-propagate |
| 'Smashing the Stack' | 1996 (Phrack 49) | popularized the exact technique |
| Code Red worm | 2001 | IIS overflow — 369,000 machines |
Intuition
The technique was published in 1996 and the Morris Worm predates that — yet memory-corruption CVEs keep appearing. Why? Because the install base of C never shrinks: kernels, libraries, and firmware accumulate.
Every new embedded device — a router, a car ECU, a smart bulb — is often more legacy C exposed to the network. The attack surface grows faster than it's rewritten.
Explain it
Discussion prompt
Explain Why a 35-year-old bug won't die to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Every new embedded device — a router, a car ECU, a smart bulb — is often more legacy C exposed to the network. The attack surface grows faster than it's rewritten.
Concept
Code Red compromised 369,000 machines through a buffer overflow in Microsoft IIS — proof the bug class scales to internet-wide damage.
Memory-corruption bugs still dominate CVE statistics. Legacy C and embedded systems (routers, IoT, cars) keep this attack surface very much alive.
Analogy
Discussion prompt
Explain Still relevant today by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Memory-corruption bugs still dominate CVE statistics. Legacy C and embedded systems (routers, IoT, cars) keep this attack surface very much alive.
Ranking
Put in order
Put the moves of The fix: bounds-checked input into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. fgets is told the buffer size, so it stops before overflowing — it never writes past buf.
Worked example
// vulnerable:
// gets(buf); // unbounded
// fixed:
fgets(buf, sizeof(buf), stdin); // reads at most sizeof(buf)-1 bytesReplace gets with fgets
Why: fgets is told the buffer size, so it stops before overflowing — it never writes past buf.
Pass sizeof(buf), not a hard-coded number
Why: If buf's size changes later, sizeof tracks it automatically; a literal would drift out of sync and reintroduce the bug.
| Call | Knows buf size? | Can overflow? |
|---|---|---|
| gets(buf) | no | yes — unbounded |
| fgets(buf, sizeof(buf), stdin) | yes | no — capped |
Verify the bound holds
Why: fgets writes at most sizeof(buf)-1 chars plus a terminating '\0' — exactly sizeof(buf) bytes worst case, which fits. The 12-byte path to the rip can no longer be crossed by input.
Trade off
Comparison matrix
From The fix: bounds-checked input: every row here is a choice with a cost. Fill the Can overflow? column, then say which row you would actually pick and what you give up for it.
| Call | Knows buf size? | Can overflow? |
|---|---|---|
| gets(buf) | no | yes — unbounded |
| fgets(buf, sizeof(buf), stdin) | yes | no — capped |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Reading a line safely into buf.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: gets takes no size and stops only at newline/EOF.
Reading a line safely into buf.
Why: gets takes no size and stops only at newline/EOF. 'Users will type short lines' is not a security boundary — the attacker won't.
Trap
Reading a line safely into buf.
Use gets(buf) and trust short input
Why: gets takes no size and stops only at newline/EOF. 'Users will type short lines' is not a security boundary — the attacker won't.
Reading a line safely into buf.
Use fgets(buf, sizeof(buf), stdin)
Why: fgets is length-limited by construction. gets is so dangerous it was removed from the C11 standard library entirely.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
[shellcode]. The skill being built is reading a frame and spotting the overflow.; In a memory-safe language an out-of-range index throws; in C it is silent and the write simply lands somewhere it shouldn't.; An out-of-bounds write is only dangerous because of what sits next to the buffer. On the stack, that neighbor is control-flow metadata — the sfp and the rip.Pattern
Predict first
The table runs: gets(buf) | fgets(buf, sizeof(buf), stdin) · strcpy(dst, src) | strncpy(dst, src, sizeof(dst))
In Other unbounded copies have safe twins, given the rows so far: what is the next one — the row where Unbounded is sprintf(buf, ...)?
Correct: sprintf(buf, ...) | snprintf(buf, sizeof(buf), ...)
| Unbounded | Bounded twin |
|---|---|
| gets(buf) | fgets(buf, sizeof(buf), stdin) |
| strcpy(dst, src) | strncpy(dst, src, sizeof(dst)) |
| sprintf(buf, ...) | snprintf(buf, sizeof(buf), ...) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Any copy that takes no destination size — strcpy, sprintf, strcat — can overflow the same way.
Worked example
strcpy(dst, src); // unbounded
strncpy(dst, src, sizeof(dst)); // bounded
sprintf(buf, "%s", s); // unbounded
snprintf(buf, sizeof(buf), "%s", s); // boundedgets is not the only offender
Why: Any copy that takes no destination size — strcpy, sprintf, strcat — can overflow the same way.
Prefer the size-aware variant
Why: The 'n' / snprintf forms take a length and stop, so the destination can't be overrun.
| Unbounded | Bounded twin |
|---|---|
| gets(buf) | fgets(buf, sizeof(buf), stdin) |
| strcpy(dst, src) | strncpy(dst, src, sizeof(dst)) |
| sprintf(buf, ...) | snprintf(buf, sizeof(buf), ...) |
Comparison
Comparison matrix
From Other unbounded copies have safe twins: refill the Bounded twin column from what you know. The rest of the table is as it appeared.
| Unbounded | Bounded twin |
|---|---|
| gets(buf) | fgets(buf, sizeof(buf), stdin) |
| strcpy(dst, src) | strncpy(dst, src, sizeof(dst)) |
| sprintf(buf, ...) | snprintf(buf, sizeof(buf), ...) |
Concept
Length-limited APIs (fgets, strncpy, snprintf) fix one call site at a time — necessary but error-prone at scale.
The real cure (Week 4 preview) is memory-safe languages plus runtime mitigations — stack canaries, non-executable stacks (NX), and ASLR — which assume the bug exists and try to stop exploitation anyway.
Counterexample
Discussion prompt
Length-limited APIs (fgets, strncpy, snprintf) fix one call site at a time — necessary but error-prone at scale.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Section
Part 6
Constraint
Discussion prompt
Run Pattern: build a stack-smashing payload with this step confiscated:
Set the rip to your target address, typed little-endian (LSB first).
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
char buf[8]: 8 + 4 = 12.Pattern
char buf[8]: 8 + 4 = 12.Edge cases
Discussion prompt
Pattern: build a stack-smashing payload works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
char buf[8]: 8 + 4 = 12.Elimination
Eliminate the wrong options
An overflow of name must travel how many bytes from the first byte of name to the first byte of the saved return address (rip)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: C
Why: From name you must cross every slot below the rip: name (16) + the int x (4) + the sfp (4) = 24 bytes before the rip's first byte. The sfp is 4 bytes because this is 32-bit.
Check
A 32-bit function has locals laid out (high→low) as: int x; then char name[16];, with the usual sfp and rip above. The frame is, high to low: rip, sfp, x, name. Solve on paper before clicking.
Check your understanding
An overflow of name must travel how many bytes from the first byte of name to the first byte of the saved return address (rip)?
Answer: C
Why: From name you must cross every slot below the rip: name (16) + the int x (4) + the sfp (4) = 24 bytes before the rip's first byte. The sfp is 4 bytes because this is 32-bit.
Concept
\xef\xbe\xad\xde.Concept
Concept
vulnerable() with gcc -m32 -O0 -fno-stack-protector, then in gdb watch the rip at ebp+4 get overwritten as input length passes 12 bytes.Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why C Overflows · The Vulnerable Function · Overwriting the Return Address · Shellcode & Placement · History & The Fix · Consolidate. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
You can explain why C never bounds-checks, draw the gets(buf) frame, compute the 8 + 4 = 12-byte offset to the rip, type an address little-endian, lay out all three shellcode strategies, and fix the bug with fgets.
| Idea | The one fact |
|---|---|
| No bounds checking | C/C++/Obj-C: the programmer owns every bound |
| Offset to rip (char buf[8]) | 8 (buf) + 4 (sfp, 32-bit) = 12 |
| Address encoding | little-endian: 0xDEADBEEF → \xef\xbe\xad\xde |
| Shellcode placement | existing code / in buf / above rip (+ NOP sled) |
| The fix | fgets(buf, sizeof(buf), stdin) |
Want this taught 1-on-1? Alexander tutors Computer Security — $55/session, free consultation.