L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript

CS 161, Lesson 36, in 57 slides. It covers the browser and protocol model that sits behind every web attack: the parts of a URL in section 18.1, the HTTP request and response model with GET against POST in sections 18.2 to 18.5, the webpage as a distributed application together with HTML and frame isolation in sections 18.6 and 18.7, and CSS, JavaScript, the DOM, and the JavaScript sandbox in sections 18.8 and 18.9. It is anchored to textbook sections 18.1 to 18.9.

Subject: Computer Security · 84 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. How the Web Actually Works

Title

CS 161 · Lesson 36 of 45

URLs · the HTTP request/response model · HTML & frame isolation · CSS, JavaScript & the DOM — the browser model behind every web attack

2. By the end of this lesson you can…

Objectives

  1. Break a URL into its three mandatory parts — protocol, location/domain, and path — and say which one picks the server.
  2. Describe the HTTP request/response model, read a request's method/path/version + headers, and contrast GET vs POST (only POST has a body).
  3. Explain the webpage as a distributed application and identify the security-relevant HTML elements, including why frames are isolated.
  4. Trace the HTML → DOM → JavaScript-modifies-DOM → render pipeline and explain why page JS runs in a sandbox.
  5. Map each piece of this model to the upcoming web attacks (SOP, sessions/CSRF, XSS, clickjacking) it makes possible.

3. What survived from L35 · SQL Injection & Code Injection?

Warm-up

Discussion prompt

Before we open L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript: without looking back, what was the main idea of L35 · SQL Injection & Code Injection, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

CS 161 Lesson 35 (50 slides): code injection as the data-becomes-code pattern (§17.1), SQL injection as its canonical instance (§17.2), the anatomy of an injection payload and the OR 1=1 login bypass (§17.3), escaping and why it is fragile (§17.4), and parameterized queries as the real fix that separates code from data (§17.5). Authorized security education with toy/sandbox examples only.

4. Why a whole lesson on web basics

Concept

Every web attack in this unit — same-origin policy, sessions, CSRF, XSS, clickjacking — is an attack on how the browser and server talk. Before we can break the web, we have to model it precisely.

Where does it live?
§18.1 URLs name every resource
How does it travel?
§18.2–18.5 HTTP request/response
What gets rendered?
§18.6–18.9 HTML, the DOM & JavaScript

5. Which is which: Why a whole lesson on web basics

Matching

Match the pairs

From Why a whole lesson on web basics — match each one to what it actually does. The descriptions have been shuffled.

  • c1. Where does it live?
  • c2. How does it travel?
  • c3. What gets rendered?
  • b1. §18.1 URLs name every resource
  • b2. §18.2–18.5 HTTP request/response
  • b3. §18.6–18.9 HTML, the DOM & JavaScript

Why: Where does it live?, How does it travel?, What gets rendered? are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

6. URLs — Naming Every Resource

Section

Part 1 · §18.1 protocol, location, path

7. §18.1 Every web resource has a URL

Concept

Concrete scenario: you type an address into the browser bar and a page loads. That address is a URL — a Uniform Resource Locator — and it's how the browser knows what to fetch and from where.

URL — A string that names a web resource. Every resource the browser loads — a page, an image, a script — is identified by a URL. It has three mandatory parts: a protocol, a location (domain), and a path.

8. Break it if you can: §18.1 Every web resource has a URL

Counterexample

Discussion prompt

Concrete scenario: you type an address into the browser bar and a page loads. That address is a URL — a Uniform Resource Locator — and it's how the browser knows what to fetch and from where.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. §18.1 The three mandatory parts

Concept

Take the running example http://www.example.com/index.html. Read it left to right and it splits into exactly three parts.

http://www.example.com/index.html
\__/   \_____________/\________/
  |            |          |
PROTOCOL    LOCATION     PATH
(scheme)    (domain)
  1. Protocol / scheme — comes before ://. For this course just HTTP and HTTPS (HTTPS = HTTP over TLS).
  2. Location / domain — www.example.com: WHICH server to contact.
  3. Path — /index.html: which resource to ask that server for.

10. By analogy: §18.1 The three mandatory parts

Analogy

Discussion prompt

Explain §18.1 The three mandatory parts by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Take the running example http://www.example.com/index.html. Read it left to right and it splits into exactly three parts.

11. §18.1 The location can carry more

Concept

The location part has two optional extras: a username@ prefix, and an explicit :port. If no port is given, the browser uses the protocol's default (80 for HTTP, 443 for HTTPS).

http://alice@www.example.com:81/index.html
       \___/ \_____________/ \_/
      userinfo    host       port

So www.example.com:81 means: contact that server on port 81 instead of the default 80. The port narrows down which service on the machine answers.

12. Teach it back: §18.1 The location can carry more

Explain it

Discussion prompt

Explain §18.1 The location can carry more to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The location part has two optional extras: a username@ prefix, and an explicit :port. If no port is given, the browser uses the protocol's default (80 for HTTP, 443 for HTTPS).

13. §18.1 A URL is a mailing address

Intuition

Think of mailing a letter. The location/domain is the building you send it to — it routes the letter to the right place. The path is the apartment or person inside that building — it only matters once the letter has arrived.

The protocol is the postal service you use (regular vs certified mail) — it decides how the delivery is handled, here HTTP vs HTTPS.

Ask yourself: if you change only the path, does the letter go to a different building? (No — same building/server, different room. Only changing the domain changes the server.)

14. What rests on this: §18.1 A URL is a mailing address

Socratic

Discussion prompt

The protocol is the postal service you use (regular vs certified mail) — it decides how the delivery is handled, here HTTP vs HTTPS.

Suppose that were not true. What is the first thing in L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

Answer:

Ask yourself: if you change only the path, does the letter go to a different building? (No — same building/server, different room. Only changing the domain changes the server.)

15. §18.1 Label every part of a URL

Worked example

Take the URL https://user@shop.example.com:8443/cart/items?id=7

Why: We'll walk left to right and assign each chunk to its role.

chunkpartwhat it does
httpsprotocol/schemeHTTP over TLS — how to talk
user@userinfo (optional)credentials prefix on the location
shop.example.comlocation/domainWHICH server to contact
:8443port (optional)which service/port on that server
/cart/itemspathwhich resource to request FROM that server
?id=7query (optional)extra parameters sent with the request

Verify: which single part decides the server? The location/domain, shop.example.com

Why: §18.1: the domain alone selects the server. The path, query, and even the port are all interpreted only AFTER the browser has reached that server — they never change which machine is contacted.

16. Fill in: part for §18.1 Label every part of a URL

Comparison

Comparison matrix

From §18.1 Label every part of a URL: refill the part column from what you know. The rest of the table is as it appeared.

chunkpartwhat it does
httpsprotocol/schemeHTTP over TLS — how to talk
user@userinfo (optional)credentials prefix on the location
shop.example.comlocation/domainWHICH server to contact
:8443port (optional)which service/port on that server
/cart/itemspathwhich resource to request FROM that server
?id=7query (optional)extra parameters sent with the request

17. Something is wrong here: 'the path tells the browser which server to contact'

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student: 'The path /index.html is what tells the browser where to go.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The path is requested FROM a server that's already been chosen.

A student: which part of the URL picks the server?

Why: The path is requested FROM a server that's already been chosen. By itself it says nothing about which machine to contact — /index.html is meaningless until you know whose /index.html.

18. Trap: 'the path tells the browser which server to contact'

Trap

The trap

A student: 'The path /index.html is what tells the browser where to go.'

Treat the path as choosing the server

Why: Wrong. The path is requested FROM a server that's already been chosen. By itself it says nothing about which machine to contact — /index.html is meaningless until you know whose /index.html.

The fix

A student: which part of the URL picks the server?

The LOCATION/domain selects the server; the path is requested from it

Why: §18.1: www.example.com decides which server to contact. The browser connects there first, THEN asks that server for the path /index.html. (This matters for the origin in L37 — origin is scheme+host+port, never the path.)

19. HTTP — The Request/Response Model

Section

Part 2 · §18.2–18.5 methods & headers

20. §18.2 The request/response model

Concept

Concrete scenario: you click a link. Under the hood the browser opens a connection to the server, sends a request, and the server sends back a response. HTTP powers the web with exactly this back-and-forth.

HTTP request/response model — HTTP is a protocol where the CLIENT (browser) initiates a connection and sends a request; the SERVER responds with a reply. The client always speaks first; the server only answers.

21. Take the definitions apart: URL vs HTTP request/response…

Definition probe

Sort into buckets

Every line below is part of the definition of URL or of HTTP request/response model — one or the other, never both. Put each where it belongs.

URL
A string that names a web resource.; Every resource the browser loads; It has three mandatory parts
HTTP request/response model
HTTP is a protocol where the CLIENT (browser) initiates a connection and sends a request; the SERVER responds with a reply.; The client always speaks first
b1
A string that names a web resource. Every resource the browser loads — a page, an image, a script — is identified by a URL. It has three mandatory parts: a protocol, a location (domain), and a path.
b2
HTTP is a protocol where the CLIENT (browser) initiates a connection and sends a request; the SERVER responds with a reply. The client always speaks first; the server only answers.

22. §18.3 HTTP/1.1 is text; HTTP/2 is binary

Concept

HTTP/1.1 is text-based: a request or response is a header section followed by an optional payload body. You can read it by eye, which is why we study it.

HTTP/2 packs the same information into a binary wire format for speed — but the concepts are identical: same methods, same headers, same request/response idea. Learn HTTP/1.1 and you understand both.

23. §18.4 The structure of a request

Concept

The very first line of a request carries three things: the method, the path, and the version. After that come the request headers, one per line.

GET /index.html HTTP/1.1
Host: www.example.com
Dnt: 1
User-Agent: Mozilla/5.0
piecevaluerole
methodGETthe action to perform
path/index.htmlwhich resource on the server
versionHTTP/1.1the protocol version
Host:www.example.comwhich domain (one IP can host many)
Dnt:1a request header — 'do not track'

24. What each one costs: §18.4 The structure of a request

Trade off

Comparison matrix

From §18.4 The structure of a request: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

piecevaluerole
methodGETthe action to perform
path/index.htmlwhich resource on the server
versionHTTP/1.1the protocol version
Host:www.example.comwhich domain (one IP can host many)
Dnt:1a request header — 'do not track'

25. §18.4 Why a separate Host header?

Intuition

The first line's path is just /index.html — no domain. But one server (one IP) can host many sites. So the request repeats the domain in the Host: header to say which site's /index.html it wants.

Ask yourself: where did the domain from the URL go? (It's in the Host: header. The browser split the URL — domain into Host:, path into the first line.)

26. §18.5 GET vs POST: two intents

Concept

Scenario: viewing a page vs submitting a login form. These are different intents, and HTTP gives them different methods.

The key structural fact: only POST has a body. GET passes its data through query parameters in the URL instead.

27. §18.5 A real GET request

Concept

A GET carries its data in the URL's query string — everything after the ?, as key=value pairs joined by &. There is no body.

GET /posts?search=security&sortby=popularity HTTP/1.1
Host: www.example.com

(no body)

The server reads search=security and sortby=popularity straight out of the path's query part. Because it's in the URL, it's visible in the address bar, bookmarks, and server logs.

28. §18.5 A real POST request

Concept

A POST puts its data in the body, below a blank line that separates headers from payload. This is how a login form sends a password.

POST /login HTTP/1.1
Host: www.example.com
Content-Type: application/x-www-form-urlencoded
Content-Length: 35

username=alice&password=hunter2

The blank line marks the end of the headers; everything after it is the body. The username=…&password=… payload modifies state — it creates a logged-in session.

29. Predict the next row: §18.4–18.5 Read the parts of an HTTP request

Pattern

Predict first

The table runs: first line | POST /login HTTP/1.1 | method + path + version · headers | Host / Content-Type / Content-Length | metadata about the request · blank line | (empty) | separates headers from body

In §18.4–18.5 Read the parts of an HTTP request, given the rows so far: what is the next one — the row where region is body?

Correct: body | username=alice&password=hunter2 | the state-changing payload (POST only)

regioncontentmeaning
first linePOST /login HTTP/1.1method + path + version
headersHost / Content-Type / Content-Lengthmetadata about the request
blank line(empty)separates headers from body
bodyusername=alice&password=hunter2the state-changing payload (POST only)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Every HTTP/1.1 request has the same shape: a first line, header lines, a blank line, then an optional body.

30. §18.4–18.5 Read the parts of an HTTP request

Worked example

Take the POST login request above and split it into its regions

Why: Every HTTP/1.1 request has the same shape: a first line, header lines, a blank line, then an optional body.

regioncontentmeaning
first linePOST /login HTTP/1.1method + path + version
headersHost / Content-Type / Content-Lengthmetadata about the request
blank line(empty)separates headers from body
bodyusername=alice&password=hunter2the state-changing payload (POST only)

Verify: if this were a GET, where would username/password go?

Why: §18.5: a GET has no body, so the data would have to ride in the URL as ?username=alice&password=hunter2 — visible in the address bar and logs. That's exactly why credentials use POST, not GET.

31. Fill in: content for §18.4–18.5 Read the parts of an HTTP request

Comparison

Comparison matrix

From §18.4–18.5 Read the parts of an HTTP request: refill the content column from what you know. The rest of the table is as it appeared.

regioncontentmeaning
first linePOST /login HTTP/1.1method + path + version
headersHost / Content-Type / Content-Lengthmetadata about the request
blank line(empty)separates headers from body
bodyusername=alice&password=hunter2the state-changing payload (POST only)

32. §18.5 The model vs reality of GET

Intuition

The HTTP spec says a GET should be 'safe' — it shouldn't change server state, just retrieve. That's the intended contract. A verifier (or a cache, or a prefetcher) is allowed to assume GETs are harmless.

But the spec only says shouldn't, not can't. Nothing in the protocol stops a server from acting on a GET — and modern apps routinely do, via query parameters like /delete?id=7.

Ask yourself: why is this gap security-relevant? (A GET that changes state can be triggered just by loading a URL — no form needed — which is the heart of CSRF in L39.)

33. Trap: 'GET requests can't change server state'

Trap

The trap

A student: 'GET is read-only by the spec, so a GET request can never modify anything on the server.'

Assume the spec's 'should' is enforced as a 'cannot'

Why: Wrong. The spec says a GET SHOULDN'T change state — it does NOT prevent it. In practice apps wire state changes to GET URLs (e.g. /transfer?to=mallory&amt=100), which is precisely what makes CSRF possible.

The fix

A student: what does the spec actually guarantee about GET?

Treat GET as 'should be safe' but assume servers may break that

Why: §18.5: GET is supposed to be side-effect-free, but the protocol can't enforce it, so real servers often change state on GET. This gap foreshadows CSRF (L39): a state-changing GET fires from just a URL load.

34. Something is wrong here: 'GET requests have a body'

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student: 'I'll send the username and password in the body of a GET request.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: In this model GET carries data via QUERY PARAMETERS in the URL, not a body — only POST has a body.

A student: where does each method put its data?

Why: In this model GET carries data via QUERY PARAMETERS in the URL, not a body — only POST has a body. Data smuggled in a GET 'body' is non-standard and routinely ignored.

35. Trap: 'GET requests have a body'

Trap

The trap

A student: 'I'll send the username and password in the body of a GET request.'

GET /login HTTP/1.1
Host: www.example.com

username=alice&password=hunter2   <-- wrong: GET has no body

Put state-changing data in a GET body

Why: Wrong. In this model GET carries data via QUERY PARAMETERS in the URL, not a body — only POST has a body. Data smuggled in a GET 'body' is non-standard and routinely ignored.

The fix

A student: where does each method put its data?

GET /login?username=alice HTTP/1.1   <-- GET: query params

POST /login HTTP/1.1                  <-- POST: body

username=alice&password=hunter2

GET → query parameters in the URL; POST → request body

Why: §18.5: only POST has a body. GET passes data through the URL's query string. Knowing which is which tells you where to look for parameters when analyzing a request.

36. Inspect it line by line: Trap: 'GET requests have a body'

Error analysis

Annotate

Walk the callouts on Trap: 'GET requests have a body'. Each one is a place this is easy to get subtly wrong.

  • Wrong. In this model GET carries data via QUERY PARAMETERS in the URL, not a body — only POST has a body. Data smuggled in a GET 'body' is non-standard and routinely ignored.
  • §18.5: only POST has a body. GET passes data through the URL's query string. Knowing which is which tells you where to look for parameters when analyzing a request.

37. The Webpage as a Distributed App + HTML

Section

Part 3 · §18.6–18.7 HTML & frame isolation

38. §18.6 A web page is a distributed application

Concept

Scenario: loading a page isn't 'downloading a document' — it's running a program split across two machines. There's a server-side component that generates the response and a browser-side component that renders it.

Distributed web application — A web page runs as two cooperating parts: a server-side component that produces the response, and a browser-side component (HTML/CSS/JavaScript) that the browser executes and renders. Both halves are part of one application.

39. §18.6 The media type tells the browser how to read it

Concept

The server's response includes a media type (a.k.a. content type) that tells the browser how to interpret the body — as HTML to render, an image to draw, JSON data, and so on.

HTTP/1.1 200 OK
Content-Type: text/html

<html> … </html>

Content-Type: text/html says 'treat the body as HTML.' The same bytes labeled differently would be handled differently — the label drives the interpretation.

40. §18.7 HTML structures the document

Concept

HTML (HyperText Markup Language) structures the page as a tree of nested elements written with tags like <tag>…</tag>. A few elements matter a lot for security.

<a href="https://example.com">a link</a>
<img src="photo.jpg">
<script>alert(1)</script>
<iframe src="https://other.com"></iframe>
elementwhat it doeswhy it matters
<a href>a hyperlinknavigates to another URL
<img src>loads an imagefires a GET to any URL on load
<script>INLINE SCRIPTruns JavaScript in the page
<iframe src>embeds another page (a FRAME)a page inside a page — a risk

41. What each one costs: §18.7 HTML structures the document

Trade off

Comparison matrix

From §18.7 HTML structures the document: every row here is a choice with a cost. Fill the what it does column, then say which row you would actually pick and what you give up for it.

elementwhat it doeswhy it matters
<a href>a hyperlinknavigates to another URL
<img src>loads an imagefires a GET to any URL on load
<script>INLINE SCRIPTruns JavaScript in the page
<iframe src>embeds another page (a FRAME)a page inside a page — a risk

42. §18.7 Why an inline <script> is special

Intuition

<img> and <a> just point at resources. But <script>alert(1)</script> is executable code embedded in the document — the browser runs whatever is between the tags as JavaScript.

Ask yourself: if an attacker could slip their own <script> tag into a page you view, what could they do? (Run arbitrary JS in that page — the entire basis of XSS in L40.)

43. §18.7 Frames embed a page inside a page

Concept

An <iframe> embeds one web page inside another. The outer page is the host; the inner page (the frame) is whatever the src points to — possibly a page from a different, untrusted site.

That's a real risk: the outer page is embedding a possibly-malicious inner page, and vice versa. So browsers enforce a protection.

Frame isolation — A browser rule that an outer page and an embedded inner frame CANNOT read or modify each other's contents. Each frame is walled off from the other, even though they're displayed together.

44. Predict the next row: §18.7 What an embedded frame can and can't do

Pattern

Predict first

The table runs: A displays B in a box | yes | embedding is the point of iframes · B reads A's form fields | NO | frame isolation blocks it · A reads B's cookies/contents | NO | frame isolation blocks it

In §18.7 What an embedded frame can and can't do, given the rows so far: what is the next one — the row where action is A resizes/positions the iframe?

Correct: A resizes/positions the iframe | yes | the box is A's; the contents aren't

actionallowed?why
A displays B in a boxyesembedding is the point of iframes
B reads A's form fieldsNOframe isolation blocks it
A reads B's cookies/contentsNOframe isolation blocks it
A resizes/positions the iframeyesthe box is A's; the contents aren't

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Now two pages render together: A is the outer page, B (b.com) is the inner frame.

45. §18.7 What an embedded frame can and can't do

Worked example

Page A from a.com embeds an iframe whose src is b.com

Why: Now two pages render together: A is the outer page, B (b.com) is the inner frame.

<!-- served by a.com -->
<h1>A's page</h1>
<iframe src="https://b.com/widget"></iframe>
actionallowed?why
A displays B in a boxyesembedding is the point of iframes
B reads A's form fieldsNOframe isolation blocks it
A reads B's cookies/contentsNOframe isolation blocks it
A resizes/positions the iframeyesthe box is A's; the contents aren't

Verify: the difference between the box and the contents

Why: §18.7: A controls the iframe's geometry (where the box sits) but NOT what's inside it. Frame isolation stops cross-frame reads/writes — that A still controls the box's POSITION is exactly the gap clickjacking exploits (L41).

46. Which is which, by allowed?

Discrimination

Sort into buckets

Sort these by allowed?, from memory, without looking back at §18.7 What an embedded frame can and can't do. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

yes
A displays B in a box; A resizes/positions the iframe
NO
B reads A's form fields; A reads B's cookies/contents
g1
allowed? is "yes" for A displays B in a box, A resizes/positions the iframe — that is what the table on "§18.7 What an embedded frame can and…" records, and it is the single property separating this group from the rest.
g2
allowed? is "NO" for B reads A's form fields, A reads B's cookies/contents — that is what the table on "§18.7 What an embedded frame can and…" records, and it is the single property separating this group from the rest.

47. Something is wrong here: 'an embedded iframe can freely read/modify the outer…

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student: 'Since the iframe is inside my page, the embedded page can reach up and read my page's contents — and I can reach into it.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Frame isolation forbids it. An embedded b.com frame cannot read a.com's fields, and a.com cannot read inside the b.com frame.

A student: what does frame isolation actually permit?

Why: Frame isolation forbids it. An embedded b.com frame cannot read a.com's fields, and a.com cannot read inside the b.com frame. They render together but are walled off.

48. Trap: 'an embedded iframe can freely read/modify the outer page'

Trap

The trap

A student: 'Since the iframe is inside my page, the embedded page can reach up and read my page's contents — and I can reach into it.'

Assume outer and inner frames share access to each other's DOM

Why: Wrong. Frame isolation forbids it. An embedded b.com frame cannot read a.com's fields, and a.com cannot read inside the b.com frame. They render together but are walled off.

The fix

A student: what does frame isolation actually permit?

Frames render together but can't touch each other's contents

Why: §18.7: the outer page can position/size the iframe box, but neither side can read or modify the other's DOM. This containment is the foundation of the same-origin policy (L37) and the reason clickjacking (L41) attacks geometry, not contents.

49. CSS, JavaScript & the DOM

Section

Part 4 · §18.8–18.9 the render pipeline

50. §18.8 CSS styles the page — and is dangerous

Concept

CSS (Cascading Style Sheets) controls how the page looks: colors, layout, positioning. It seems harmless next to JavaScript — but it isn't.

Used maliciously, CSS is as powerful as JavaScript: forcing a victim to load attacker-controlled CSS is roughly as dangerous as forcing them to run malicious JS. Styling can move, hide, and overlay elements to trick or exfiltrate.

51. §18.9 JavaScript runs in the browser

Concept

JavaScript is the programming language the browser runs. Crucially, page JavaScript can arbitrarily modify any HTML or CSS on the page — add elements, change text, restyle, rewrite the whole document.

JavaScript on a page — Code the browser executes as part of rendering a page. It can read and arbitrarily modify the page's HTML and CSS at runtime — which is why injecting attacker JavaScript (XSS) is so powerful.

52. §18.9 The DOM is the parsed tree

Concept

JavaScript doesn't edit the raw HTML text. The browser first parses the HTML into an internal tree structure — the DOM — and JS manipulates that tree.

DOM (Document Object Model) — The browser's internal, in-memory TREE representation of a parsed HTML document. JavaScript reads and modifies the DOM (not the original HTML text), and the browser renders the DOM to the user.

53. §18.9 The render pipeline

Concept

Putting it together, a page is rendered in a pipeline. HTML comes in as text; the browser builds the DOM; JavaScript modifies the DOM; the browser renders the final DOM to the screen.

HTML text  --parse-->  DOM tree  --JS modifies-->  DOM'  --render-->  pixels

Every change JavaScript makes happens on the DOM tree, and the user only ever sees the rendered DOM — not the original HTML the server sent.

54. §18.9 HTML is the recipe; the DOM is the cake

Intuition

The HTML text is like a recipe the server hands over. The browser 'bakes' it into the DOM — a live, structured object. JavaScript can then change the cake after it's baked (add a layer, restyle the frosting).

If you only read the recipe (the raw HTML), you might miss what JS did afterward. The user eats the cake (the rendered DOM), which can differ from the recipe.

Ask yourself: when an attacker injects a <script>, does it attack the recipe or the cake? (Both — it rides in as HTML, then runs and rewrites the live DOM that the user actually sees.)

55. What rests on this: §18.9 HTML is the recipe; the DOM is the cake

Socratic

Discussion prompt

If you only read the recipe (the raw HTML), you might miss what JS did afterward. The user eats the cake (the rendered DOM), which can differ from the recipe.

Suppose that were not true. What is the first thing in L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript that would stop working?

Hint: Follow it one step downstream. The answer is whatever was quietly relying on it.

Answer:

Ask yourself: when an attacker injects a <script>, does it attack the recipe or the cake? (Both — it rides in as HTML, then runs and rewrites the live DOM that the user actually sees.)

56. §18.9 JavaScript runs in a sandbox

Concept

Because page JS is so powerful, browsers run it in a sandbox: page JavaScript can't touch your files or read other sites' data. It's confined to its own page's world.

JavaScript sandbox — The restricted environment the browser runs page JavaScript in. Sandboxed JS cannot access the local filesystem or other origins' data — it can only act within its own page/origin.

The sandbox is why visiting a web page is (mostly) safe even though the page runs arbitrary code: that code is boxed in.

57. Teach it back: §18.9 JavaScript runs in a sandbox

Explain it

Discussion prompt

Explain §18.9 JavaScript runs in a sandbox to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Because page JS is so powerful, browsers run it in a sandbox: page JavaScript can't touch your files or read other sites' data. It's confined to its own page's world.

58. §18.9 JS is the browser's attack surface

Concept

Most browser exploits need JavaScript. Either the JS engine itself is the target (a bug in how it runs code), or JS is used to shape memory so a separate bug becomes exploitable.

And JS is JIT-compiled to machine code for speed — the just-in-time compiler is itself complex attack surface. So 'just turn off JavaScript' really does shrink the attack surface, at the cost of a broken web.

59. By analogy: §18.9 JS is the browser's attack surface

Analogy

Discussion prompt

Explain §18.9 JS is the browser's attack surface by analogy to something with no Computer Security in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Most browser exploits need JavaScript. Either the JS engine itself is the target (a bug in how it runs code), or JS is used to shape memory so a separate bug becomes exploitable.

60. What has to happen first: §18.9 HTML → DOM → JS modifies DOM → render

Ranking

Put in order

Put the moves of §18.9 HTML → DOM → JS modifies DOM → render into the order they have to happen.

  1. Server sends HTML text with an empty greeting span and a script
  2. The browser parses the HTML into a DOM tree
  3. JavaScript runs and modifies the DOM node's text
  4. Verify: the screen shows 'Hi, Alice', which never appeared in the server's HTML

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. This is the recipe the browser receives — nothing is rendered yet.

61. §18.9 HTML → DOM → JS modifies DOM → render

Worked example

Server sends HTML text with an empty greeting span and a script

Why: This is the recipe the browser receives — nothing is rendered yet.

<p id="greet"></p>
<script>
  document.getElementById('greet').textContent = 'Hi, Alice';
</script>

The browser parses the HTML into a DOM tree

Why: The <p> becomes a node in the tree; at this point its text content is empty.

JavaScript runs and modifies the DOM node's text

Why: getElementById('greet') finds the node; setting .textContent changes the live DOM, not the original HTML text.

stagethe <p id=greet> nodeon screen
HTML receivedemptynothing yet
parsed to DOMnode exists, text = ''blank paragraph
JS modifies DOMtext = 'Hi, Alice'(not yet repainted)
render DOMtext = 'Hi, Alice'shows: Hi, Alice

Verify: the screen shows 'Hi, Alice', which never appeared in the server's HTML

Why: §18.9: the user sees the RENDERED DOM after JS modified it — not the original HTML, which had an empty <p>. Reading the raw HTML alone would miss the greeting entirely.

62. Which is which, by the <p id=greet> node

Discrimination

Sort into buckets

Sort these by the <p id=greet> node, from memory, without looking back at §18.9 HTML → DOM → JS modifies DOM → render. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

empty
HTML received
node exists, text = ''
parsed to DOM
text = 'Hi, Alice'
JS modifies DOM; render DOM
g1
the <p id=greet> node is "empty" for HTML received — that is what the table on "§18.9 HTML → DOM → JS modifies DOM →…" records, and it is the single property separating this group from the rest.
g2
the <p id=greet> node is "node exists, text = ''" for parsed to DOM — that is what the table on "§18.9 HTML → DOM → JS modifies DOM →…" records, and it is the single property separating this group from the rest.
g3
the <p id=greet> node is "text = 'Hi, Alice'" for JS modifies DOM, render DOM — that is what the table on "§18.9 HTML → DOM → JS modifies DOM →…" records, and it is the single property separating this group from the rest.

63. Trap: 'page JavaScript can read any file on your computer'

Trap

The trap

A student: 'A web page runs JavaScript, so a malicious site could read my files, like C:\\secrets.txt.'

Assume page JS has filesystem access

Why: Wrong. Page JavaScript runs in a SANDBOX: it cannot touch your local files or read other sites' data. If it could read arbitrary files, every site you visited could steal your documents.

The fix

A student: what is page JavaScript actually allowed to touch?

JS is confined to its own page/origin by the sandbox

Why: §18.9: sandboxed JS can manipulate its own page's DOM but cannot reach the filesystem or other origins. Browser exploits that DO escape usually need a separate JS-engine bug, not a feature of normal JS.

64. Something is wrong here: 'the DOM is the same as the raw HTML text'

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student: 'I read the page's HTML source, so I've seen exactly what the user sees — the DOM is just that text.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The DOM is the PARSED TREE the browser built, and JavaScript may have modified it after parsing.

A student: what's the relationship between HTML and the DOM?

Why: The DOM is the PARSED TREE the browser built, and JavaScript may have modified it after parsing. The rendered DOM can differ from the original HTML — content can appear, vanish, or change.

65. Trap: 'the DOM is the same as the raw HTML text'

Trap

The trap

A student: 'I read the page's HTML source, so I've seen exactly what the user sees — the DOM is just that text.'

Equate the raw HTML with what's on screen

Why: Wrong. The DOM is the PARSED TREE the browser built, and JavaScript may have modified it after parsing. The rendered DOM can differ from the original HTML — content can appear, vanish, or change.

The fix

A student: what's the relationship between HTML and the DOM?

HTML is parsed INTO the DOM tree, which JS then manipulates

Why: §18.9: HTML text → parse → DOM → JS modifies DOM → render. To see what the user actually gets, inspect the live DOM (dev tools), not just the raw HTML source.

66. Why This Model Powers the Attacks

Section

Part 5 · the bridge to L37–L42

67. Each piece maps to a coming attack

Concept

Everything in this lesson is load-bearing for the rest of the web unit. Here's the map from mechanism to attack.

mechanism (this lesson)enableslesson
origin = scheme + host + port (from the URL)Same-Origin PolicyL37
HTTP is stateless → cookies carry identitysessionsL38
cookies ride on every HTTP requestCSRFL39
JS can modify the DOM + injected <script>XSSL40
iframes + frame isolation (geometry gap)clickjackingL41

68. Fill in: enables for Each piece maps to a coming attack

Comparison

Comparison matrix

From Each piece maps to a coming attack: refill the enables column from what you know. The rest of the table is as it appeared.

mechanism (this lesson)enableslesson
origin = scheme + host + port (from the URL)Same-Origin PolicyL37
HTTP is stateless → cookies carry identitysessionsL38
cookies ride on every HTTP requestCSRFL39
JS can modify the DOM + injected <script>XSSL40
iframes + frame isolation (geometry gap)clickjackingL41

69. Statelessness is why cookies exist

Intuition

HTTP has no built-in memory: each request stands alone, with no notion of 'who sent the last one.' That's statelessness.

So to stay logged in, the browser attaches a cookie to each request — that's how the server recognizes you across requests. Sessions (L38) and CSRF (L39) both flow directly from this patch.

Ask yourself: if cookies ride on EVERY request to a site, what happens when a different site triggers a request to it? (Your cookie goes along — which is exactly the CSRF problem.)

70. Break it if you can: Statelessness is why cookies exist

Counterexample

Discussion prompt

HTTP has no built-in memory: each request stands alone, with no notion of 'who sent the last one.' That's statelessness.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Ask yourself: if cookies ride on EVERY request to a site, what happens when a different site triggers a request to it? (Your cookie goes along — which is exactly the CSRF problem.)

71. ⊕ Supplemental tools to flag (beyond §18)

Concept

⊕ Supplemental — beyond §18.1–18.9. A few practical pieces you'll meet when we use this model in attacks and defenses.

72. Something is wrong here: 'HTTP remembers who you are between requests'

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student: 'Once I log in, the server knows it's me on the next request — HTTP keeps the connection's identity.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: HTTP is STATELESS: each request is independent and carries no memory of previous ones.

A student: how does the server actually recognize a returning user?

Why: HTTP is STATELESS: each request is independent and carries no memory of previous ones. Identity has to be re-supplied every time — there's nothing in the protocol that 'remembers' you.

73. Trap: 'HTTP remembers who you are between requests'

Trap

The trap

A student: 'Once I log in, the server knows it's me on the next request — HTTP keeps the connection's identity.'

Assume HTTP itself tracks the logged-in user

Why: Wrong. HTTP is STATELESS: each request is independent and carries no memory of previous ones. Identity has to be re-supplied every time — there's nothing in the protocol that 'remembers' you.

The fix

A student: how does the server actually recognize a returning user?

Identity rides in a cookie attached to each request

Why: Because HTTP is stateless, the browser re-sends a session cookie on every request (L38). That the cookie auto-attaches even to cross-site-triggered requests is the root of CSRF (L39).

74. Which of these survive contact with L36 · Web Basics: URLs, HTTP, HTML, the DOM…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Take the running example http://www.example.com/index.html. Read it left to right and it splits into exactly three parts.; The protocol is the postal service you use (regular vs certified mail) — it decides how the delivery is handled, here HTTP vs HTTPS.; The very first line of a request carries three things: the method, the path, and the version. After that come the request headers, one per line.
Breaks
A student: 'The path /index.html is what tells the browser where to go.'; A student: 'GET is read-only by the spec, so a GET request can never modify anything on the server.'
sound
These are stated as this lesson states them — each one survives the edge cases L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

75. Without one step: The web-model playbook

Constraint

Discussion prompt

Run The web-model playbook with this step confiscated:

Render pipeline: server response (with a media type) → HTML → browser parses to the DOM tree → JavaScript arbitrarily modifies the DOM → browser renders. The user sees the DOM, not the raw HTML.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Parse the URL: protocol (HTTP/HTTPS) · location/domain (which server) · path (which resource). The DOMAIN picks the server; optional user@ and :port…
  2. Read the HTTP request: first line = method + path + version; then headers (Host: repeats the domain); blank line; optional body.
  3. GET vs POST: GET gets info and should be side-effect-free, data in URL query params, NO body. POST sends state-changing data in a BODY. Only POST has a…
  4. Render pipeline: server response (with a media type) → HTML → browser parses to the DOM tree → JavaScript arbitrarily modifies the DOM → browser renders…
  5. Containment: frames are isolated (no cross-frame read/write, but the outer page controls geometry); page JS is sandboxed (no files, no other origins). CSS…
  6. Map to attacks: origin (scheme+host+port) → SOP (L37); stateless HTTP + cookies → sessions/CSRF (L38/L39); DOM + injected <script> → XSS (L40); iframe…

76. The web-model playbook

Pattern

  1. Parse the URL: protocol (HTTP/HTTPS) · location/domain (which server) · path (which resource). The DOMAIN picks the server; optional user@ and :port extend the location.
  2. Read the HTTP request: first line = method + path + version; then headers (Host: repeats the domain); blank line; optional body.
  3. GET vs POST: GET gets info and should be side-effect-free, data in URL query params, NO body. POST sends state-changing data in a BODY. Only POST has a body.
  4. Render pipeline: server response (with a media type) → HTML → browser parses to the DOM tree → JavaScript arbitrarily modifies the DOM → browser renders. The user sees the DOM, not the raw HTML.
  5. Containment: frames are isolated (no cross-frame read/write, but the outer page controls geometry); page JS is sandboxed (no files, no other origins). CSS is as powerful as JS when forced on a victim.
  6. Map to attacks: origin (scheme+host+port) → SOP (L37); stateless HTTP + cookies → sessions/CSRF (L38/L39); DOM + injected <script> → XSS (L40); iframe geometry → clickjacking (L41).

77. Where does it stop working: The web-model playbook

Edge cases

Discussion prompt

The web-model playbook works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Parse the URL: protocol (HTTP/HTTPS) · location/domain (which server) · path (which resource). The DOMAIN picks the server; optional user@ and :port…
  2. Read the HTTP request: first line = method + path + version; then headers (Host: repeats the domain); blank line; optional body.
  3. GET vs POST: GET gets info and should be side-effect-free, data in URL query params, NO body. POST sends state-changing data in a BODY. Only POST has a…
  4. Render pipeline: server response (with a media type) → HTML → browser parses to the DOM tree → JavaScript arbitrarily modifies the DOM → browser renders…
  5. Containment: frames are isolated (no cross-frame read/write, but the outer page controls geometry); page JS is sandboxed (no files, no other origins). CSS…
  6. Map to attacks: origin (scheme+host+port) → SOP (L37); stateless HTTP + cookies → sessions/CSRF (L38/L39); DOM + injected <script> → XSS (L40); iframe…

78. Rule out three: Checkpoint — which statement is TRUE?

Elimination

Eliminate the wrong options

Which statement about this page is TRUE?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The path /cart is what tells the browser which server to contact.
  • B. The login credentials must travel in the body of a GET request.
  • C. The page's JavaScript can read arbitrary files like C:\\secrets.txt from the user's disk.
  • D. The browser parses the HTML into the DOM, JavaScript modifies that DOM tree, and the user sees the rendered DOM — which can differ from the original HTML.

Survives elimination: D

Why: §18.6–18.9: the browser parses the server's HTML into the DOM (an internal tree), JavaScript arbitrarily modifies that DOM, and the browser renders the modified DOM to the user — so what's on screen can differ from the raw HTML the server sent. The other claims fail: the LOCATION/domain (shop.example.com), not the path, selects the server; only POST has a body, so credentials go in a POST body (a GET would use URL query params, not a body); and page JavaScript runs in a sandbox that cannot touch the local filesystem.

79. Checkpoint — which statement is TRUE?

Check

A browser loads https://shop.example.com/cart?id=7, the page's JavaScript rewrites part of the page, and a login form POSTs a password. Think through URL parts, GET vs POST, the DOM, and the JS sandbox before choosing.

Check your understanding

Which statement about this page is TRUE?

  • A. The path /cart is what tells the browser which server to contact.
  • B. The login credentials must travel in the body of a GET request.
  • C. The page's JavaScript can read arbitrary files like C:\\secrets.txt from the user's disk.
  • D. The browser parses the HTML into the DOM, JavaScript modifies that DOM tree, and the user sees the rendered DOM — which can differ from the original HTML. (correct)

Answer: D

Why: §18.6–18.9: the browser parses the server's HTML into the DOM (an internal tree), JavaScript arbitrarily modifies that DOM, and the browser renders the modified DOM to the user — so what's on screen can differ from the raw HTML the server sent. The other claims fail: the LOCATION/domain (shop.example.com), not the path, selects the server; only POST has a body, so credentials go in a POST body (a GET would use URL query params, not a body); and page JavaScript runs in a sandbox that cannot touch the local filesystem.

Why A tempts people
The path is requested FROM a server that's already chosen by the location/domain (shop.example.com). The path /cart is meaningless until the browser has reached that server — it never selects the machine. (Origin in L37 is scheme+host+port, never the path.)
Why B tempts people
Only POST has a body; a GET carries data via query parameters in the URL, not a body. Credentials are sent in a POST request body precisely so they don't sit in the URL (visible in the address bar and logs).
Why C tempts people
Page JavaScript runs in a sandbox: it can manipulate its own page's DOM but cannot read the local filesystem or other origins' data. If it could read arbitrary files, every site you visited could steal your documents.

80. Misconceptions to retire

Concept

81. Synthesis — the browser is a TCB running untrusted code

Concept

82. Primary sources & where to read more

Concept

83. Connect it up: L36 · Web Basics: URLs, HTTP, HTML, the DOM & JavaScript

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — URLs — Naming Every Resource · HTTP — The Request/Response Model · The Webpage as a Distributed App + HTML · CSS, JavaScript & the DOM · Why This Model Powers the Attacks. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

84. Recap — Lesson 36

Recap

You can now split a URL into protocol/location/path and name which part picks the server, read an HTTP request and tell GET from POST (only POST has a body), explain the page as a distributed app with isolated frames, trace HTML → DOM → JS-modifies-DOM → render, and say why page JS is sandboxed — and map every piece to the web attacks ahead.

Idea§The one-line version
URL parts18.1protocol · location/domain (picks the server) · path
HTTP model18.2–18.4client sends request (method/path/version + headers); server responds
GET vs POST18.5GET = info, query params, NO body; POST = state change, body
Distributed app18.6server component + browser component; media type drives interpretation
HTML & frames18.7<script> runs JS; frames are isolated (no cross-frame access)
DOM pipeline18.9HTML → parse → DOM → JS modifies → render; user sees the DOM
Sandbox18.9page JS can't read files or other origins; JIT is attack surface
Bridge—origin→SOP, cookies→sessions/CSRF, <script>→XSS, iframe→clickjacking

Sources

  1. CS 161 Computer Security Textbook §18.1–18.9 — Wagner, Weaver, Kao, Shakir, Law & Ngai, UC Berkeley — URLs and their three parts (§18.1), the HTTP request/response model and HTTP/1.1 vs HTTP/2 (§18.2–18.3), HTTP requests and headers (§18.4), GET vs POST (§18.5), the webpage as a distributed application (§18.6), HTML elements and frame isolation (§18.7), CSS as powerful as JavaScript (§18.8), and JavaScript, the DOM, the sandbox and JIT (§18.9)
  2. RFC 9110 — HTTP Semantics — R. Fielding, M. Nottingham & J. Reschke (eds.), IETF, June 2022 — defines HTTP methods (incl. the safe/idempotent semantics of GET), request/response messages, and headers
  3. RFC 3986 — Uniform Resource Identifier (URI): Generic Syntax — T. Berners-Lee, R. Fielding & L. Masinter, IETF, January 2005 — defines the scheme / authority (userinfo, host, port) / path / query / fragment structure of URIs and URLs
  4. WHATWG DOM Living Standard — WHATWG — defines the Document Object Model: the parsed tree the browser builds from HTML and that JavaScript manipulates
  5. MDN Web Docs — HTTP overview & Introduction to the DOM ⊕ — Mozilla — supplemental reference for the HTTP request/response model, methods, and the DOM as the in-memory tree representation of a document

Want this taught 1-on-1? Alexander tutors Computer Security — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108