USAAIO Lesson 94, from Phase 3. It covers message passing on graphs and derives the normalized adjacency D^{-1/2} Ã D^{-1/2} from first principles on a 4-cycle, then works a GCN forward pass with real computed values. It covers Graph Attention Networks, with the attention weights verified to sum to 1, and global readout pooling by mean, max, or sum, ending with a TinyGCN of 48 parameters trained to 100% accuracy on a synthetic 8-node community graph. All the numbers were verified with torch 2.7.1+cpu. The lesson runs to 28 slides.
Subject: Machine Learning · 53 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 94 · Phase 3
Message passing on graphs: normalize the adjacency, aggregate neighbor features, optionally weight edges with learned attention. Every formula derived and every number verified in PyTorch — from a 4-node cycle to a trained community-graph classifier.
Objectives
h_v^(l+1) = UPDATE(h_v^(l), AGGREGATE({m_{uv}})) and name what GCN substitutes for each pieceD^{-1/2} Ã D^{-1/2} from scratch and compute it on a toy graphH^(l+1) = σ(Ã_norm H^(l) W^(l))) with real tensor shapes and valuesα_{uv} and verify that they sum to 1 over each neighborhoodWarm-up
Discussion prompt
Before we open Lesson 94: Graph Neural Networks (GCN & GAT): without looking back, what was the main idea of BERT Variants — XLNet, RoBERTa, T5, DeBERTa, ALBERT, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
why bidirectional BERT is slow at inference and how each variant fixes a different limitation — XLNet permutation LM (two-stream attention, content vs query masks), RoBERTa (dynamic masking, no NSP, larger batches), T5 text-to-text unified seq2seq framing, DeBERTa disentangled content+position attention, and ALBERT cross-layer parameter sharing (6x reduction verified in PyTorch at d=64).
Section
Part 1 of 4
Concept
A graph G = (V, E) pairs a set of nodes V (each carrying a feature vector h_v) with edges E (which may carry features e_{uv}). Tasks: node classification (label each node), link prediction (predict edges), graph classification (one label per graph).
| graph | nodes | edges | node features | typical task |
|---|---|---|---|---|
| social network | users | friendships | profile text embeddings | node classification |
| molecular graph | atoms | bonds | atom type, charge | graph classification |
| citation network | papers | citations | bag-of-words TF-IDF | node classification |
| knowledge graph | entities | relations | type one-hot | link prediction |
Discrimination
Sort into buckets
Sort these by typical task, from memory, without looking back at Graphs: nodes, edges, features. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
All major GNNs (GCN, GAT, GraphSAGE, GIN) are instances of the same message-passing template — only the choices of MESSAGE, AGGREGATE, and UPDATE differ.
\[ m_{uv}^{(l)} = \text{MESSAGE}(h_u^{(l)},\; h_v^{(l)},\; e_{uv}) \]
\[ h_v^{(l+1)} = \text{UPDATE}\!\left(h_v^{(l)},\; \text{AGGREGATE}\bigl(\{m_{uv}^{(l)} : u \in \mathcal{N}(v)\}\bigr)\right) \]
| model | MESSAGE | AGGREGATE | UPDATE |
|---|---|---|---|
| GCN | W h_u (linear) | normalized sum | σ(·) |
| GAT | W h_u | attention-weighted sum | σ(·) |
| GraphSAGE | W h_u | mean/max/LSTM | concat + σ(·) |
| GIN | MLP(h_u) | sum (injective) | MLP(·) |
Comparison
Comparison matrix
From The message-passing framework: refill the AGGREGATE column from what you know. The rest of the table is as it appeared.
| model | MESSAGE | AGGREGATE | UPDATE |
|---|---|---|---|
| GCN | W h_u (linear) | normalized sum | σ(·) |
| GAT | W h_u | attention-weighted sum | σ(·) |
| GraphSAGE | W h_u | mean/max/LSTM | concat + σ(·) |
| GIN | MLP(h_u) | sum (injective) | MLP(·) |
Section
Part 2 of 4
Concept
Naively summing neighbor features gives high-degree nodes much larger activation magnitudes than low-degree nodes — a scale instability that makes training hard. GCN (Kipf & Welling 2017) fixes this with a symmetric normalization.
\[ \tilde{A} = A + I_N \quad\text{(add self-loops so each node aggregates itself)} \]
\[ \hat{A} = \tilde{D}^{-1/2}\, \tilde{A}\, \tilde{D}^{-1/2} \qquad \tilde{D}_{ii} = \sum_j \tilde{A}_{ij} \]
The (u,v) entry of  equals 1 / sqrt(d̃_u · d̃_v). Each edge contribution is scaled by the geometric mean of the degree-with-self-loop of both endpoints — so high-degree hubs are down-weighted and low-degree nodes are not dominated.
Estimation
Predict first
Graph: nodes 0-3, edges {(0,1),(1,2),(2,3),(0,3)} — a 4-cycle. Add self-loops to get Ã. Compute degree matrix D̃ and then  = D̃^{-1/2} à D̃^{-1/2}. Predict each diagonal entry before proceeding.
Commit before you compute: what does Compute  on a 4-cycle come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Â[i,i] = 1/sqrt(3·3) = 1/3 ≈ 0.3333; Â[i,j] (connected) = 1/sqrt(3·3) ≈ 0.3333 too
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Because all degrees are equal (regular graph), all non-zero entries of  are the same: 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/3.
Worked example
Graph: nodes 0-3, edges {(0,1),(1,2),(2,3),(0,3)} — a 4-cycle. Add self-loops to get Ã. Compute degree matrix D̃ and then  = D̃^{-1/2} à D̃^{-1/2}. Predict each diagonal entry before proceeding.
import torch
A = torch.tensor([
[0,1,0,1],
[1,0,1,0],
[0,1,0,1],
[1,0,1,0]
], dtype=torch.float32)
A_tilde = A + torch.eye(4) # add self-loops
d = A_tilde.sum(dim=1) # [3, 3, 3, 3]
D_inv_sqrt = torch.diag(d ** -0.5) # 1/sqrt(3) on diagonal
A_norm = D_inv_sqrt @ A_tilde @ D_inv_sqrt
print('d:', d.tolist())
print('A_norm:\n', A_norm.round(decimals=4))Each node has degree 2 (two cycle edges); adding self-loop makes d̃ = 3 for all nodes
Why: Ã row sums: node 0 has edges to 1, 3, and itself → 3 non-zeros → d̃_0 = 3. Symmetric graph so all degrees equal.
Â[i,i] = 1/sqrt(3·3) = 1/3 ≈ 0.3333; Â[i,j] (connected) = 1/sqrt(3·3) ≈ 0.3333 too
Why: Because all degrees are equal (regular graph), all non-zero entries of  are the same: 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/3.
| entry | formula | value |
|---|---|---|
| d̃_0 = d̃_1 = d̃_2 = d̃_3 | 2 edges + self-loop | 3 |
| D̃^{-1/2} diagonal | 1/sqrt(3) | 0.5774 |
| Â[i,i] (self) | 1/sqrt(3·3) = 1/3 | 0.3333 |
| Â[i,j] (connected) | 1/sqrt(3·3) = 1/3 | 0.3333 |
| Â[i,j] (disconnected) | 0 (no edge) | 0.0000 |
Trade off
Comparison matrix
From Compute  on a 4-cycle: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| entry | formula | value |
|---|---|---|
| d̃_0 = d̃_1 = d̃_2 = d̃_3 | 2 edges + self-loop | 3 |
| D̃^{-1/2} diagonal | 1/sqrt(3) | 0.5774 |
| Â[i,i] (self) | 1/sqrt(3·3) = 1/3 | 0.3333 |
| Â[i,j] (connected) | 1/sqrt(3·3) = 1/3 | 0.3333 |
| Â[i,j] (disconnected) | 0 (no edge) | 0.0000 |
Estimation
Predict first
One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.
Commit before you compute: what does GCN forward pass: Â H W then ReLU come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For node 0, neighbors are {1, 3} plus self, so row 0 of A_norm@H = (h_0 + h_1 + h_3)/3.
Worked example
One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.
import torch, torch.nn.functional as F
torch.manual_seed(42)
H = torch.randn(4, 2) # node features (4 nodes, 2 dims)
W = torch.randn(2, 3) # weight matrix (project to 3 dims)
# A_norm from previous slide
A_norm = torch.full((4,4), 1/3) * torch.tensor([
[1,1,0,1],[1,1,1,0],[0,1,1,1],[1,0,1,1]
], dtype=torch.float32)
H1 = F.relu(A_norm @ H @ W)
print('H shape:', H.shape, ' W shape:', W.shape)
print('A_norm @ H @ W before ReLU shape:', (A_norm @ H @ W).shape)
print('H^(1):\n', H1.round(decimals=4))A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)
Why: For node 0, neighbors are {1, 3} plus self, so row 0 of A_norm@H = (h_0 + h_1 + h_3)/3. This is symmetric mean-of-neighbors aggregation weighted by degree.
| node | H^(1)[node, :] | note |
|---|---|---|
| 0 | [0.3525, 0.1445, 0.6526] | neighbors: {0,1,3}; all pre-ReLU > 0 |
| 1 | [0.0000, 0.0148, 0.0000] | neighbors: {0,1,2}; mostly clipped by ReLU |
| 2 | [0.0428, 0.0000, 0.5699] | neighbors: {1,2,3} |
| 3 | [0.0312, 0.0000, 0.6453] | neighbors: {0,2,3} |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use the raw adjacency A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself.
Always add self-loops first: Ã = A + I. Compute degree on Ã, then normalize: Â = D̃^{-1/2} Ã D̃^{-1/2}.
Why: Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.
Trap
Use the raw adjacency A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.
A_norm = D_inv_sqrt @ A @ D_inv_sqrt — skipping A + I
Why: This is wrong. Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.
Always add self-loops first: Ã = A + I. Compute degree on Ã, then normalize: Â = D̃^{-1/2} Ã D̃^{-1/2}.
A_tilde = A + torch.eye(n); d = A_tilde.sum(1); A_norm = diag(d-0.5) @ A_tilde @ diag(d-0.5)
Why: With self-loops, row i of ·H includes h_i itself at weight 1/d̃_i. This ensures the node's own representation propagates forward — critical for learning node-level properties.
Break the constraint
Discussion prompt
The rule this trap just fixed:
With self-loops, row i of ·H includes h_i itself at weight 1/d̃_i. This ensures the node's own representation propagates forward — critical for learning node-level properties.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.
Section
Part 3 of 4
Concept
GCN treats all neighbors equally: weight = 1/sqrt(d̃_u · d̃_v). GAT (Velickovic 2018) lets the model learn how much attention to pay each neighbor, conditioned on the features of both endpoints.
\[ e_{uv} = \text{LeakyReLU}\!\left(\mathbf{a}^\top [\mathbf{W}h_u \| \mathbf{W}h_v]\right) \]
\[ \alpha_{uv} = \text{softmax}_u(e_{uv}) = \frac{\exp(e_{uv})}{\sum_{k \in \mathcal{N}(v) \cup \{v\}} \exp(e_{kv})} \]
The softmax is over the neighborhood of v (including v itself). So Σ_{u ∈ N(v)∪{v}} α_{uv} = 1 for every node v — the attention weights are a proper probability distribution over neighbors.
Counterexample
Discussion prompt
GCN treats all neighbors equally: weight = 1/sqrt(d̃_u · d̃_v). GAT (Velickovic 2018) lets the model learn how much attention to pay each neighbor, conditioned on the features of both endpoints.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The softmax is over the neighborhood of v (including v itself). So Σ_{u ∈ N(v)∪{v}} α_{uv} = 1 for every node v — the attention weights are a proper probability distribution over neighbors.
Estimation
Predict first
Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.
Commit before you compute: what does GAT attention weights on a toy neighborhood come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each e_{0,j} = LeakyReLU(a^T [Wh_0 || Wh_j]).
Worked example
Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.
import torch, torch.nn.functional as F
torch.manual_seed(42)
H_gat = torch.tensor([[0.1,0.2,0.3],[0.4,0.5,0.6],
[0.7,0.8,0.9],[0.2,0.3,0.1]], dtype=torch.float32)
# Wh = H @ W_gat (3->3 linear, random W)
W_gat = torch.randn(3, 3)
a_vec = torch.randn(6) # attention param: 2*d_out = 6
Wh = H_gat @ W_gat # (4, 3)
def attn_score(hi, hj):
return F.leaky_relu((a_vec @ torch.cat([hi, hj])).unsqueeze(0),
negative_slope=0.2).item()
# Node 0 neighbors: {0,1,3}
neighbors = [0, 1, 3]
raw = [attn_score(Wh[0], Wh[j]) for j in neighbors]
alpha = F.softmax(torch.tensor(raw), dim=0)
print('raw e_{0,j}:', [round(x,4) for x in raw])
print('alpha_{0,j}:', alpha.round(decimals=4).tolist())
print('sum:', alpha.sum().item())Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771
Why: Each e_{0,j} = LeakyReLU(a^T [Wh_0 || Wh_j]). Neighbor 1 gets the highest raw score — the model discovers that node 1's features are most aligned with node 0's attention vector.
| neighbor j | raw score e_{0,j} | α_{0,j} (softmax) | interpretation |
|---|---|---|---|
| 0 (self) | 1.4656 | 0.1507 | self-attention: 15% weight |
| 1 | 3.0661 | 0.7470 | dominant neighbor: 74.7% weight |
| 3 | 1.0771 | 0.1022 | weakest neighbor: 10.2% weight |
| sum | — | 1.0000 | valid probability distribution |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.
Mask non-neighbors with -inf before softmax so they receive exactly zero attention weight.
Why: Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.
Trap
Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).
alpha = F.softmax(e_all_nodes, dim=0) — no masking of non-neighbors
Why: This is wrong. Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.
Mask non-neighbors with -inf before softmax so they receive exactly zero attention weight.
e_masked = e.masked_fill(adj_mask == 0, float('-inf')); alpha = F.softmax(e_masked, dim=1)
Why: softmax(-inf) = 0. Only neighbors (and self) have finite scores, so all weight concentrates on them. Row sums are 1 only over the valid neighborhood — verified by alpha.sum(dim=1) all equals 1.0 in Part 8 of the verification script.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
MESSAGE, AGGREGATE, and UPDATE differ.; Node classification uses H^(L)[v] directly. For graph classification, you need one fixed-size vector per graph — a readout (global pooling) over all node embeddings.; Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.; Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).Section
Part 4 of 4
Concept
Node classification uses H^(L)[v] directly. For graph classification, you need one fixed-size vector per graph — a readout (global pooling) over all node embeddings.
\[ h_G = \text{READOUT}\bigl(\{h_v^{(L)} : v \in V\}\bigr) \]
| readout | formula | properties | differentiable? |
|---|---|---|---|
| mean pool | mean over rows of H^(L) | permutation-invariant, bounded | yes |
| max pool | max per dim over all rows | selects most activated feature per dim | yes (subgradient) |
| sum pool | sum over rows of H^(L) | sensitive to graph size | yes |
| attn pool | softmax(gate)^T H^(L) | learned soft gate over nodes | yes |
Comparison
Comparison matrix
From From node embeddings to a graph vector: refill the differentiable? column from what you know. The rest of the table is as it appeared.
| readout | formula | properties | differentiable? |
|---|---|---|---|
| mean pool | mean over rows of H^(L) | permutation-invariant, bounded | yes |
| max pool | max per dim over all rows | selects most activated feature per dim | yes (subgradient) |
| sum pool | sum over rows of H^(L) | sensitive to graph size | yes |
| attn pool | softmax(gate)^T H^(L) | learned soft gate over nodes | yes |
Constraint
Discussion prompt
Run The GCN/GAT recipe with this step confiscated:
Depth: 2-3 layers typical; L layers → L-hop neighborhood receptive field
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
à = A + I; degree d̃_i = Σ_j Ã_{ij};  = D̃^{-1/2} à D̃^{-1/2}H^(l+1) = σ( H^(l) W^(l)) — one matrix multiply, no Python loop over nodesW (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges…mean/max/sum over H^(L) → fixed-size graph vector → MLP classifierPattern
à = A + I; degree d̃_i = Σ_j Ã_{ij};  = D̃^{-1/2} à D̃^{-1/2}H^(l+1) = σ( H^(l) W^(l)) — one matrix multiply, no Python loop over nodesW (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges with -inf); aggregate h_v^{(l+1)} = σ(Σ_{u∈N(v)} α_{uv} W h_u)mean/max/sum over H^(L) → fixed-size graph vector → MLP classifierEdge cases
Discussion prompt
The GCN/GAT recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
à = A + I; degree d̃_i = Σ_j Ã_{ij};  = D̃^{-1/2} à D̃^{-1/2}H^(l+1) = σ( H^(l) W^(l)) — one matrix multiply, no Python loop over nodesW (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges…mean/max/sum over H^(L) → fixed-size graph vector → MLP classifierElimination
Eliminate the wrong options
In a graph where node u has degree 4 (after adding self-loop, d̃_u = 5) and node v has degree 2 (d̃_v = 3), and (u,v) is an edge, what is Â[u,v] in the GCN normalized adjacency?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Â[u,v] = 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/sqrt(5×3) = 1/sqrt(15) ≈ 0.2582. The formula uses the degree after adding self-loops (d̃), not the original degree (d).
Check
Work out the formula before selecting.
Check your understanding
In a graph where node u has degree 4 (after adding self-loop, d̃_u = 5) and node v has degree 2 (d̃_v = 3), and (u,v) is an edge, what is Â[u,v] in the GCN normalized adjacency?
Answer: A
Why: Â[u,v] = 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/sqrt(5×3) = 1/sqrt(15) ≈ 0.2582. The formula uses the degree after adding self-loops (d̃), not the original degree (d).
Prediction
Predict first
Node v has 3 neighbors (plus itself, so 4 total in N(v)∪{v}). After computing GAT attention weights α_{uv} for all u ∈ N(v)∪{v}, what must be true?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Σ_{u ∈ N(v)∪{v}} α_{uv} = 1
Why: The softmax in GAT is applied only over the neighborhood N(v)∪{v}, making the 4 weights a proper probability distribution that sums to 1. Non-neighbors receive exactly 0 weight (via -inf masking). This was verified in code: all row sums equal 1.0.
Check
Think about the softmax domain.
Check your understanding
Node v has 3 neighbors (plus itself, so 4 total in N(v)∪{v}). After computing GAT attention weights α_{uv} for all u ∈ N(v)∪{v}, what must be true?
Answer: A
Why: The softmax in GAT is applied only over the neighborhood N(v)∪{v}, making the 4 weights a proper probability distribution that sums to 1. Non-neighbors receive exactly 0 weight (via -inf masking). This was verified in code: all row sums equal 1.0.
Elimination
Eliminate the wrong options
You are building a citation-network node classifier. Each paper has a bag-of-words feature vector. Some papers cite far more relevant works than others. Which statement correctly describes when GAT has an advantage over GCN?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: GAT learns feature-conditioned attention weights α_{uv}, so it can assign high weight to topically relevant neighbors and low weight to off-topic citations — exactly what fixed degree-normalization cannot do. On Cora, GAT historically outperforms GCN by ~1–2% because the benefit of selective attention outweighs the added parameters.
Check
Apply the inductive bias reasoning.
Check your understanding
You are building a citation-network node classifier. Each paper has a bag-of-words feature vector. Some papers cite far more relevant works than others. Which statement correctly describes when GAT has an advantage over GCN?
Answer: A
Why: GAT learns feature-conditioned attention weights α_{uv}, so it can assign high weight to topically relevant neighbors and low weight to off-topic citations — exactly what fixed degree-normalization cannot do. On Cora, GAT historically outperforms GCN by ~1–2% because the benefit of selective attention outweighs the added parameters.
Section
Project
Concept
Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.
| # | milestone | key check |
|---|---|---|
| 1 | Build normalized adjacency Â; verify row sums ≤ 1 (irregular graph) | A_norm.sum(1) |
| 2 | Implement TinyGCN (2 layers, 48 params); trace shapes through each A_norm @ H @ W | print(H1.shape, H2.shape) |
| 3 | Train 200 epochs with Adam lr=0.01; hit 100% node accuracy on epoch≥50 | (logits.argmax(1)==y).float().mean() |
Build rules: precompute  once before the training loop (it is fixed); use bias=False in both Linear layers; verify that 2-hop message passing (2 GCN layers) is sufficient for nodes to receive information from across the bridge.
Counterexample
Discussion prompt
Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: construct the 8-node two-clique adjacency. Predict: are all row sums of  the same? (Hint: what is the degree of the bridge node vs an interior clique node?)
Hint: d = A_tilde.sum(1) gives a vector (not a scalar). Interior nodes have degree 7 + self-loop = 8; the bridge node has degree 7 + 1 bridge edge + self-loop = 9. So  is not a uniform matrix unlike the 4-cycle example.
import torch
torch.manual_seed(0)
n = 8
A = torch.zeros(n, n)
# Two 4-cliques
for i in range(4):
for j in range(4):
if i != j: A[i,j] = 1
for i in range(4, 8):
for j in range(4, 8):
if i != j: A[i,j] = 1
A[3,4] = A[4,3] = 1 # bridge between node 3 and node 4
A_tilde = A + torch.eye(n)
d = A_tilde.sum(1) # degree with self-loops
D_inv_sqrt = torch.diag(d**-0.5)
A_norm = D_inv_sqrt @ A_tilde @ D_inv_sqrt
print('degrees d_tilde:', d.tolist())
print('A_norm row sums:', A_norm.sum(1).round(decimals=4).tolist())| node | d̃ (self-loop) | role | Â row sum |
|---|---|---|---|
| 0, 1, 2 | 4 | interior clique-A | 1.0 (all edges normalized) |
| 3 | 5 | bridge node (clique-A side) | 1.0 |
| 4 | 5 | bridge node (clique-B side) | 1.0 |
| 5, 6, 7 | 4 | interior clique-B | 1.0 |
Worked example
Your turn: implement TinyGCN with layers 4→8→2. Predict the shape of the intermediate hidden matrix H1 and the final logits H2. How many total parameters?
Hint: nn.Linear(4, 8, bias=False) has 32 weights; nn.Linear(8, 2, bias=False) has 16 weights. Total = 48 parameters. The forward pass is: H1 = ReLU(Â @ W1(X)), then H2 = Â @ W2(H1) (no activation on the final layer — CrossEntropyLoss expects raw logits).
import torch, torch.nn as nn, torch.nn.functional as F
class TinyGCN(nn.Module):
def __init__(self, in_f, h_f, out_f):
super().__init__()
self.W1 = nn.Linear(in_f, h_f, bias=False)
self.W2 = nn.Linear(h_f, out_f, bias=False)
def forward(self, X, A_norm):
H1 = F.relu(A_norm @ self.W1(X)) # (n, h_f)
H2 = A_norm @ self.W2(H1) # (n, out_f)
return H2
torch.manual_seed(42)
X = torch.randn(8, 4) # 8 nodes, 4 features each
y = torch.tensor([0,0,0,0,1,1,1,1], dtype=torch.long)
model = TinyGCN(4, 8, 2)
logits = model(X, A_norm) # A_norm from milestone 1
print('params:', sum(p.numel() for p in model.parameters()))
print('H2 shape:', logits.shape)| layer | input shape | output shape | params |
|---|---|---|---|
| W1 (Linear 4→8) | (8, 4) | (8, 8) | 32 |
| ReLU + A_norm@ | (8, 8) | (8, 8) | 0 |
| W2 (Linear 8→2) | (8, 8) | (8, 2) | 16 |
| A_norm@ (readout) | (8, 2) | (8, 2) | 0 |
| TOTAL | — | — | 48 |
Discrimination
Sort into buckets
Sort these by output shape, from memory, without looking back at Milestone 2 — TinyGCN forward pass. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Your turn: run the training loop for 200 epochs. Predict: does the model converge faster or slower than a vanilla MLP on the same features? (Recall (Lesson 40): 5-step loop is zero_grad → forward → loss → backward → step.)
Hint: precompute A_norm outside the loop — it is a fixed graph property, not a model parameter. Use Adam lr=0.01 (higher than transformers, typical for GNNs on small graphs).
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
X = torch.randn(8, 4)
y = torch.tensor([0,0,0,0,1,1,1,1], dtype=torch.long)
model = TinyGCN(4, 8, 2) # defined in milestone 2
opt = torch.optim.Adam(model.parameters(), lr=0.01)
loss_fn = nn.CrossEntropyLoss()
for epoch in range(201):
opt.zero_grad()
logits = model(X, A_norm) # A_norm from milestone 1
loss = loss_fn(logits, y)
loss.backward()
opt.step()
if epoch % 50 == 0:
acc = (logits.argmax(1)==y).float().mean()
print(f'ep {epoch:3d} loss={loss.item():.4f} acc={acc*100:.0f}%')| epoch | loss | accuracy | note |
|---|---|---|---|
| 0 | 0.7640 | 0% | random init, predictions all wrong |
| 50 | 0.2027 | 100% | graph structure learned by epoch 50 |
| 100 | 0.0644 | 100% | loss still decreasing, accuracy saturated |
| 150 | 0.0280 | 100% | well separated logits |
| 200 | 0.0134 | 100% | converged — 48 params sufficient |
Trade off
Comparison matrix
From Milestone 3 — train to 100% accuracy: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.
| epoch | loss | accuracy | note |
|---|---|---|---|
| 0 | 0.7640 | 0% | random init, predictions all wrong |
| 50 | 0.2027 | 100% | graph structure learned by epoch 50 |
| 100 | 0.0644 | 100% | loss still decreasing, accuracy saturated |
| 150 | 0.0280 | 100% | well separated logits |
| 200 | 0.0134 | 100% | converged — 48 params sufficient |
Concept
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
# --- Graph: two 4-cliques + bridge ---
n = 8
A = torch.zeros(n, n)
for i in range(4):
for j in range(4):
if i!=j: A[i,j]=1
for i in range(4,8):
for j in range(4,8):
if i!=j: A[i,j]=1
A[3,4]=A[4,3]=1
A_t = A+torch.eye(n); d=A_t.sum(1)
A_norm = torch.diag(d**-0.5)@A_t@torch.diag(d**-0.5)
# --- Features & labels ---
X = torch.randn(n,4)
y = torch.tensor([0,0,0,0,1,1,1,1],dtype=torch.long)
# --- Model ---
class TinyGCN(nn.Module):
def __init__(self,i,h,o):
super().__init__()
self.W1=nn.Linear(i,h,bias=False)
self.W2=nn.Linear(h,o,bias=False)
def forward(self,X,A):
return A@self.W2(F.relu(A@self.W1(X)))
model=TinyGCN(4,8,2)
opt=torch.optim.Adam(model.parameters(),lr=0.01)
for _ in range(200):
opt.zero_grad()
loss=F.cross_entropy(model(X,A_norm),y)
loss.backward(); opt.step()
with torch.no_grad():
acc=(model(X,A_norm).argmax(1)==y).float().mean()
print(f'{acc.item()*100:.0f}% node accuracy') # 100%
print('params:', sum(p.numel() for p in model.parameters())) # 48| design choice | value | rationale |
|---|---|---|
| layers | 2 GCN | 2-hop receptive field crosses the bridge |
| hidden dim | 8 | small enough to train in <1 s on CPU |
| bias | False | simpler; normalization handles scale |
| lr | 0.01 | GNNs on small graphs tolerate higher lr than transformers |
| epochs | 200 | 100% accuracy reached by epoch 50 |
| total params | 48 | 4×8 + 8×2 = 32+16 |
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| design choice | value | rationale |
|---|---|---|
| layers | 2 GCN | 2-hop receptive field crosses the bridge |
| hidden dim | 8 | small enough to train in <1 s on CPU |
| bias | False | simpler; normalization handles scale |
| lr | 0.01 | GNNs on small graphs tolerate higher lr than transformers |
| epochs | 200 | 100% accuracy reached by epoch 50 |
| total params | 48 | 4×8 + 8×2 = 32+16 |
Concept
Out loud, slides closed: (1) derive the GCN normalized adjacency formula from scratch — explain each step and why you need both +I and both D^{-1/2} factors; (2) trace the attention weight computation for one node in GAT, stating why softmax is restricted to the neighborhood; (3) explain when you would choose mean-pool vs max-pool readout for graph classification.
Stretch (lesson-plan homework): implement SingleHeadGAT from scratch — learn W and a, mask non-edges with -inf, apply softmax per row, aggregate. Verify alpha.sum(1) is all 1.0. Then compare TinyGCN vs a linear MLP (same feature matrix, no adjacency) on the community graph — does removing graph structure hurt accuracy?
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Message passing — the GNN abstraction · GCN — normalized adjacency · GAT — learned edge attention · Readout — graph-level pooling · Your turn: TinyGCN from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
 = D̃^{-1/2} à D̃^{-1/2} from scratch and compute it on any toy graphH^(l+1) = σ( H^(l) W^(l)) with correct tensor shapese_{uv} from [Wh_u ‖ Wh_v], mask non-edges, softmax over N(v)∪{v} — weights sum to 1TinyGCN (48 params, 2 layers) to 100% node accuracy in 200 epochs| concept | the one thing to remember |
|---|---|
| GCN normalization | Always use Ã=A+I; entry = 1/sqrt(d̃_u · d̃_v) |
| GAT attention | softmax over N(v)∪{v} only — mask non-edges with -inf |
| message passing | L layers = L-hop receptive field |
| readout | mean/max/sum pool over node dim for graph tasks |
| GCN vs GAT | GAT wins when neighbor quality varies; GCN simpler and fast |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.