Lesson 94: Graph Neural Networks (GCN & GAT)

USAAIO Lesson 94, from Phase 3. It covers message passing on graphs and derives the normalized adjacency D^{-1/2} Ã D^{-1/2} from first principles on a 4-cycle, then works a GCN forward pass with real computed values. It covers Graph Attention Networks, with the attention weights verified to sum to 1, and global readout pooling by mean, max, or sum, ending with a TinyGCN of 48 parameters trained to 100% accuracy on a synthetic 8-node community graph. All the numbers were verified with torch 2.7.1+cpu. The lesson runs to 28 slides.

Subject: Machine Learning · 53 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Graph Neural Networks GCN & GAT from Scratch

Title

USAAIO · Lesson 94 · Phase 3

Message passing on graphs: normalize the adjacency, aggregate neighbor features, optionally weight edges with learned attention. Every formula derived and every number verified in PyTorch — from a 4-node cycle to a trained community-graph classifier.

2. By the end of this lesson you can

Objectives

  1. Write the message-passing update rule h_v^(l+1) = UPDATE(h_v^(l), AGGREGATE({m_{uv}})) and name what GCN substitutes for each piece
  2. Derive the GCN normalized adjacency D^{-1/2} Ã D^{-1/2} from scratch and compute it on a toy graph
  3. Trace a GCN forward pass (matrix form H^(l+1) = σ(Ã_norm H^(l) W^(l))) with real tensor shapes and values
  4. Explain how GAT replaces uniform 1/deg weights with learned α_{uv} and verify that they sum to 1 over each neighborhood
  5. Implement global readout (mean/max/sum pool) to produce a fixed-size graph-level vector from node embeddings

3. What survived from BERT Variants — XLNet, RoBERTa, T5, DeBERTa, ALBERT?

Warm-up

Discussion prompt

Before we open Lesson 94: Graph Neural Networks (GCN & GAT): without looking back, what was the main idea of BERT Variants — XLNet, RoBERTa, T5, DeBERTa, ALBERT, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

why bidirectional BERT is slow at inference and how each variant fixes a different limitation — XLNet permutation LM (two-stream attention, content vs query masks), RoBERTa (dynamic masking, no NSP, larger batches), T5 text-to-text unified seq2seq framing, DeBERTa disentangled content+position attention, and ALBERT cross-layer parameter sharing (6x reduction verified in PyTorch at d=64).

4. Message passing — the GNN abstraction

Section

Part 1 of 4

5. Graphs: nodes, edges, features

Concept

A graph G = (V, E) pairs a set of nodes V (each carrying a feature vector h_v) with edges E (which may carry features e_{uv}). Tasks: node classification (label each node), link prediction (predict edges), graph classification (one label per graph).

graphnodesedgesnode featurestypical task
social networkusersfriendshipsprofile text embeddingsnode classification
molecular graphatomsbondsatom type, chargegraph classification
citation networkpaperscitationsbag-of-words TF-IDFnode classification
knowledge graphentitiesrelationstype one-hotlink prediction

6. Which is which, by typical task

Discrimination

Sort into buckets

Sort these by typical task, from memory, without looking back at Graphs: nodes, edges, features. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

node classification
social network; citation network
graph classification
molecular graph
link prediction
knowledge graph
g1
typical task is "node classification" for social network, citation network — that is what the table on "Graphs: nodes, edges, features" records, and it is the single property separating this group from the rest.
g2
typical task is "graph classification" for molecular graph — that is what the table on "Graphs: nodes, edges, features" records, and it is the single property separating this group from the rest.
g3
typical task is "link prediction" for knowledge graph — that is what the table on "Graphs: nodes, edges, features" records, and it is the single property separating this group from the rest.

7. The message-passing framework

Concept

All major GNNs (GCN, GAT, GraphSAGE, GIN) are instances of the same message-passing template — only the choices of MESSAGE, AGGREGATE, and UPDATE differ.

\[ m_{uv}^{(l)} = \text{MESSAGE}(h_u^{(l)},\; h_v^{(l)},\; e_{uv}) \]

\[ h_v^{(l+1)} = \text{UPDATE}\!\left(h_v^{(l)},\; \text{AGGREGATE}\bigl(\{m_{uv}^{(l)} : u \in \mathcal{N}(v)\}\bigr)\right) \]

modelMESSAGEAGGREGATEUPDATE
GCNW h_u (linear)normalized sumσ(·)
GATW h_uattention-weighted sumσ(·)
GraphSAGEW h_umean/max/LSTMconcat + σ(·)
GINMLP(h_u)sum (injective)MLP(·)

8. Fill in: AGGREGATE for The message-passing framework

Comparison

Comparison matrix

From The message-passing framework: refill the AGGREGATE column from what you know. The rest of the table is as it appeared.

modelMESSAGEAGGREGATEUPDATE
GCNW h_u (linear)normalized sumσ(·)
GATW h_uattention-weighted sumσ(·)
GraphSAGEW h_umean/max/LSTMconcat + σ(·)
GINMLP(h_u)sum (injective)MLP(·)

9. GCN — normalized adjacency

Section

Part 2 of 4

10. Why normalize the adjacency?

Concept

Naively summing neighbor features gives high-degree nodes much larger activation magnitudes than low-degree nodes — a scale instability that makes training hard. GCN (Kipf & Welling 2017) fixes this with a symmetric normalization.

\[ \tilde{A} = A + I_N \quad\text{(add self-loops so each node aggregates itself)} \]

\[ \hat{A} = \tilde{D}^{-1/2}\, \tilde{A}\, \tilde{D}^{-1/2} \qquad \tilde{D}_{ii} = \sum_j \tilde{A}_{ij} \]

The (u,v) entry of  equals 1 / sqrt(d̃_u · d̃_v). Each edge contribution is scaled by the geometric mean of the degree-with-self-loop of both endpoints — so high-degree hubs are down-weighted and low-degree nodes are not dominated.

11. Guess the shape of the answer: Compute  on a 4-cycle

Estimation

Predict first

Graph: nodes 0-3, edges {(0,1),(1,2),(2,3),(0,3)} — a 4-cycle. Add self-loops to get Ã. Compute degree matrix D̃ and then  = D̃^{-1/2} à D̃^{-1/2}. Predict each diagonal entry before proceeding.

Commit before you compute: what does Compute  on a 4-cycle come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Â[i,i] = 1/sqrt(3·3) = 1/3 ≈ 0.3333; Â[i,j] (connected) = 1/sqrt(3·3) ≈ 0.3333 too

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Because all degrees are equal (regular graph), all non-zero entries of  are the same: 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/3.

12. Compute  on a 4-cycle

Worked example

Graph: nodes 0-3, edges {(0,1),(1,2),(2,3),(0,3)} — a 4-cycle. Add self-loops to get Ã. Compute degree matrix D̃ and then  = D̃^{-1/2} à D̃^{-1/2}. Predict each diagonal entry before proceeding.

import torch

A = torch.tensor([
    [0,1,0,1],
    [1,0,1,0],
    [0,1,0,1],
    [1,0,1,0]
], dtype=torch.float32)

A_tilde = A + torch.eye(4)         # add self-loops
d = A_tilde.sum(dim=1)             # [3, 3, 3, 3]
D_inv_sqrt = torch.diag(d ** -0.5) # 1/sqrt(3) on diagonal
A_norm = D_inv_sqrt @ A_tilde @ D_inv_sqrt
print('d:', d.tolist())
print('A_norm:\n', A_norm.round(decimals=4))

Each node has degree 2 (two cycle edges); adding self-loop makes d̃ = 3 for all nodes

Why: Ã row sums: node 0 has edges to 1, 3, and itself → 3 non-zeros → d̃_0 = 3. Symmetric graph so all degrees equal.

Â[i,i] = 1/sqrt(3·3) = 1/3 ≈ 0.3333; Â[i,j] (connected) = 1/sqrt(3·3) ≈ 0.3333 too

Why: Because all degrees are equal (regular graph), all non-zero entries of  are the same: 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/3.

entryformulavalue
d̃_0 = d̃_1 = d̃_2 = d̃_32 edges + self-loop3
D̃^{-1/2} diagonal1/sqrt(3)0.5774
Â[i,i] (self)1/sqrt(3·3) = 1/30.3333
Â[i,j] (connected)1/sqrt(3·3) = 1/30.3333
Â[i,j] (disconnected)0 (no edge)0.0000

13. What each one costs: Compute  on a 4-cycle

Trade off

Comparison matrix

From Compute  on a 4-cycle: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

entryformulavalue
d̃_0 = d̃_1 = d̃_2 = d̃_32 edges + self-loop3
D̃^{-1/2} diagonal1/sqrt(3)0.5774
Â[i,i] (self)1/sqrt(3·3) = 1/30.3333
Â[i,j] (connected)1/sqrt(3·3) = 1/30.3333
Â[i,j] (disconnected)0 (no edge)0.0000

14. Guess the shape of the answer: GCN forward pass: Â H W then ReLU

Estimation

Predict first

One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.

Commit before you compute: what does GCN forward pass: Â H W then ReLU come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For node 0, neighbors are {1, 3} plus self, so row 0 of A_norm@H = (h_0 + h_1 + h_3)/3.

15. GCN forward pass: Â H W then ReLU

Worked example

One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.

import torch, torch.nn.functional as F

torch.manual_seed(42)
H = torch.randn(4, 2)   # node features (4 nodes, 2 dims)
W = torch.randn(2, 3)   # weight matrix (project to 3 dims)

# A_norm from previous slide
A_norm = torch.full((4,4), 1/3) * torch.tensor([
    [1,1,0,1],[1,1,1,0],[0,1,1,1],[1,0,1,1]
], dtype=torch.float32)

H1 = F.relu(A_norm @ H @ W)
print('H shape:', H.shape, '  W shape:', W.shape)
print('A_norm @ H @ W before ReLU shape:', (A_norm @ H @ W).shape)
print('H^(1):\n', H1.round(decimals=4))

A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)

Why: For node 0, neighbors are {1, 3} plus self, so row 0 of A_norm@H = (h_0 + h_1 + h_3)/3. This is symmetric mean-of-neighbors aggregation weighted by degree.

nodeH^(1)[node, :]note
0[0.3525, 0.1445, 0.6526]neighbors: {0,1,3}; all pre-ReLU > 0
1[0.0000, 0.0148, 0.0000]neighbors: {0,1,2}; mostly clipped by ReLU
2[0.0428, 0.0000, 0.5699]neighbors: {1,2,3}
3[0.0312, 0.0000, 0.6453]neighbors: {0,2,3}

16. Work backwards from the answer: GCN forward pass: Â H W then ReLU

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

A_norm @ H: each row i of the output is (1/3)(h_i + sum of neighbor features)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

One GCN layer: H^(1) = ReLU( H^(0) W). Use the  from the 4-cycle, random H^(0) (4×2), and random W (2×3). Trace node 0's output vector.

17. Something is wrong here: using raw A instead of à (forgetting self-loops)

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use the raw adjacency A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself.

Always add self-loops first: Ã = A + I. Compute degree on Ã, then normalize: Â = D̃^{-1/2} Ã D̃^{-1/2}.

Why: Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.

18. Trap: using raw A instead of à (forgetting self-loops)

Trap

The trap

Use the raw adjacency A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.

A_norm = D_inv_sqrt @ A @ D_inv_sqrt — skipping A + I

Why: This is wrong. Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.

The fix

Always add self-loops first: Ã = A + I. Compute degree on Ã, then normalize: Â = D̃^{-1/2} Ã D̃^{-1/2}.

A_tilde = A + torch.eye(n); d = A_tilde.sum(1); A_norm = diag(d-0.5) @ A_tilde @ diag(d-0.5)

Why: With self-loops, row i of ·H includes h_i itself at weight 1/d̃_i. This ensures the node's own representation propagates forward — critical for learning node-level properties.

19. Break it on purpose: using raw A instead of à (forgetting…

Break the constraint

Discussion prompt

The rule this trap just fixed:

With self-loops, row i of ·H includes h_i itself at weight 1/d̃_i. This ensures the node's own representation propagates forward — critical for learning node-level properties.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Without self-loops, node v's own features h_v are NOT included in the aggregation — the new embedding h_v^{(l+1)} depends only on neighbors, not on v itself. The node forgets where it started.

20. GAT — learned edge attention

Section

Part 3 of 4

21. GAT: replace 1/d with learned α

Concept

GCN treats all neighbors equally: weight = 1/sqrt(d̃_u · d̃_v). GAT (Velickovic 2018) lets the model learn how much attention to pay each neighbor, conditioned on the features of both endpoints.

\[ e_{uv} = \text{LeakyReLU}\!\left(\mathbf{a}^\top [\mathbf{W}h_u \| \mathbf{W}h_v]\right) \]

\[ \alpha_{uv} = \text{softmax}_u(e_{uv}) = \frac{\exp(e_{uv})}{\sum_{k \in \mathcal{N}(v) \cup \{v\}} \exp(e_{kv})} \]

The softmax is over the neighborhood of v (including v itself). So Σ_{u ∈ N(v)∪{v}} α_{uv} = 1 for every node v — the attention weights are a proper probability distribution over neighbors.

22. Break it if you can: GAT: replace 1/d with learned α

Counterexample

Discussion prompt

GCN treats all neighbors equally: weight = 1/sqrt(d̃_u · d̃_v). GAT (Velickovic 2018) lets the model learn how much attention to pay each neighbor, conditioned on the features of both endpoints.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The softmax is over the neighborhood of v (including v itself). So Σ_{u ∈ N(v)∪{v}} α_{uv} = 1 for every node v — the attention weights are a proper probability distribution over neighbors.

23. Guess the shape of the answer: GAT attention weights on a toy neighborhood

Estimation

Predict first

Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.

Commit before you compute: what does GAT attention weights on a toy neighborhood come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each e_{0,j} = LeakyReLU(a^T [Wh_0 || Wh_j]).

24. GAT attention weights on a toy neighborhood

Worked example

Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.

import torch, torch.nn.functional as F

torch.manual_seed(42)
H_gat = torch.tensor([[0.1,0.2,0.3],[0.4,0.5,0.6],
                       [0.7,0.8,0.9],[0.2,0.3,0.1]], dtype=torch.float32)
# Wh = H @ W_gat  (3->3 linear, random W)
W_gat = torch.randn(3, 3)
a_vec = torch.randn(6)        # attention param: 2*d_out = 6
Wh = H_gat @ W_gat            # (4, 3)

def attn_score(hi, hj):
    return F.leaky_relu((a_vec @ torch.cat([hi, hj])).unsqueeze(0),
                        negative_slope=0.2).item()

# Node 0 neighbors: {0,1,3}
neighbors = [0, 1, 3]
raw = [attn_score(Wh[0], Wh[j]) for j in neighbors]
alpha = F.softmax(torch.tensor(raw), dim=0)
print('raw e_{0,j}:', [round(x,4) for x in raw])
print('alpha_{0,j}:', alpha.round(decimals=4).tolist())
print('sum:', alpha.sum().item())

Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771

Why: Each e_{0,j} = LeakyReLU(a^T [Wh_0 || Wh_j]). Neighbor 1 gets the highest raw score — the model discovers that node 1's features are most aligned with node 0's attention vector.

neighbor jraw score e_{0,j}α_{0,j} (softmax)interpretation
0 (self)1.46560.1507self-attention: 15% weight
13.06610.7470dominant neighbor: 74.7% weight
31.07710.1022weakest neighbor: 10.2% weight
sum—1.0000valid probability distribution

25. Work backwards from the answer: GAT attention weights on a toy neighborhood

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Raw scores: e_{0,0}=1.4656, e_{0,1}=3.0661, e_{0,3}=1.0771

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Node 0 has neighbors {0 (self), 1, 3} in the 4-cycle. Compute raw attention scores e_{0j} for each neighbor, then apply softmax. Verify the weights sum to 1.

26. Something is wrong here: computing softmax over ALL nodes, not just the…

Anomaly

Predict first

A student writes this, and it looks reasonable:

Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.

Mask non-neighbors with -inf before softmax so they receive exactly zero attention weight.

Why: Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.

27. Trap: computing softmax over ALL nodes, not just the neighborhood

Trap

The trap

Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).

alpha = F.softmax(e_all_nodes, dim=0) — no masking of non-neighbors

Why: This is wrong. Each node has only a few neighbors, but softmax over all N nodes will assign tiny but non-zero weights to non-connected nodes, leaking information across non-edges and violating the local aggregation principle of GNNs.

The fix

Mask non-neighbors with -inf before softmax so they receive exactly zero attention weight.

e_masked = e.masked_fill(adj_mask == 0, float('-inf')); alpha = F.softmax(e_masked, dim=1)

Why: softmax(-inf) = 0. Only neighbors (and self) have finite scores, so all weight concentrates on them. Row sums are 1 only over the valid neighborhood — verified by alpha.sum(dim=1) all equals 1.0 in Part 8 of the verification script.

28. Which of these survive contact with Lesson 94: Graph Neural Networks (GCN & GAT)?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
All major GNNs (GCN, GAT, GraphSAGE, GIN) are instances of the same message-passing template — only the choices of MESSAGE, AGGREGATE, and UPDATE differ.; Node classification uses H^(L)[v] directly. For graph classification, you need one fixed-size vector per graph — a readout (global pooling) over all node embeddings.; Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.
Breaks
Use the raw adjacency A (without self-loop) for normalization: Â = D^{-1/2} A D^{-1/2}.; Apply softmax over all N nodes to get α_{uv}: alpha = softmax(e[v, :]) where e has shape (N,).
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 94: Graph Neural Networks (GCN & GAT) puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

29. Readout — graph-level pooling

Section

Part 4 of 4

30. From node embeddings to a graph vector

Concept

Node classification uses H^(L)[v] directly. For graph classification, you need one fixed-size vector per graph — a readout (global pooling) over all node embeddings.

\[ h_G = \text{READOUT}\bigl(\{h_v^{(L)} : v \in V\}\bigr) \]

readoutformulapropertiesdifferentiable?
mean poolmean over rows of H^(L)permutation-invariant, boundedyes
max poolmax per dim over all rowsselects most activated feature per dimyes (subgradient)
sum poolsum over rows of H^(L)sensitive to graph sizeyes
attn poolsoftmax(gate)^T H^(L)learned soft gate over nodesyes

31. Fill in: differentiable? for From node embeddings to a graph vector

Comparison

Comparison matrix

From From node embeddings to a graph vector: refill the differentiable? column from what you know. The rest of the table is as it appeared.

readoutformulapropertiesdifferentiable?
mean poolmean over rows of H^(L)permutation-invariant, boundedyes
max poolmax per dim over all rowsselects most activated feature per dimyes (subgradient)
sum poolsum over rows of H^(L)sensitive to graph sizeyes
attn poolsoftmax(gate)^T H^(L)learned soft gate over nodesyes

32. Without one step: The GCN/GAT recipe

Constraint

Discussion prompt

Run The GCN/GAT recipe with this step confiscated:

Depth: 2-3 layers typical; L layers → L-hop neighborhood receptive field

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Prepare adjacency: Ã = A + I; degree d̃_i = Σ_j Ã_{ij}; Â = D̃^{-1/2} Ã D̃^{-1/2}
  2. GCN layer (matrix form): H^(l+1) = σ(Â H^(l) W^(l)) — one matrix multiply, no Python loop over nodes
  3. GAT layer: learn W (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges…
  4. Depth: 2-3 layers typical; L layers → L-hop neighborhood receptive field
  5. Readout (graph tasks): mean/max/sum over H^(L) → fixed-size graph vector → MLP classifier
  6. Node tasks: no readout; apply linear head directly to each row of H^(L)
  7. Key hyperparameter: attention heads K; multi-head output = concat (intermediate) or mean (final layer)

33. The GCN/GAT recipe

Pattern

  1. Prepare adjacency: Ã = A + I; degree d̃_i = Σ_j Ã_{ij}; Â = D̃^{-1/2} Ã D̃^{-1/2}
  2. GCN layer (matrix form): H^(l+1) = σ(Â H^(l) W^(l)) — one matrix multiply, no Python loop over nodes
  3. GAT layer: learn W (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges with -inf); aggregate h_v^{(l+1)} = σ(Σ_{u∈N(v)} α_{uv} W h_u)
  4. Depth: 2-3 layers typical; L layers → L-hop neighborhood receptive field
  5. Readout (graph tasks): mean/max/sum over H^(L) → fixed-size graph vector → MLP classifier
  6. Node tasks: no readout; apply linear head directly to each row of H^(L)
  7. Key hyperparameter: attention heads K; multi-head output = concat (intermediate) or mean (final layer)

34. Where does it stop working: The GCN/GAT recipe

Edge cases

Discussion prompt

The GCN/GAT recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Prepare adjacency: Ã = A + I; degree d̃_i = Σ_j Ã_{ij}; Â = D̃^{-1/2} Ã D̃^{-1/2}
  2. GCN layer (matrix form): H^(l+1) = σ(Â H^(l) W^(l)) — one matrix multiply, no Python loop over nodes
  3. GAT layer: learn W (transform) and a (attention); compute e_{uv} = LeakyReLU(a^T [Wh_u ‖ Wh_v]); apply softmax only over N(v) (mask non-edges…
  4. Depth: 2-3 layers typical; L layers → L-hop neighborhood receptive field
  5. Readout (graph tasks): mean/max/sum over H^(L) → fixed-size graph vector → MLP classifier
  6. Node tasks: no readout; apply linear head directly to each row of H^(L)
  7. Key hyperparameter: attention heads K; multi-head output = concat (intermediate) or mean (final layer)

35. Rule out three: Check yourself — normalized adjacency entry

Elimination

Eliminate the wrong options

In a graph where node u has degree 4 (after adding self-loop, d̃_u = 5) and node v has degree 2 (d̃_v = 3), and (u,v) is an edge, what is Â[u,v] in the GCN normalized adjacency?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1 / sqrt(5 × 3) ≈ 0.2582
  • B. 1 / (5 + 3) = 0.125
  • C. 2 / (5 + 3) = 0.25
  • D. 1 / sqrt(4 × 2) ≈ 0.3536

Survives elimination: A

Why: Â[u,v] = 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/sqrt(5×3) = 1/sqrt(15) ≈ 0.2582. The formula uses the degree after adding self-loops (d̃), not the original degree (d).

36. Check yourself — normalized adjacency entry

Check

Work out the formula before selecting.

Check your understanding

In a graph where node u has degree 4 (after adding self-loop, d̃_u = 5) and node v has degree 2 (d̃_v = 3), and (u,v) is an edge, what is Â[u,v] in the GCN normalized adjacency?

  • A. 1 / sqrt(5 × 3) ≈ 0.2582 (correct)
  • B. 1 / (5 + 3) = 0.125
  • C. 2 / (5 + 3) = 0.25
  • D. 1 / sqrt(4 × 2) ≈ 0.3536

Answer: A

Why: Â[u,v] = 1/(d̃_u^{1/2} · d̃_v^{1/2}) = 1/sqrt(5×3) = 1/sqrt(15) ≈ 0.2582. The formula uses the degree after adding self-loops (d̃), not the original degree (d).

Why B tempts people
1/(d̃_u + d̃_v) is the arithmetic-mean normalization, not the symmetric geometric-mean used by GCN. GCN uses the product under a square root.
Why C tempts people
2/(d̃_u + d̃_v) has no derivation from the Kipf & Welling formula; the factor of 2 appears nowhere in D̃^{-1/2} Ã D̃^{-1/2}.
Why D tempts people
Using the original degrees (4 and 2) instead of the self-loop-augmented degrees (5 and 3) is the 'forgot self-loop in denominator' mistake — the trap slide covers exactly this error.

37. Answer it before you see the options: Check yourself — GAT attention…

Prediction

Predict first

Node v has 3 neighbors (plus itself, so 4 total in N(v)∪{v}). After computing GAT attention weights α_{uv} for all u ∈ N(v)∪{v}, what must be true?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Σ_{u ∈ N(v)∪{v}} α_{uv} = 1

Why: The softmax in GAT is applied only over the neighborhood N(v)∪{v}, making the 4 weights a proper probability distribution that sums to 1. Non-neighbors receive exactly 0 weight (via -inf masking). This was verified in code: all row sums equal 1.0.

38. Check yourself — GAT attention constraint

Check

Think about the softmax domain.

Check your understanding

Node v has 3 neighbors (plus itself, so 4 total in N(v)∪{v}). After computing GAT attention weights α_{uv} for all u ∈ N(v)∪{v}, what must be true?

  • A. Σ_{u ∈ N(v)∪{v}} α_{uv} = 1 (correct)
  • B. Each α_{uv} ∈ [0, 1] and Σ_{u ∈ V} α_{uv} = 1 (sum over all N nodes)
  • C. Σ_{u ∈ N(v)∪{v}} α_{uv} = 1/4 (one contribution per neighbor)
  • D. Σ_{u ∈ N(v)∪{v}} α_{uv} = K where K is the number of attention heads

Answer: A

Why: The softmax in GAT is applied only over the neighborhood N(v)∪{v}, making the 4 weights a proper probability distribution that sums to 1. Non-neighbors receive exactly 0 weight (via -inf masking). This was verified in code: all row sums equal 1.0.

Why B tempts people
Summing over all N nodes would mean non-neighbors could receive positive weight, violating the local aggregation principle. The correct softmax domain is N(v)∪{v}, not all of V.
Why C tempts people
α_{uv} = 1/4 for each neighbor would be uniform weighting — identical to mean pooling. GAT's entire purpose is to learn non-uniform weights; a uniform constraint defeats that.
Why D tempts people
Multi-head GAT runs K independent attention functions and concatenates their outputs, but each head's weights still sum to 1 independently. K appears in the output dimension, not in the weight sum.

39. Rule out three: Check yourself — GCN vs GAT tradeoffs

Elimination

Eliminate the wrong options

You are building a citation-network node classifier. Each paper has a bag-of-words feature vector. Some papers cite far more relevant works than others. Which statement correctly describes when GAT has an advantage over GCN?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. GAT can up-weight the most informative neighboring papers and ignore noisy citations, while GCN averages all neighbors equally weighted by degree.
  • B. GAT is always better than GCN because attention adds parameters, and more parameters always improve accuracy.
  • C. GCN is better when neighbor features vary widely; GAT is better when all neighbors are equally informative.
  • D. GAT removes the need for the self-loop trick because attention automatically down-weights the node's own features.

Survives elimination: A

Why: GAT learns feature-conditioned attention weights α_{uv}, so it can assign high weight to topically relevant neighbors and low weight to off-topic citations — exactly what fixed degree-normalization cannot do. On Cora, GAT historically outperforms GCN by ~1–2% because the benefit of selective attention outweighs the added parameters.

40. Check yourself — GCN vs GAT tradeoffs

Check

Apply the inductive bias reasoning.

Check your understanding

You are building a citation-network node classifier. Each paper has a bag-of-words feature vector. Some papers cite far more relevant works than others. Which statement correctly describes when GAT has an advantage over GCN?

  • A. GAT can up-weight the most informative neighboring papers and ignore noisy citations, while GCN averages all neighbors equally weighted by degree. (correct)
  • B. GAT is always better than GCN because attention adds parameters, and more parameters always improve accuracy.
  • C. GCN is better when neighbor features vary widely; GAT is better when all neighbors are equally informative.
  • D. GAT removes the need for the self-loop trick because attention automatically down-weights the node's own features.

Answer: A

Why: GAT learns feature-conditioned attention weights α_{uv}, so it can assign high weight to topically relevant neighbors and low weight to off-topic citations — exactly what fixed degree-normalization cannot do. On Cora, GAT historically outperforms GCN by ~1–2% because the benefit of selective attention outweighs the added parameters.

Why B tempts people
More parameters do not automatically help — they can overfit. GAT's advantage is the inductive bias of edge-wise attention, not raw parameter count. On very small graphs GAT can overfit where GCN does not.
Why C tempts people
This is backwards. GCN is best when neighbors are roughly equally informative (uniform weights are near-optimal). GAT shines precisely when neighbor quality varies — it can selectively up-weight informative neighbors.
Why D tempts people
Both GCN and GAT typically use self-loops. GAT's attention weight for the self-edge (u=v) is computed the same way as for other edges and is learned, but the self-loop is still explicitly included in N(v)∪{v}.

41. Your turn: TinyGCN from scratch

Section

Project

42. Project: train TinyGCN on a community graph

Concept

Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.

#milestonekey check
1Build normalized adjacency Â; verify row sums ≤ 1 (irregular graph)A_norm.sum(1)
2Implement TinyGCN (2 layers, 48 params); trace shapes through each A_norm @ H @ Wprint(H1.shape, H2.shape)
3Train 200 epochs with Adam lr=0.01; hit 100% node accuracy on epoch≥50(logits.argmax(1)==y).float().mean()

Build rules: precompute  once before the training loop (it is fixed); use bias=False in both Linear layers; verify that 2-hop message passing (2 GCN layers) is sufficient for nodes to receive information from across the bridge.

43. Break it if you can: Project: train TinyGCN on a community graph

Counterexample

Discussion prompt

Build TinyGCN — two GCN layers (no bias, 4→8→2) on an 8-node synthetic graph with two 4-cliques connected by a bridge. Task: classify each node into community 0 or 1.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

44. Milestone 1 — build and verify Â

Worked example

Your turn: construct the 8-node two-clique adjacency. Predict: are all row sums of  the same? (Hint: what is the degree of the bridge node vs an interior clique node?)

Hint: d = A_tilde.sum(1) gives a vector (not a scalar). Interior nodes have degree 7 + self-loop = 8; the bridge node has degree 7 + 1 bridge edge + self-loop = 9. So  is not a uniform matrix unlike the 4-cycle example.

import torch

torch.manual_seed(0)
n = 8
A = torch.zeros(n, n)
# Two 4-cliques
for i in range(4):
    for j in range(4):
        if i != j: A[i,j] = 1
for i in range(4, 8):
    for j in range(4, 8):
        if i != j: A[i,j] = 1
A[3,4] = A[4,3] = 1          # bridge between node 3 and node 4

A_tilde = A + torch.eye(n)
d = A_tilde.sum(1)            # degree with self-loops
D_inv_sqrt = torch.diag(d**-0.5)
A_norm = D_inv_sqrt @ A_tilde @ D_inv_sqrt
print('degrees d_tilde:', d.tolist())
print('A_norm row sums:', A_norm.sum(1).round(decimals=4).tolist())
noded̃ (self-loop)role row sum
0, 1, 24interior clique-A1.0 (all edges normalized)
35bridge node (clique-A side)1.0
45bridge node (clique-B side)1.0
5, 6, 74interior clique-B1.0

45. Milestone 2 — TinyGCN forward pass

Worked example

Your turn: implement TinyGCN with layers 4→8→2. Predict the shape of the intermediate hidden matrix H1 and the final logits H2. How many total parameters?

Hint: nn.Linear(4, 8, bias=False) has 32 weights; nn.Linear(8, 2, bias=False) has 16 weights. Total = 48 parameters. The forward pass is: H1 = ReLU(Â @ W1(X)), then H2 = Â @ W2(H1) (no activation on the final layer — CrossEntropyLoss expects raw logits).

import torch, torch.nn as nn, torch.nn.functional as F

class TinyGCN(nn.Module):
    def __init__(self, in_f, h_f, out_f):
        super().__init__()
        self.W1 = nn.Linear(in_f, h_f, bias=False)
        self.W2 = nn.Linear(h_f, out_f, bias=False)
    def forward(self, X, A_norm):
        H1 = F.relu(A_norm @ self.W1(X))   # (n, h_f)
        H2 = A_norm @ self.W2(H1)           # (n, out_f)
        return H2

torch.manual_seed(42)
X = torch.randn(8, 4)             # 8 nodes, 4 features each
y = torch.tensor([0,0,0,0,1,1,1,1], dtype=torch.long)
model = TinyGCN(4, 8, 2)
logits = model(X, A_norm)         # A_norm from milestone 1
print('params:', sum(p.numel() for p in model.parameters()))
print('H2 shape:', logits.shape)
layerinput shapeoutput shapeparams
W1 (Linear 4→8)(8, 4)(8, 8)32
ReLU + A_norm@(8, 8)(8, 8)0
W2 (Linear 8→2)(8, 8)(8, 2)16
A_norm@ (readout)(8, 2)(8, 2)0
TOTAL——48

46. Which is which, by output shape

Discrimination

Sort into buckets

Sort these by output shape, from memory, without looking back at Milestone 2 — TinyGCN forward pass. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(8, 8)
W1 (Linear 4→8); ReLU + A_norm@
(8, 2)
W2 (Linear 8→2); A_norm@ (readout)
—
TOTAL
g1
output shape is "(8, 8)" for W1 (Linear 4→8), ReLU + A_norm@ — that is what the table on "Milestone 2 — TinyGCN forward pass" records, and it is the single property separating this group from the rest.
g2
output shape is "(8, 2)" for W2 (Linear 8→2), A_norm@ (readout) — that is what the table on "Milestone 2 — TinyGCN forward pass" records, and it is the single property separating this group from the rest.
g3
output shape is "—" for TOTAL — that is what the table on "Milestone 2 — TinyGCN forward pass" records, and it is the single property separating this group from the rest.

47. Milestone 3 — train to 100% accuracy

Worked example

Your turn: run the training loop for 200 epochs. Predict: does the model converge faster or slower than a vanilla MLP on the same features? (Recall (Lesson 40): 5-step loop is zero_grad → forward → loss → backward → step.)

Hint: precompute A_norm outside the loop — it is a fixed graph property, not a model parameter. Use Adam lr=0.01 (higher than transformers, typical for GNNs on small graphs).

import torch, torch.nn as nn, torch.nn.functional as F

torch.manual_seed(42)
X = torch.randn(8, 4)
y = torch.tensor([0,0,0,0,1,1,1,1], dtype=torch.long)
model = TinyGCN(4, 8, 2)        # defined in milestone 2
opt = torch.optim.Adam(model.parameters(), lr=0.01)
loss_fn = nn.CrossEntropyLoss()

for epoch in range(201):
    opt.zero_grad()
    logits = model(X, A_norm)   # A_norm from milestone 1
    loss = loss_fn(logits, y)
    loss.backward()
    opt.step()
    if epoch % 50 == 0:
        acc = (logits.argmax(1)==y).float().mean()
        print(f'ep {epoch:3d}  loss={loss.item():.4f}  acc={acc*100:.0f}%')
epochlossaccuracynote
00.76400%random init, predictions all wrong
500.2027100%graph structure learned by epoch 50
1000.0644100%loss still decreasing, accuracy saturated
1500.0280100%well separated logits
2000.0134100%converged — 48 params sufficient

48. What each one costs: Milestone 3 — train to 100% accuracy

Trade off

Comparison matrix

From Milestone 3 — train to 100% accuracy: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.

epochlossaccuracynote
00.76400%random init, predictions all wrong
500.2027100%graph structure learned by epoch 50
1000.0644100%loss still decreasing, accuracy saturated
1500.0280100%well separated logits
2000.0134100%converged — 48 params sufficient

49. The full program

Concept

import torch, torch.nn as nn, torch.nn.functional as F

torch.manual_seed(42)

# --- Graph: two 4-cliques + bridge ---
n = 8
A = torch.zeros(n, n)
for i in range(4):
    for j in range(4):
        if i!=j: A[i,j]=1
for i in range(4,8):
    for j in range(4,8):
        if i!=j: A[i,j]=1
A[3,4]=A[4,3]=1
A_t = A+torch.eye(n); d=A_t.sum(1)
A_norm = torch.diag(d**-0.5)@A_t@torch.diag(d**-0.5)

# --- Features & labels ---
X = torch.randn(n,4)
y = torch.tensor([0,0,0,0,1,1,1,1],dtype=torch.long)

# --- Model ---
class TinyGCN(nn.Module):
    def __init__(self,i,h,o):
        super().__init__()
        self.W1=nn.Linear(i,h,bias=False)
        self.W2=nn.Linear(h,o,bias=False)
    def forward(self,X,A):
        return A@self.W2(F.relu(A@self.W1(X)))

model=TinyGCN(4,8,2)
opt=torch.optim.Adam(model.parameters(),lr=0.01)
for _ in range(200):
    opt.zero_grad()
    loss=F.cross_entropy(model(X,A_norm),y)
    loss.backward(); opt.step()
with torch.no_grad():
    acc=(model(X,A_norm).argmax(1)==y).float().mean()
print(f'{acc.item()*100:.0f}% node accuracy')  # 100%
print('params:', sum(p.numel() for p in model.parameters()))  # 48
design choicevaluerationale
layers2 GCN2-hop receptive field crosses the bridge
hidden dim8small enough to train in <1 s on CPU
biasFalsesimpler; normalization handles scale
lr0.01GNNs on small graphs tolerate higher lr than transformers
epochs200100% accuracy reached by epoch 50
total params484×8 + 8×2 = 32+16

50. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

design choicevaluerationale
layers2 GCN2-hop receptive field crosses the bridge
hidden dim8small enough to train in <1 s on CPU
biasFalsesimpler; normalization handles scale
lr0.01GNNs on small graphs tolerate higher lr than transformers
epochs200100% accuracy reached by epoch 50
total params484×8 + 8×2 = 32+16

51. Show it off

Concept

Out loud, slides closed: (1) derive the GCN normalized adjacency formula from scratch — explain each step and why you need both +I and both D^{-1/2} factors; (2) trace the attention weight computation for one node in GAT, stating why softmax is restricted to the neighborhood; (3) explain when you would choose mean-pool vs max-pool readout for graph classification.

Stretch (lesson-plan homework): implement SingleHeadGAT from scratch — learn W and a, mask non-edges with -inf, apply softmax per row, aggregate. Verify alpha.sum(1) is all 1.0. Then compare TinyGCN vs a linear MLP (same feature matrix, no adjacency) on the community graph — does removing graph structure hurt accuracy?

52. Connect it up: Lesson 94: Graph Neural Networks (GCN & GAT)

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Message passing — the GNN abstraction · GCN — normalized adjacency · GAT — learned edge attention · Readout — graph-level pooling · Your turn: TinyGCN from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

53. What you can do now

Recap

conceptthe one thing to remember
GCN normalizationAlways use Ã=A+I; entry = 1/sqrt(d̃_u · d̃_v)
GAT attentionsoftmax over N(v)∪{v} only — mask non-edges with -inf
message passingL layers = L-hop receptive field
readoutmean/max/sum pool over node dim for graph tasks
GCN vs GATGAT wins when neighbor quality varies; GCN simpler and fast

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 94 — Graph Neural Networks — Barron · USAAIO Round 2 Preparation, 2026
  2. Kipf & Welling 'Semi-Supervised Classification with Graph Convolutional Networks' (ICLR 2017) — arXiv:1609.02907
  3. Velickovic et al. 'Graph Attention Networks' (ICLR 2018) — arXiv:1710.10903
  4. TinyGCN, normalized adjacency, GAT attention verification with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108