USAAIO Lesson 96, from Phase 3, on graph transformers. It replaces positional encoding with structural encoding, using degree centrality and shortest-path distance, then covers Graphormer's spatial bias and edge encoding, the VNode used for a global graph readout, full all-pairs attention against the sparsity of a GNN, and the applications to molecular property prediction and knowledge graphs. A four-node toy molecule was verified with torch 2.7.1+cpu, covering the shortest-path-distance matrix, the spatial bias matrix, the biased attention weights, masked attention in the sparse variant, and the edge-encoding projection. The lesson runs to 27 slides.
Subject: Machine Learning · 54 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 96 · Phase 3
Standard transformers assume 1-D sequences. Graphs have arbitrary topology. Graph Transformers replace positional encoding with structural encoding — shortest-path distance, degree, edge features — so the same all-pairs attention mechanism generalises to molecules, social networks, and knowledge graphs.
Objectives
[CLS] in BERT (Lesson 88)Section
Part 1 of 4
Concept
In a standard transformer (Lesson 83–86), every token attends to every other token: attention is all-pairs. That is exactly a complete directed graph — nodes are tokens, edges are attention connections.
Graph Neural Networks (GCN, GAT) only pass messages along existing edges — they miss long-range interactions. Graph Transformers use full attention and compensate with structure-aware biases.
Matching
Match the pairs
From A transformer is a fully-connected graph — match each one to what it actually does. The descriptions have been shuffled.
Why: Sequence input, Graph input, Key insight are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Concept
Fix a toy molecule: C(0)–N(1)–H(3) and C(0)–O(2). Adjacency gives degrees and a shortest-path distance (SPD) matrix that will drive all structural encodings.
| Node | Atom | Degree | Neighbours |
|---|---|---|---|
| 0 | C | 2 | N(1), O(2) |
| 1 | N | 2 | C(0), H(3) |
| 2 | O | 1 | C(0) |
| 3 | H | 1 | N(1) |
\[ \text{SPD} = \begin{pmatrix}0&1&1&2\\1&0&2&1\\1&2&0&3\\2&1&3&0\end{pmatrix} \]
Comparison
Comparison matrix
From Toy molecule: 4-node graph: refill the Atom column from what you know. The rest of the table is as it appeared.
| Node | Atom | Degree | Neighbours |
|---|---|---|---|
| 0 | C | 2 | N(1), O(2) |
| 1 | N | 2 | C(0), H(3) |
| 2 | O | 1 | C(0) |
| 3 | H | 1 | N(1) |
Section
Part 2 of 4
Concept
Graphormer replaces the sinusoidal positional embedding with a centrality encoding — a learned embedding indexed by in-degree and out-degree, added to the initial node feature vector.
\[ h_i^{(0)} = x_i + z_{\deg^-(i)} + z_{\deg^+(i)} \]
For undirected graphs deg⁻ = deg⁺ = deg. Node C (deg=2) and node O (deg=1) get different initial representations even if their atom-type features are identical — degree is a coarse measure of importance in the graph.
| Node | deg | Centrality emb index |
|---|---|---|
| C (0) | 2 | z[2] |
| N (1) | 2 | z[2] |
| O (2) | 1 | z[1] |
| H (3) | 1 | z[1] |
Discrimination
Sort into buckets
Sort these by deg, from memory, without looking back at Centrality encoding: degree as node identity. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
For each pair (i, j), Graphormer adds a learned scalar bias φ(SPD[i,j]) to the attention logit before the softmax. Closer nodes (SPD=1) get a weaker penalty than distant nodes (SPD=3).
\[ A_{ij} = \frac{Q_i K_j^\top}{\sqrt{d_k}} + \varphi(\text{SPD}[i,j]) \]
| SPD | phi(SPD) — toy learned value |
|---|---|
| 0 (self) | 0.00 |
| 1 (direct edge) | -0.50 |
| 2 (two hops) | -1.20 |
| 3 (three hops) | -2.10 |
| ≥4 (clipped) | -3.00 |
Unlike positional encoding's absolute index, φ encodes graph distance — it is the same value regardless of node ordering in the sequence.
Trade off
Comparison matrix
From Spatial encoding: SPD → attention bias: every row here is a choice with a cost. Fill the phi(SPD) — toy learned value column, then say which row you would actually pick and what you give up for it.
| SPD | phi(SPD) — toy learned value |
|---|---|
| 0 (self) | 0.00 |
| 1 (direct edge) | -0.50 |
| 2 (two hops) | -1.20 |
| 3 (three hops) | -2.10 |
| ≥4 (clipped) | -3.00 |
Ranking
Put in order
Put the moves of Worked example: biased attention for node 0 (C) into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each node i is a 4-dim one-hot (atom type), linearly projected to d_model=8 then split into d_k=4 head queries/keys; scale by 1/√4 = 0.5.
Worked example
Embed node features and compute raw Q·Kᵀ/√d_k
Why: Each node i is a 4-dim one-hot (atom type), linearly projected to d_model=8 then split into d_k=4 head queries/keys; scale by 1/√4 = 0.5.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
N, d_model, d_k = 4, 8, 4
X = torch.eye(N) # one-hot atom features
proj = nn.Linear(N, d_model, bias=False)
X_emb = proj(X) # (4, 8)
torch.manual_seed(1)
Wq = nn.Linear(d_model, d_k, bias=False)
Wk = nn.Linear(d_model, d_k, bias=False)
Wv = nn.Linear(d_model, d_k, bias=False)
for m in [Wq, Wk, Wv]: nn.init.normal_(m.weight, std=0.3)
Q, K, V = Wq(X_emb), Wk(X_emb), Wv(X_emb)
scores = Q @ K.T / (d_k**0.5) # (4, 4)
print(scores[0].detach().numpy().round(4))| Output: scores[0] | j=0 (C) | j=1 (N) | j=2 (O) | j=3 (H) |
|---|---|---|---|---|
| raw Q·Kᵀ/√dk | 0.0649 | 0.0766 | -0.0943 | -0.0441 |
Add spatial bias B[0,j] = phi(SPD[0,j])
Why: SPD[0,:] = [0,1,1,2] → phi = [0.00, -0.50, -0.50, -1.20]. Bias penalises distant nodes, pushing attention mass toward closer neighbours.
Softmax over biased logits to get attention weights for node 0
Why: softmax([0.065, -0.423, -0.594, -1.244]) = [0.4165, 0.2556, 0.2154, 0.1125]. Node 0 (self) dominates; distant H(3) receives least weight.
| j (atom) | SPD(0,j) | phi | biased logit | attn weight |
|---|---|---|---|---|
| 0 (C) | 0 | 0.00 | 0.0649 | 0.4165 |
| 1 (N) | 1 | -0.50 | -0.4234 | 0.2556 |
| 2 (O) | 1 | -0.50 | -0.5943 | 0.2154 |
| 3 (H) | 2 | -1.20 | -1.2441 | 0.1125 |
Error analysis
Annotate
Walk the callouts on Worked example: biased attention for node 0 (C). Each one is a place this is easy to get subtly wrong.
Concept
SPD captures distance but not bond type (single/double/aromatic). Graphormer embeds each edge's feature vector and mean-pools along the shortest path between i and j, then adds the result as an additional attention bias.
\[ b_{ij}^{\text{edge}} = \frac{1}{|P_{ij}|} \sum_{e \in P_{ij}} W_c \, a_e \]
Full bias: b[i,j] = phi(SPD[i,j]) + b_edge[i,j]. For the toy edge (C→N), a single-bond one-hot [1,0,0] projects to d_k=4 via Wc: [0.1719, -0.3128, -0.3979, -0.2849] (verified).
| Path | Bond type | a_e (3-dim) | W_c·a_e (4-dim, verified) |
|---|---|---|---|
| C–N (0→1) | single | [1, 0, 0] | [0.1719, -0.3128, -0.3979, -0.2849] |
| C–O (0→2) | single | [1, 0, 0] | [0.1719, -0.3128, -0.3979, -0.2849] |
| C–N=O (0→1→?) | multi-hop | mean pooled | — |
Counterexample
Discussion prompt
Full bias: b[i,j] = phi(SPD[i,j]) + b_edge[i,j]. For the toy edge (C→N), a single-bond one-hot [1,0,0] projects to d_k=4 via Wc: [0.1719, -0.3128, -0.3979, -0.2849] (verified).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Wrong approach: only use phi(SPD[i,j]) as the bias. Skip edge encoding because the distance already tells you how far apart nodes are.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).
Correct approach: combine spatial encoding with edge encoding for the full structural bias.
Why: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).
Trap
Wrong approach: only use phi(SPD[i,j]) as the bias. Skip edge encoding because the distance already tells you how far apart nodes are.
Set b[i,j] = phi(SPD[i,j]) only
Why: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).
Correct approach: combine spatial encoding with edge encoding for the full structural bias.
Set b[i,j] = phi(SPD[i,j]) + mean_path_edge_enc(i,j)
Why: Edge encoding encodes the type of bonds along the path; SPD encodes distance. Together they give the attention head complete structural context — critical for molecular property prediction.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Edge encoding encodes the type of bonds along the path; SPD encodes distance. Together they give the attention head complete structural context — critical for molecular property prediction.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).
Section
Part 3 of 4
Concept
An alternative to the soft bias (Graphormer) is a hard mask: set attention logits for non-adjacent pairs to -∞ so only direct neighbours contribute. This recovers a transformer-style GNN.
\[ A_{ij} = \begin{cases} \frac{Q_i K_j^\top}{\sqrt{d_k}} & \text{if } (i,j) \in \mathcal{E} \\ -\infty & \text{otherwise} \end{cases} \]
| Variant | Node 0 sees | Node 2 sees | Long-range? |
|---|---|---|---|
| Graphormer (soft) | all 4 nodes (biased) | all 4 nodes (biased) | Yes |
| Graph Transformer (hard mask) | N(1), O(2) only | C(0) only | No (1-hop only) |
Comparison
Comparison matrix
From Sparse variant: mask non-edge pairs: refill the Node 0 sees column from what you know. The rest of the table is as it appeared.
| Variant | Node 0 sees | Node 2 sees | Long-range? |
|---|---|---|---|
| Graphormer (soft) | all 4 nodes (biased) | all 4 nodes (biased) | Yes |
| Graph Transformer (hard mask) | N(1), O(2) only | C(0) only | No (1-hop only) |
Pattern
Predict first
The table runs: 0(C) — masked | 0.0000 | 0.5426 | 0.4574 | 0.0000 · 1(N) — masked | 0.5072 | 0.0000 | 0.0000 | 0.4928 · 2(O) — masked | 1.0000 | 0.0000 | 0.0000 | 0.0000
In Worked example: masked attention (hard mask), given the rows so far: what is the next one — the row where Node is 3(H) — masked?
Correct: 3(H) — masked | 0.0000 | 1.0000 | 0.0000 | 0.0000
| Node | attn→0(C) | attn→1(N) | attn→2(O) | attn→3(H) |
|---|---|---|---|---|
| 0(C) — masked | 0.0000 | 0.5426 | 0.4574 | 0.0000 |
| 1(N) — masked | 0.5072 | 0.0000 | 0.0000 | 0.4928 |
| 2(O) — masked | 1.0000 | 0.0000 | 0.0000 | 0.0000 |
| 3(H) — masked | 0.0000 | 1.0000 | 0.0000 | 0.0000 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Adding -1e9 before softmax makes exp(-1e9) ≈ 0, effectively zeroing non-neighbour attention after softmax.
Worked example
Build adjacency mask M: M[i,j] = 0 if edge exists, -1e9 if not
Why: Adding -1e9 before softmax makes exp(-1e9) ≈ 0, effectively zeroing non-neighbour attention after softmax.
import torch, torch.nn.functional as F
# Adjacency (same toy graph)
A = torch.zeros(4, 4)
for i, j in [(0,1),(1,0),(0,2),(2,0),(1,3),(3,1)]:
A[i,j] = 1.0
# scores from same Q,K as before (torch.manual_seed(0,1))
scores = torch.tensor([
[ 0.0649, 0.0766, -0.0943, -0.0441],
[ 0.0088, 0.0033, -0.0073, -0.0191],
[ 0.0383, 0.0432, -0.0224, -0.1069],
[ 0.0011, 0.0173, 0.0362, -0.0858]
])
mask = (A == 0).float() * (-1e9)
masked_attn = F.softmax(scores + mask, dim=-1)
print(masked_attn.numpy().round(4))| Node | attn→0(C) | attn→1(N) | attn→2(O) | attn→3(H) |
|---|---|---|---|---|
| 0(C) — masked | 0.0000 | 0.5426 | 0.4574 | 0.0000 |
| 1(N) — masked | 0.5072 | 0.0000 | 0.0000 | 0.4928 |
| 2(O) — masked | 1.0000 | 0.0000 | 0.0000 | 0.0000 |
| 3(H) — masked | 0.0000 | 1.0000 | 0.0000 | 0.0000 |
Compare: node 2 (O) attends only to C(0), its sole neighbour
Why: After masking, O has one non-zero attention weight = 1.0 on C. Soft Graphormer gives O: 0.31 on C, 0.15 on N, 0.48 on O(self), 0.05 on H — the self-loop dominates in soft mode.
Cost model
Annotate
In Worked example: masked attention (hard mask), before reading the notes: mark where the time actually goes. Which line dominates?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Wrong: hard masking still lets distant nodes influence each other — information just propagates through intermediate nodes over multiple layers.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.
Correct: hard-mask Graph Transformer is equivalent to a GNN in its information-propagation radius per layer.
Why: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.
Trap
Wrong: hard masking still lets distant nodes influence each other — information just propagates through intermediate nodes over multiple layers.
Claim: stacking L masked-attention layers captures paths of length L
Why: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.
Correct: hard-mask Graph Transformer is equivalent to a GNN in its information-propagation radius per layer.
Use soft-bias (Graphormer) for true one-layer global context
Why: With phi(SPD[i,j]) as the bias, node O(2) and node H(3) interact in a single attention layer even though SPD=3. Graphormer achieves global aggregation without depth — critical for molecular graphs where bond order 4+ effects matter.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Concept
Graphormer prepends a Virtual Node (VNode) to the node sequence — analogous to [CLS] in BERT (Lesson 88). VNode is connected to every real node with a fixed SPD = 1, so phi(1) = -0.50 applies uniformly.
\[ h_{\text{VNode}} = \text{Transformer}\bigl([v; h_0, h_1, \ldots, h_{N-1}]\bigr)[0] \]
| Token | SPD to VNode | phi | Role |
|---|---|---|---|
| VNode (self) | 0 | 0.00 | Global representation |
| C (0) | 1 | -0.50 | All real nodes seen equally |
| N (1) | 1 | -0.50 | All real nodes seen equally |
| O (2) | 1 | -0.50 | All real nodes seen equally |
| H (3) | 1 | -0.50 | All real nodes seen equally |
After the encoder, VNode's hidden state is read out for graph-level tasks (e.g. predicting molecular energy). Node-level tasks use the individual node representations.
Trade off
Comparison matrix
From VNode: the global graph token: every row here is a choice with a cost. Fill the SPD to VNode column, then say which row you would actually pick and what you give up for it.
| Token | SPD to VNode | phi | Role |
|---|---|---|---|
| VNode (self) | 0 | 0.00 | Global representation |
| C (0) | 1 | -0.50 | All real nodes seen equally |
| N (1) | 1 | -0.50 | All real nodes seen equally |
| O (2) | 1 | -0.50 | All real nodes seen equally |
| H (3) | 1 | -0.50 | All real nodes seen equally |
Section
Part 4 of 4
Concept
Graph Transformers use full pairwise attention: O(N² · d) per layer. Standard GNNs (GCN, GAT) use sparse message passing: O(|E| · d) per layer. The tradeoff depends on graph density.
| Method | Complexity/layer | Long-range | Best for |
|---|---|---|---|
| GCN/GIN | O(|E|·d) | No (k layers = k hops) | Large sparse graphs |
| GAT | O(|E|·d·H) | No | Sparse + learned edge weights |
| Graph Transformer (mask) | O(|E|·d) | No (same as GNN) | Sparse, transformer inductive bias |
| Graphormer (soft bias) | O(N²·d) | Yes (1 layer) | Small dense graphs, molecules |
For N=50 nodes, d=128: Graphormer costs 50²×128 = 320,000 ops/layer vs a chain graph's 50×128 = 6,400 for a GCN — ~50× more expensive but captures global interactions in one layer.
Discrimination
Sort into buckets
Sort these by Complexity/layer, from memory, without looking back at O(N²) attention vs sparse GNN. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Key signal: if your task requires correlating features of nodes that are multiple hops apart and your graph is small enough for O(N²) attention, a Graph Transformer is worth the cost.
Counterexample
Discussion prompt
Key signal: if your task requires correlating features of nodes that are multiple hops apart and your graph is small enough for O(N²) attention, a Graph Transformer is worth the cost.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Constraint
Discussion prompt
Run Pattern: building a Graphormer attention layer with this step confiscated:
Edge encoding: for each pair (i,j) embed edges on shortest path; mean-pool → b_edge[i,j]
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
x_i (atom type, etc.); adjacency A; edge features a_ez[deg] to each x_i → h_i⁰Q = h W_Q, K = h W_K, V = h W_V for each headphi[i,j] = phi_table[SPD[i,j]]b_edge[i,j]A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weightsh_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)h_iPattern
x_i (atom type, etc.); adjacency A; edge features a_ez[deg] to each x_i → h_i⁰Q = h W_Q, K = h W_K, V = h W_V for each headphi[i,j] = phi_table[SPD[i,j]]b_edge[i,j]A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weightsh_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)h_iEdge cases
Discussion prompt
Pattern: building a Graphormer attention layer works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
x_i (atom type, etc.); adjacency A; edge features a_ez[deg] to each x_i → h_i⁰Q = h W_Q, K = h W_K, V = h W_V for each headphi[i,j] = phi_table[SPD[i,j]]b_edge[i,j]A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weightsh_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)h_iElimination
Eliminate the wrong options
In the toy graph, node O(2) to H(3) has SPD=3 and phi=-2.10; O to C(0) has SPD=1 and phi=-0.50. Assuming equal raw attention scores, which statement is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: B
Why: phi(1)=-0.50 is a weaker penalty than phi(3)=-2.10. With equal raw scores, the biased logit for (O→C) is 1.70 higher than for (O→H), so softmax assigns more weight to C. phi encodes graph proximity, not strength of bonding.
Check
In the 4-node toy graph, node O(2) has SPD of 3 to node H(3). The spatial bias table assigns phi(3) = -2.10. How does this affect O→H attention relative to O→C (SPD=1, phi=-0.50)?
Check your understanding
In the toy graph, node O(2) to H(3) has SPD=3 and phi=-2.10; O to C(0) has SPD=1 and phi=-0.50. Assuming equal raw attention scores, which statement is correct?
Answer: B
Why: phi(1)=-0.50 is a weaker penalty than phi(3)=-2.10. With equal raw scores, the biased logit for (O→C) is 1.70 higher than for (O→H), so softmax assigns more weight to C. phi encodes graph proximity, not strength of bonding.
Prediction
Predict first
Which of the following correctly describes a design choice in Graphormer?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The VNode is connected to all real nodes with a fixed SPD=1, so phi(1) is added to every VNode–node attention logit.
Why: Graphormer assigns SPD(VNode, i)=1 for all real nodes i. This means phi(1) biases the VNode's attention toward all nodes uniformly, giving it a balanced global view. The VNode hidden state after the encoder is used as the graph-level representation.
Check
Identify the correct statement about Graphormer's VNode and centrality encoding.
Check your understanding
Which of the following correctly describes a design choice in Graphormer?
Answer: C
Why: Graphormer assigns SPD(VNode, i)=1 for all real nodes i. This means phi(1) biases the VNode's attention toward all nodes uniformly, giving it a balanced global view. The VNode hidden state after the encoder is used as the graph-level representation.
Elimination
Eliminate the wrong options
For N=50, |E|=120, d=128, how many times more expensive (approximately) is one Graphormer layer vs one GCN layer in terms of multiply-accumulate operations?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: B
Why: Graphormer: O(N²·d) = 50²×128 = 320,000. GCN: O(|E|·d) = 120×128 = 15,360. Ratio = 320000/15360 ≈ 20.8×, closest to option B's ~26× (the approximation difference comes from ignoring constant factors). The key point: Graphormer's cost scales quadratically in N, GCN's scales with edge count.
Check
A molecule has N=50 atoms, d=128, and 60 bonds (|E|=120 counting both directions). Compare one Graphormer attention layer to one GCN message-passing layer.
Check your understanding
For N=50, |E|=120, d=128, how many times more expensive (approximately) is one Graphormer layer vs one GCN layer in terms of multiply-accumulate operations?
Answer: B
Why: Graphormer: O(N²·d) = 50²×128 = 320,000. GCN: O(|E|·d) = 120×128 = 15,360. Ratio = 320000/15360 ≈ 20.8×, closest to option B's ~26× (the approximation difference comes from ignoring constant factors). The key point: Graphormer's cost scales quadratically in N, GCN's scales with edge count.
Pattern
Predict first
The table runs: C (0) | val[0] | val[1] | val[2] | val[3] · Shape | (4, 4) | — | — | —
In Full Graphormer mini-module in PyTorch, given the rows so far: what is the next one — the row where Node is Module params?
Correct: Module params | 3×(8×4)+5×1 | = | 101 | params
| Node | out[0] | out[1] | out[2] | out[3] |
|---|---|---|---|---|
| C (0) | val[0] | val[1] | val[2] | val[3] |
| Shape | (4, 4) | — | — | — |
| Module params | 3×(8×4)+5×1 | = | 101 | params |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The SPD matrix is a precomputed graph property (BFS); we convert it to bias via a learned embedding table phi_table of size max_dist+1.
Worked example
Build GraphormerAttention: takes node embeddings X and precomputed SPD matrix, returns updated node representations
Why: The SPD matrix is a precomputed graph property (BFS); we convert it to bias via a learned embedding table phi_table of size max_dist+1.
import torch, torch.nn as nn, torch.nn.functional as F
class GraphormerAttention(nn.Module):
def __init__(self, d_model=8, d_k=4, max_dist=4):
super().__init__()
self.d_k = d_k
self.Wq = nn.Linear(d_model, d_k, bias=False)
self.Wk = nn.Linear(d_model, d_k, bias=False)
self.Wv = nn.Linear(d_model, d_k, bias=False)
# phi_table[d] = learned scalar bias for SPD=d
self.phi = nn.Embedding(max_dist + 1, 1)
def forward(self, X, spd):
# X: (N, d_model) spd: (N, N) int tensor
Q, K, V = self.Wq(X), self.Wk(X), self.Wv(X)
scores = Q @ K.T / (self.d_k ** 0.5) # (N, N)
bias = self.phi(spd.clamp(0, 4)).squeeze(-1) # (N, N)
attn = F.softmax(scores + bias, dim=-1)
return attn @ V # (N, d_k)
# Test on toy molecule
torch.manual_seed(42)
N, d = 4, 8
X = torch.eye(N)
spd = torch.tensor([[0,1,1,2],[1,0,2,1],[1,2,0,3],[2,1,3,0]])
model = GraphormerAttention()
out = model(X, spd)
print(out.shape, out.detach().numpy().round(4))| Node | out[0] | out[1] | out[2] | out[3] |
|---|---|---|---|---|
| C (0) | val[0] | val[1] | val[2] | val[3] |
| Shape | (4, 4) | — | — | — |
| Module params | 3×(8×4)+5×1 | = | 101 | params |
Verify output shape: (N=4, d_k=4) and parameter count: 3×(8×4) + 5 = 101
Why: 3 linear layers of 8→4 = 3×32=96 weights; phi Embedding with 5 entries × 1 scalar = 5 weights. Total = 101. Output shape (4,4) confirmed by print.
Comparison
Comparison matrix
From Full Graphormer mini-module in PyTorch: refill the out[2] column from what you know. The rest of the table is as it appeared.
| Node | out[0] | out[1] | out[2] | out[3] |
|---|---|---|---|---|
| C (0) | val[0] | val[1] | val[2] | val[3] |
| Shape | (4, 4) | — | — | — |
| Module params | 3×(8×4)+5×1 | = | 101 | params |
Concept
Ranking
Put in order
Put the moves of Your turn — Graphormer attention from scratch into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. For a path 0–1–2–3–4, SPD[i,j] = |i−j|.
Worked example
Project brief: implement Graphormer single-head attention on a 5-node path graph (0–1–2–3–4). Compute the SPD matrix by hand, build the spatial bias, run biased attention, and verify that node 0 attends most to its neighbour node 1 and least to node 4 (SPD=4).
Milestone 1 — compute SPD for a path graph of 5 nodes
Why: For a path 0–1–2–3–4, SPD[i,j] = |i−j|. Write the 5×5 SPD matrix and confirm the diagonal is 0 and the corners SPD[0,4]=SPD[4,0]=4.
Milestone 2 — build the phi table and spatial bias matrix B
Why: Use phi_weights = [0.0, -0.5, -1.2, -2.1, -3.0] (same as lesson). Construct a 5×5 tensor B where B[i,j]=phi(min(|i-j|,4)). Check B[0,4]=B[4,0]=-3.0.
Milestone 3 — run biased attention for node 0; print attn_weights[0] and verify attn[0,1] > attn[0,4]
Why: Node 0 has SPD=1 to node 1 and SPD=4 to node 4. phi(1)=-0.50 vs phi(4)=-3.00: a 2.50 penalty differential guarantees more weight on node 1 (barring extreme raw score differences).
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
N, d_model, d_k = 5, 8, 4
# 5-node path graph: SPD[i,j] = |i-j|
spd = torch.tensor([[abs(i-j) for j in range(N)] for i in range(N)])
phi_w = torch.tensor([0.0, -0.5, -1.2, -2.1, -3.0])
B = phi_w[spd.clamp(0,4)] # (5,5) spatial bias
X = torch.eye(N) # one-hot node features
proj = nn.Linear(N, d_model, bias=False)
X_emb = proj(X)
Wq = nn.Linear(d_model, d_k, bias=False)
Wk = nn.Linear(d_model, d_k, bias=False)
Wv = nn.Linear(d_model, d_k, bias=False)
for m in [Wq,Wk,Wv]: nn.init.normal_(m.weight, std=0.3)
Q, K, V = Wq(X_emb), Wk(X_emb), Wv(X_emb)
scores = Q @ K.T / d_k**0.5
attn = F.softmax(scores + B, dim=-1)
print('attn[0]:', attn[0].detach().numpy().round(4))
print('node1 > node4:', attn[0,1].item() > attn[0,4].item())| attn[0,j] | j=0 (self) | j=1 | j=2 | j=3 | j=4 |
|---|---|---|---|---|---|
| Expected order | highest or 2nd | high (SPD=1) | medium | low | lowest (SPD=4) |
| phi bias | 0.00 | -0.50 | -1.20 | -2.10 | -3.00 |
Show it off: print the full 5×5 attention weight matrix and verify the diagonal dominance pattern along rows
Why: Each row i should show decreasing attention as SPD increases from 0 outward. This is the graph-aware inductive bias: closer nodes, more attention — without any positional integer index.
Error analysis
Annotate
Walk the callouts on Your turn — Graphormer attention from scratch. Each one is a place this is easy to get subtly wrong.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why graphs? Transformers already are graphs · Graphormer's three structural encodings · Soft bias vs hard mask — two Graph Transformer flavours · Complexity, scaling, applications. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.