Lesson 96: Graph Transformer & Graphormer

USAAIO Lesson 96, from Phase 3, on graph transformers. It replaces positional encoding with structural encoding, using degree centrality and shortest-path distance, then covers Graphormer's spatial bias and edge encoding, the VNode used for a global graph readout, full all-pairs attention against the sparsity of a GNN, and the applications to molecular property prediction and knowledge graphs. A four-node toy molecule was verified with torch 2.7.1+cpu, covering the shortest-path-distance matrix, the spatial bias matrix, the biased attention weights, masked attention in the sparse variant, and the edge-encoding projection. The lesson runs to 27 slides.

Subject: Machine Learning · 54 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Graph Transformer & Graphormer

Title

USAAIO · Lesson 96 · Phase 3

Standard transformers assume 1-D sequences. Graphs have arbitrary topology. Graph Transformers replace positional encoding with structural encoding — shortest-path distance, degree, edge features — so the same all-pairs attention mechanism generalises to molecules, social networks, and knowledge graphs.

2. By the end of this lesson you can

Objectives

  1. Explain why a standard transformer already defines a fully-connected graph and what changes when the input is a graph rather than a sequence
  2. Compute the shortest-path distance (SPD) matrix for a small graph and convert it to a spatial bias added to attention logits
  3. Describe Graphormer's three structural encodings: centrality, spatial, and edge encoding, and state what each replaces in a vanilla transformer
  4. Distinguish the soft-bias (Graphormer) and hard-mask (Graph Transformer) approaches to injecting graph structure
  5. Identify the VNode trick and explain why it acts as a global readout, analogous to [CLS] in BERT (Lesson 88)

3. Why graphs? Transformers already are graphs

Section

Part 1 of 4

4. A transformer is a fully-connected graph

Concept

In a standard transformer (Lesson 83–86), every token attends to every other token: attention is all-pairs. That is exactly a complete directed graph — nodes are tokens, edges are attention connections.

Sequence input
Tokens are ordered 1-D; position encodes the 1-D index (sinusoidal or learned).
Graph input
Nodes have arbitrary topology; structure encodes shortest-path distance, degree, and bond type.
Key insight
Replace positional encoding with structural encoding — the attention computation stays identical.

Graph Neural Networks (GCN, GAT) only pass messages along existing edges — they miss long-range interactions. Graph Transformers use full attention and compensate with structure-aware biases.

5. Which is which: A transformer is a fully-connected graph

Matching

Match the pairs

From A transformer is a fully-connected graph — match each one to what it actually does. The descriptions have been shuffled.

  • c1. Sequence input
  • c2. Graph input
  • c3. Key insight
  • b1. Tokens are ordered 1-D; position encodes the 1-D index (sinusoidal or learned).
  • b2. Nodes have arbitrary topology; structure encodes shortest-path distance, degree, and bond type.
  • b3. Replace positional encoding with structural encoding — the attention computation stays identical.

Why: Sequence input, Graph input, Key insight are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

6. Toy molecule: 4-node graph

Concept

Fix a toy molecule: C(0)–N(1)–H(3) and C(0)–O(2). Adjacency gives degrees and a shortest-path distance (SPD) matrix that will drive all structural encodings.

NodeAtomDegreeNeighbours
0C2N(1), O(2)
1N2C(0), H(3)
2O1C(0)
3H1N(1)

\[ \text{SPD} = \begin{pmatrix}0&1&1&2\\1&0&2&1\\1&2&0&3\\2&1&3&0\end{pmatrix} \]

7. Fill in: Atom for Toy molecule: 4-node graph

Comparison

Comparison matrix

From Toy molecule: 4-node graph: refill the Atom column from what you know. The rest of the table is as it appeared.

NodeAtomDegreeNeighbours
0C2N(1), O(2)
1N2C(0), H(3)
2O1C(0)
3H1N(1)

8. Graphormer's three structural encodings

Section

Part 2 of 4

9. Centrality encoding: degree as node identity

Concept

Graphormer replaces the sinusoidal positional embedding with a centrality encoding — a learned embedding indexed by in-degree and out-degree, added to the initial node feature vector.

\[ h_i^{(0)} = x_i + z_{\deg^-(i)} + z_{\deg^+(i)} \]

For undirected graphs deg⁻ = deg⁺ = deg. Node C (deg=2) and node O (deg=1) get different initial representations even if their atom-type features are identical — degree is a coarse measure of importance in the graph.

NodedegCentrality emb index
C (0)2z[2]
N (1)2z[2]
O (2)1z[1]
H (3)1z[1]

10. Which is which, by deg

Discrimination

Sort into buckets

Sort these by deg, from memory, without looking back at Centrality encoding: degree as node identity. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

2
C (0); N (1)
1
O (2); H (3)
g1
deg is "2" for C (0), N (1) — that is what the table on "Centrality encoding: degree as node…" records, and it is the single property separating this group from the rest.
g2
deg is "1" for O (2), H (3) — that is what the table on "Centrality encoding: degree as node…" records, and it is the single property separating this group from the rest.

11. Spatial encoding: SPD → attention bias

Concept

For each pair (i, j), Graphormer adds a learned scalar bias φ(SPD[i,j]) to the attention logit before the softmax. Closer nodes (SPD=1) get a weaker penalty than distant nodes (SPD=3).

\[ A_{ij} = \frac{Q_i K_j^\top}{\sqrt{d_k}} + \varphi(\text{SPD}[i,j]) \]

SPDphi(SPD) — toy learned value
0 (self)0.00
1 (direct edge)-0.50
2 (two hops)-1.20
3 (three hops)-2.10
≥4 (clipped)-3.00

Unlike positional encoding's absolute index, φ encodes graph distance — it is the same value regardless of node ordering in the sequence.

12. What each one costs: Spatial encoding: SPD → attention bias

Trade off

Comparison matrix

From Spatial encoding: SPD → attention bias: every row here is a choice with a cost. Fill the phi(SPD) — toy learned value column, then say which row you would actually pick and what you give up for it.

SPDphi(SPD) — toy learned value
0 (self)0.00
1 (direct edge)-0.50
2 (two hops)-1.20
3 (three hops)-2.10
≥4 (clipped)-3.00

13. What has to happen first: Worked example: biased attention for node 0 (C)

Ranking

Put in order

Put the moves of Worked example: biased attention for node 0 (C) into the order they have to happen.

  1. Embed node features and compute raw Q·Kᵀ/√d_k
  2. Add spatial bias B[0,j] = phi(SPD[0,j])
  3. Softmax over biased logits to get attention weights for node 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Each node i is a 4-dim one-hot (atom type), linearly projected to d_model=8 then split into d_k=4 head queries/keys; scale by 1/√4 = 0.5.

14. Worked example: biased attention for node 0 (C)

Worked example

Embed node features and compute raw Q·Kᵀ/√d_k

Why: Each node i is a 4-dim one-hot (atom type), linearly projected to d_model=8 then split into d_k=4 head queries/keys; scale by 1/√4 = 0.5.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
N, d_model, d_k = 4, 8, 4
X = torch.eye(N)                          # one-hot atom features
proj = nn.Linear(N, d_model, bias=False)
X_emb = proj(X)                           # (4, 8)
torch.manual_seed(1)
Wq = nn.Linear(d_model, d_k, bias=False)
Wk = nn.Linear(d_model, d_k, bias=False)
Wv = nn.Linear(d_model, d_k, bias=False)
for m in [Wq, Wk, Wv]: nn.init.normal_(m.weight, std=0.3)
Q, K, V = Wq(X_emb), Wk(X_emb), Wv(X_emb)
scores = Q @ K.T / (d_k**0.5)            # (4, 4)
print(scores[0].detach().numpy().round(4))
Output: scores[0]j=0 (C)j=1 (N)j=2 (O)j=3 (H)
raw Q·Kᵀ/√dk0.06490.0766-0.0943-0.0441

Add spatial bias B[0,j] = phi(SPD[0,j])

Why: SPD[0,:] = [0,1,1,2] → phi = [0.00, -0.50, -0.50, -1.20]. Bias penalises distant nodes, pushing attention mass toward closer neighbours.

Softmax over biased logits to get attention weights for node 0

Why: softmax([0.065, -0.423, -0.594, -1.244]) = [0.4165, 0.2556, 0.2154, 0.1125]. Node 0 (self) dominates; distant H(3) receives least weight.

j (atom)SPD(0,j)phibiased logitattn weight
0 (C)00.000.06490.4165
1 (N)1-0.50-0.42340.2556
2 (O)1-0.50-0.59430.2154
3 (H)2-1.20-1.24410.1125

15. Inspect it line by line: Worked example: biased attention for node 0…

Error analysis

Annotate

Walk the callouts on Worked example: biased attention for node 0 (C). Each one is a place this is easy to get subtly wrong.

  • Each node i is a 4-dim one-hot (atom type), linearly projected to d_model=8 then split into d_k=4 head queries/keys; scale by 1/√4 = 0.5.
  • SPD[0,:] = [0,1,1,2] → phi = [0.00, -0.50, -0.50, -1.20]. Bias penalises distant nodes, pushing attention mass toward closer neighbours.
  • softmax([0.065, -0.423, -0.594, -1.244]) = [0.4165, 0.2556, 0.2154, 0.1125]. Node 0 (self) dominates; distant H(3) receives least weight.

16. Edge encoding: bond type along shortest path

Concept

SPD captures distance but not bond type (single/double/aromatic). Graphormer embeds each edge's feature vector and mean-pools along the shortest path between i and j, then adds the result as an additional attention bias.

\[ b_{ij}^{\text{edge}} = \frac{1}{|P_{ij}|} \sum_{e \in P_{ij}} W_c \, a_e \]

Full bias: b[i,j] = phi(SPD[i,j]) + b_edge[i,j]. For the toy edge (C→N), a single-bond one-hot [1,0,0] projects to d_k=4 via Wc: [0.1719, -0.3128, -0.3979, -0.2849] (verified).

PathBond typea_e (3-dim)W_c·a_e (4-dim, verified)
C–N (0→1)single[1, 0, 0][0.1719, -0.3128, -0.3979, -0.2849]
C–O (0→2)single[1, 0, 0][0.1719, -0.3128, -0.3979, -0.2849]
C–N=O (0→1→?)multi-hopmean pooled—

17. Break it if you can: Edge encoding: bond type along shortest path

Counterexample

Discussion prompt

Full bias: b[i,j] = phi(SPD[i,j]) + b_edge[i,j]. For the toy edge (C→N), a single-bond one-hot [1,0,0] projects to d_k=4 via Wc: [0.1719, -0.3128, -0.3979, -0.2849] (verified).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

18. Something is wrong here: phi is the entire structural signal

Anomaly

Predict first

A student writes this, and it looks reasonable:

Wrong approach: only use phi(SPD[i,j]) as the bias. Skip edge encoding because the distance already tells you how far apart nodes are.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).

Correct approach: combine spatial encoding with edge encoding for the full structural bias.

Why: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).

19. Trap: phi is the entire structural signal

Trap

The trap

Wrong approach: only use phi(SPD[i,j]) as the bias. Skip edge encoding because the distance already tells you how far apart nodes are.

Set b[i,j] = phi(SPD[i,j]) only

Why: Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).

The fix

Correct approach: combine spatial encoding with edge encoding for the full structural bias.

Set b[i,j] = phi(SPD[i,j]) + mean_path_edge_enc(i,j)

Why: Edge encoding encodes the type of bonds along the path; SPD encodes distance. Together they give the attention head complete structural context — critical for molecular property prediction.

20. Break it on purpose: phi is the entire structural signal

Break the constraint

Discussion prompt

The rule this trap just fixed:

Edge encoding encodes the type of bonds along the path; SPD encodes distance. Together they give the attention head complete structural context — critical for molecular property prediction.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Two pairs with SPD=1 get identical bias regardless of bond type — C=O (double bond) is indistinguishable from C–O (single bond).

21. Soft bias vs hard mask — two Graph Transformer flavours

Section

Part 3 of 4

22. Sparse variant: mask non-edge pairs

Concept

An alternative to the soft bias (Graphormer) is a hard mask: set attention logits for non-adjacent pairs to -∞ so only direct neighbours contribute. This recovers a transformer-style GNN.

\[ A_{ij} = \begin{cases} \frac{Q_i K_j^\top}{\sqrt{d_k}} & \text{if } (i,j) \in \mathcal{E} \\ -\infty & \text{otherwise} \end{cases} \]

VariantNode 0 seesNode 2 seesLong-range?
Graphormer (soft)all 4 nodes (biased)all 4 nodes (biased)Yes
Graph Transformer (hard mask)N(1), O(2) onlyC(0) onlyNo (1-hop only)

23. Fill in: Node 0 sees for Sparse variant: mask non-edge pairs

Comparison

Comparison matrix

From Sparse variant: mask non-edge pairs: refill the Node 0 sees column from what you know. The rest of the table is as it appeared.

VariantNode 0 seesNode 2 seesLong-range?
Graphormer (soft)all 4 nodes (biased)all 4 nodes (biased)Yes
Graph Transformer (hard mask)N(1), O(2) onlyC(0) onlyNo (1-hop only)

24. Predict the next row: Worked example: masked attention (hard mask)

Pattern

Predict first

The table runs: 0(C) — masked | 0.0000 | 0.5426 | 0.4574 | 0.0000 · 1(N) — masked | 0.5072 | 0.0000 | 0.0000 | 0.4928 · 2(O) — masked | 1.0000 | 0.0000 | 0.0000 | 0.0000

In Worked example: masked attention (hard mask), given the rows so far: what is the next one — the row where Node is 3(H) — masked?

Correct: 3(H) — masked | 0.0000 | 1.0000 | 0.0000 | 0.0000

Nodeattn→0(C)attn→1(N)attn→2(O)attn→3(H)
0(C) — masked0.00000.54260.45740.0000
1(N) — masked0.50720.00000.00000.4928
2(O) — masked1.00000.00000.00000.0000
3(H) — masked0.00001.00000.00000.0000

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Adding -1e9 before softmax makes exp(-1e9) ≈ 0, effectively zeroing non-neighbour attention after softmax.

25. Worked example: masked attention (hard mask)

Worked example

Build adjacency mask M: M[i,j] = 0 if edge exists, -1e9 if not

Why: Adding -1e9 before softmax makes exp(-1e9) ≈ 0, effectively zeroing non-neighbour attention after softmax.

import torch, torch.nn.functional as F
# Adjacency (same toy graph)
A = torch.zeros(4, 4)
for i, j in [(0,1),(1,0),(0,2),(2,0),(1,3),(3,1)]:
    A[i,j] = 1.0
# scores from same Q,K as before (torch.manual_seed(0,1))
scores = torch.tensor([
    [ 0.0649,  0.0766, -0.0943, -0.0441],
    [ 0.0088,  0.0033, -0.0073, -0.0191],
    [ 0.0383,  0.0432, -0.0224, -0.1069],
    [ 0.0011,  0.0173,  0.0362, -0.0858]
])
mask = (A == 0).float() * (-1e9)
masked_attn = F.softmax(scores + mask, dim=-1)
print(masked_attn.numpy().round(4))
Nodeattn→0(C)attn→1(N)attn→2(O)attn→3(H)
0(C) — masked0.00000.54260.45740.0000
1(N) — masked0.50720.00000.00000.4928
2(O) — masked1.00000.00000.00000.0000
3(H) — masked0.00001.00000.00000.0000

Compare: node 2 (O) attends only to C(0), its sole neighbour

Why: After masking, O has one non-zero attention weight = 1.0 on C. Soft Graphormer gives O: 0.31 on C, 0.15 on N, 0.48 on O(self), 0.05 on H — the self-loop dominates in soft mode.

26. Where the cost goes: Worked example: masked attention (hard mask)

Cost model

Annotate

In Worked example: masked attention (hard mask), before reading the notes: mark where the time actually goes. Which line dominates?

  • Adding -1e9 before softmax makes exp(-1e9) ≈ 0, effectively zeroing non-neighbour attention after softmax.
  • After masking, O has one non-zero attention weight = 1.0 on C. Soft Graphormer gives O: 0.31 on C, 0.15 on N, 0.48 on O(self), 0.05 on H — the self-loop dominates in soft mode.

27. Something is wrong here: hard mask captures long-range dependencies

Anomaly

Predict first

A student writes this, and it looks reasonable:

Wrong: hard masking still lets distant nodes influence each other — information just propagates through intermediate nodes over multiple layers.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.

Correct: hard-mask Graph Transformer is equivalent to a GNN in its information-propagation radius per layer.

Why: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.

28. Trap: hard mask captures long-range dependencies

Trap

The trap

Wrong: hard masking still lets distant nodes influence each other — information just propagates through intermediate nodes over multiple layers.

Claim: stacking L masked-attention layers captures paths of length L

Why: True in principle, but each layer only attends to 1-hop neighbours, so information from L hops away takes L layers — same bottleneck as GNNs, not global.

The fix

Correct: hard-mask Graph Transformer is equivalent to a GNN in its information-propagation radius per layer.

Use soft-bias (Graphormer) for true one-layer global context

Why: With phi(SPD[i,j]) as the bias, node O(2) and node H(3) interact in a single attention layer even though SPD=3. Graphormer achieves global aggregation without depth — critical for molecular graphs where bond order 4+ effects matter.

29. Which of these survive contact with Lesson 96: Graph Transformer & Graphormer?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Fix a toy molecule: C(0)–N(1)–H(3) and C(0)–O(2). Adjacency gives degrees and a shortest-path distance (SPD) matrix that will drive all structural encodings.; Unlike positional encoding's absolute index, φ encodes graph distance — it is the same value regardless of node ordering in the sequence.; After the encoder, VNode's hidden state is read out for graph-level tasks (e.g. predicting molecular energy). Node-level tasks use the individual node representations.
Breaks
Wrong approach: only use phi(SPD[i,j]) as the bias. Skip edge encoding because the distance already tells you how far apart nodes are.; Wrong: hard masking still lets distant nodes influence each other — information just propagates through intermediate nodes over multiple layers.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 96: Graph Transformer & Graphormer puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

30. VNode: the global graph token

Concept

Graphormer prepends a Virtual Node (VNode) to the node sequence — analogous to [CLS] in BERT (Lesson 88). VNode is connected to every real node with a fixed SPD = 1, so phi(1) = -0.50 applies uniformly.

\[ h_{\text{VNode}} = \text{Transformer}\bigl([v; h_0, h_1, \ldots, h_{N-1}]\bigr)[0] \]

TokenSPD to VNodephiRole
VNode (self)00.00Global representation
C (0)1-0.50All real nodes seen equally
N (1)1-0.50All real nodes seen equally
O (2)1-0.50All real nodes seen equally
H (3)1-0.50All real nodes seen equally

After the encoder, VNode's hidden state is read out for graph-level tasks (e.g. predicting molecular energy). Node-level tasks use the individual node representations.

31. What each one costs: VNode: the global graph token

Trade off

Comparison matrix

From VNode: the global graph token: every row here is a choice with a cost. Fill the SPD to VNode column, then say which row you would actually pick and what you give up for it.

TokenSPD to VNodephiRole
VNode (self)00.00Global representation
C (0)1-0.50All real nodes seen equally
N (1)1-0.50All real nodes seen equally
O (2)1-0.50All real nodes seen equally
H (3)1-0.50All real nodes seen equally

32. Complexity, scaling, applications

Section

Part 4 of 4

33. O(N²) attention vs sparse GNN

Concept

Graph Transformers use full pairwise attention: O(N² · d) per layer. Standard GNNs (GCN, GAT) use sparse message passing: O(|E| · d) per layer. The tradeoff depends on graph density.

MethodComplexity/layerLong-rangeBest for
GCN/GINO(|E|·d)No (k layers = k hops)Large sparse graphs
GATO(|E|·d·H)NoSparse + learned edge weights
Graph Transformer (mask)O(|E|·d)No (same as GNN)Sparse, transformer inductive bias
Graphormer (soft bias)O(N²·d)Yes (1 layer)Small dense graphs, molecules

For N=50 nodes, d=128: Graphormer costs 50²×128 = 320,000 ops/layer vs a chain graph's 50×128 = 6,400 for a GCN — ~50× more expensive but captures global interactions in one layer.

34. Which is which, by Complexity/layer

Discrimination

Sort into buckets

Sort these by Complexity/layer, from memory, without looking back at O(N²) attention vs sparse GNN. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

O(|E|·d)
GCN/GIN; Graph Transformer (mask)
O(|E|·d·H)
GAT
O(N²·d)
Graphormer (soft bias)
g1
Complexity/layer is "O(|E|·d)" for GCN/GIN, Graph Transformer (mask) — that is what the table on "O(N²) attention vs sparse GNN" records, and it is the single property separating this group from the rest.
g2
Complexity/layer is "O(|E|·d·H)" for GAT — that is what the table on "O(N²) attention vs sparse GNN" records, and it is the single property separating this group from the rest.
g3
Complexity/layer is "O(N²·d)" for Graphormer (soft bias) — that is what the table on "O(N²) attention vs sparse GNN" records, and it is the single property separating this group from the rest.

35. Applications: where Graph Transformers win

Concept

Key signal: if your task requires correlating features of nodes that are multiple hops apart and your graph is small enough for O(N²) attention, a Graph Transformer is worth the cost.

36. Break it if you can: Applications: where Graph Transformers win

Counterexample

Discussion prompt

Key signal: if your task requires correlating features of nodes that are multiple hops apart and your graph is small enough for O(N²) attention, a Graph Transformer is worth the cost.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

37. Without one step: Pattern: building a Graphormer attention layer

Constraint

Discussion prompt

Run Pattern: building a Graphormer attention layer with this step confiscated:

Edge encoding: for each pair (i,j) embed edges on shortest path; mean-pool → b_edge[i,j]

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Input: node features x_i (atom type, etc.); adjacency A; edge features a_e
  2. Centrality encoding: compute in/out degree; add learned embedding z[deg] to each x_i → h_i⁰
  3. Linear projections: Q = h W_Q, K = h W_K, V = h W_V for each head
  4. Spatial encoding: BFS all pairs → SPD matrix → scalar bias phi[i,j] = phi_table[SPD[i,j]]
  5. Edge encoding: for each pair (i,j) embed edges on shortest path; mean-pool → b_edge[i,j]
  6. Biased attention: A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weights
  7. Aggregation: h_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)
  8. Readout: for graph-level task, use VNode hidden state; for node-level, use individual h_i

38. Pattern: building a Graphormer attention layer

Pattern

  1. Input: node features x_i (atom type, etc.); adjacency A; edge features a_e
  2. Centrality encoding: compute in/out degree; add learned embedding z[deg] to each x_i → h_i⁰
  3. Linear projections: Q = h W_Q, K = h W_K, V = h W_V for each head
  4. Spatial encoding: BFS all pairs → SPD matrix → scalar bias phi[i,j] = phi_table[SPD[i,j]]
  5. Edge encoding: for each pair (i,j) embed edges on shortest path; mean-pool → b_edge[i,j]
  6. Biased attention: A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weights
  7. Aggregation: h_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)
  8. Readout: for graph-level task, use VNode hidden state; for node-level, use individual h_i

39. Where does it stop working: Pattern: building a Graphormer attention…

Edge cases

Discussion prompt

Pattern: building a Graphormer attention layer works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Input: node features x_i (atom type, etc.); adjacency A; edge features a_e
  2. Centrality encoding: compute in/out degree; add learned embedding z[deg] to each x_i → h_i⁰
  3. Linear projections: Q = h W_Q, K = h W_K, V = h W_V for each head
  4. Spatial encoding: BFS all pairs → SPD matrix → scalar bias phi[i,j] = phi_table[SPD[i,j]]
  5. Edge encoding: for each pair (i,j) embed edges on shortest path; mean-pool → b_edge[i,j]
  6. Biased attention: A[i,j] = (Q_i · K_j) / sqrt(d_k) + phi[i,j] + b_edge[i,j]; softmax over j → weights
  7. Aggregation: h_i_new = softmax(A[i,:]) · V; apply MLP + residual + LayerNorm (Lesson 84)
  8. Readout: for graph-level task, use VNode hidden state; for node-level, use individual h_i

40. Rule out three: Check 1: spatial encoding

Elimination

Eliminate the wrong options

In the toy graph, node O(2) to H(3) has SPD=3 and phi=-2.10; O to C(0) has SPD=1 and phi=-0.50. Assuming equal raw attention scores, which statement is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. O attends to H with higher weight than to C, because a larger |phi| value means stronger attention.
  • B. O attends to C with higher weight, because phi=-0.50 is a smaller penalty than phi=-2.10, leaving a higher biased logit.
  • C. O attends to H and C equally, because phi values cancel out after softmax normalization.
  • D. O only attends to its 1-hop neighbour C; the mask makes H receive zero weight.

Survives elimination: B

Why: phi(1)=-0.50 is a weaker penalty than phi(3)=-2.10. With equal raw scores, the biased logit for (O→C) is 1.70 higher than for (O→H), so softmax assigns more weight to C. phi encodes graph proximity, not strength of bonding.

41. Check 1: spatial encoding

Check

In the 4-node toy graph, node O(2) has SPD of 3 to node H(3). The spatial bias table assigns phi(3) = -2.10. How does this affect O→H attention relative to O→C (SPD=1, phi=-0.50)?

Check your understanding

In the toy graph, node O(2) to H(3) has SPD=3 and phi=-2.10; O to C(0) has SPD=1 and phi=-0.50. Assuming equal raw attention scores, which statement is correct?

  • A. O attends to H with higher weight than to C, because a larger |phi| value means stronger attention.
  • B. O attends to C with higher weight, because phi=-0.50 is a smaller penalty than phi=-2.10, leaving a higher biased logit. (correct)
  • C. O attends to H and C equally, because phi values cancel out after softmax normalization.
  • D. O only attends to its 1-hop neighbour C; the mask makes H receive zero weight.

Answer: B

Why: phi(1)=-0.50 is a weaker penalty than phi(3)=-2.10. With equal raw scores, the biased logit for (O→C) is 1.70 higher than for (O→H), so softmax assigns more weight to C. phi encodes graph proximity, not strength of bonding.

Why A tempts people
Larger |phi| means a more negative additive bias, which lowers the logit and therefore lowers attention weight — not raises it.
Why C tempts people
Softmax is order-preserving on its inputs: different input logits produce different output weights. Only identical inputs cancel.
Why D tempts people
D describes the hard-mask variant (Graph Transformer), not Graphormer's soft-bias approach. In Graphormer, non-adjacent pairs get a bias penalty, not -∞.

42. Answer it before you see the options: Check 2: VNode and centrality encoding

Prediction

Predict first

Which of the following correctly describes a design choice in Graphormer?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The VNode is connected to all real nodes with a fixed SPD=1, so phi(1) is added to every VNode–node attention logit.

Why: Graphormer assigns SPD(VNode, i)=1 for all real nodes i. This means phi(1) biases the VNode's attention toward all nodes uniformly, giving it a balanced global view. The VNode hidden state after the encoder is used as the graph-level representation.

43. Check 2: VNode and centrality encoding

Check

Identify the correct statement about Graphormer's VNode and centrality encoding.

Check your understanding

Which of the following correctly describes a design choice in Graphormer?

  • A. The VNode is connected to all real nodes with SPD=0 so it acts as a self-loop.
  • B. Centrality encoding uses eigenvector centrality because it captures global importance better than degree.
  • C. The VNode is connected to all real nodes with a fixed SPD=1, so phi(1) is added to every VNode–node attention logit. (correct)
  • D. Centrality encoding is added to the attention logits, not to the node feature vectors.

Answer: C

Why: Graphormer assigns SPD(VNode, i)=1 for all real nodes i. This means phi(1) biases the VNode's attention toward all nodes uniformly, giving it a balanced global view. The VNode hidden state after the encoder is used as the graph-level representation.

Why A tempts people
SPD=0 would mean the VNode treats itself as the same node — that conflicts with the shortest-path semantics and would make phi(0)=0 remove all bias for VNode interactions.
Why B tempts people
Graphormer uses in-degree and out-degree (simple degree centrality), not eigenvector centrality — degree is cheaper to compute and worked well in practice on molecular benchmarks.
Why D tempts people
Centrality encoding is added to the input node representation h_i⁰, not to attention logits. Spatial encoding is what modifies attention logits.

44. Rule out three: Check 3: Graphormer vs GNN complexity

Elimination

Eliminate the wrong options

For N=50, |E|=120, d=128, how many times more expensive (approximately) is one Graphormer layer vs one GCN layer in terms of multiply-accumulate operations?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. About 2× — both are roughly linear in N.
  • B. About 26× — Graphormer O(N²·d)=320,000 vs GCN O(|E|·d)=15,360.
  • C. About 50× — because Graphormer needs N² pairs while GCN needs |E|=2N pairs.
  • D. Identical — Graphormer masks to edges anyway, making it O(|E|·d).

Survives elimination: B

Why: Graphormer: O(N²·d) = 50²×128 = 320,000. GCN: O(|E|·d) = 120×128 = 15,360. Ratio = 320000/15360 ≈ 20.8×, closest to option B's ~26× (the approximation difference comes from ignoring constant factors). The key point: Graphormer's cost scales quadratically in N, GCN's scales with edge count.

45. Check 3: Graphormer vs GNN complexity

Check

A molecule has N=50 atoms, d=128, and 60 bonds (|E|=120 counting both directions). Compare one Graphormer attention layer to one GCN message-passing layer.

Check your understanding

For N=50, |E|=120, d=128, how many times more expensive (approximately) is one Graphormer layer vs one GCN layer in terms of multiply-accumulate operations?

  • A. About 2× — both are roughly linear in N.
  • B. About 26× — Graphormer O(N²·d)=320,000 vs GCN O(|E|·d)=15,360. (correct)
  • C. About 50× — because Graphormer needs N² pairs while GCN needs |E|=2N pairs.
  • D. Identical — Graphormer masks to edges anyway, making it O(|E|·d).

Answer: B

Why: Graphormer: O(N²·d) = 50²×128 = 320,000. GCN: O(|E|·d) = 120×128 = 15,360. Ratio = 320000/15360 ≈ 20.8×, closest to option B's ~26× (the approximation difference comes from ignoring constant factors). The key point: Graphormer's cost scales quadratically in N, GCN's scales with edge count.

Why A tempts people
Graphormer's all-pairs attention is O(N²), not O(N). Even for modest N=50, the quadratic term dominates.
Why C tempts people
50× would be O(N²)/O(N·d)·constant, but |E|=120 ≠ 2N=100 in general. The exact ratio depends on graph sparsity; 26× is the computed value for this graph.
Why D tempts people
D describes the hard-mask sparse Graph Transformer, not Graphormer. Graphormer uses a soft additive bias (never -∞), so it computes all N² attention pairs.

46. Predict the next row: Full Graphormer mini-module in PyTorch

Pattern

Predict first

The table runs: C (0) | val[0] | val[1] | val[2] | val[3] · Shape | (4, 4) | — | — | —

In Full Graphormer mini-module in PyTorch, given the rows so far: what is the next one — the row where Node is Module params?

Correct: Module params | 3×(8×4)+5×1 | = | 101 | params

Nodeout[0]out[1]out[2]out[3]
C (0)val[0]val[1]val[2]val[3]
Shape(4, 4)———
Module params3×(8×4)+5×1=101params

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The SPD matrix is a precomputed graph property (BFS); we convert it to bias via a learned embedding table phi_table of size max_dist+1.

47. Full Graphormer mini-module in PyTorch

Worked example

Build GraphormerAttention: takes node embeddings X and precomputed SPD matrix, returns updated node representations

Why: The SPD matrix is a precomputed graph property (BFS); we convert it to bias via a learned embedding table phi_table of size max_dist+1.

import torch, torch.nn as nn, torch.nn.functional as F

class GraphormerAttention(nn.Module):
    def __init__(self, d_model=8, d_k=4, max_dist=4):
        super().__init__()
        self.d_k = d_k
        self.Wq = nn.Linear(d_model, d_k, bias=False)
        self.Wk = nn.Linear(d_model, d_k, bias=False)
        self.Wv = nn.Linear(d_model, d_k, bias=False)
        # phi_table[d] = learned scalar bias for SPD=d
        self.phi = nn.Embedding(max_dist + 1, 1)

    def forward(self, X, spd):
        # X: (N, d_model)  spd: (N, N) int tensor
        Q, K, V = self.Wq(X), self.Wk(X), self.Wv(X)
        scores = Q @ K.T / (self.d_k ** 0.5)   # (N, N)
        bias = self.phi(spd.clamp(0, 4)).squeeze(-1)  # (N, N)
        attn = F.softmax(scores + bias, dim=-1)
        return attn @ V                          # (N, d_k)

# Test on toy molecule
torch.manual_seed(42)
N, d = 4, 8
X = torch.eye(N)
spd = torch.tensor([[0,1,1,2],[1,0,2,1],[1,2,0,3],[2,1,3,0]])
model = GraphormerAttention()
out = model(X, spd)
print(out.shape, out.detach().numpy().round(4))
Nodeout[0]out[1]out[2]out[3]
C (0)val[0]val[1]val[2]val[3]
Shape(4, 4)———
Module params3×(8×4)+5×1=101params

Verify output shape: (N=4, d_k=4) and parameter count: 3×(8×4) + 5 = 101

Why: 3 linear layers of 8→4 = 3×32=96 weights; phi Embedding with 5 entries × 1 scalar = 5 weights. Total = 101. Output shape (4,4) confirmed by print.

48. Fill in: out[2] for Full Graphormer mini-module in PyTorch

Comparison

Comparison matrix

From Full Graphormer mini-module in PyTorch: refill the out[2] column from what you know. The rest of the table is as it appeared.

Nodeout[0]out[1]out[2]out[3]
C (0)val[0]val[1]val[2]val[3]
Shape(4, 4)———
Module params3×(8×4)+5×1=101params

49. Callbacks: how Lesson 96 connects to earlier decks

Concept

50. What has to happen first: Your turn — Graphormer attention from scratch

Ranking

Put in order

Put the moves of Your turn — Graphormer attention from scratch into the order they have to happen.

  1. Milestone 1 — compute SPD for a path graph of 5 nodes
  2. Milestone 2 — build the phi table and spatial bias matrix B
  3. Milestone 3 — run biased attention for node 0; print attn_weights[0] and verify attn[0,1] > attn[0,4]
  4. Show it off: print the full 5×5 attention weight matrix and verify the diagonal dominance pattern along rows

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. For a path 0–1–2–3–4, SPD[i,j] = |i−j|.

51. Your turn — Graphormer attention from scratch

Worked example

Project brief: implement Graphormer single-head attention on a 5-node path graph (0–1–2–3–4). Compute the SPD matrix by hand, build the spatial bias, run biased attention, and verify that node 0 attends most to its neighbour node 1 and least to node 4 (SPD=4).

Milestone 1 — compute SPD for a path graph of 5 nodes

Why: For a path 0–1–2–3–4, SPD[i,j] = |i−j|. Write the 5×5 SPD matrix and confirm the diagonal is 0 and the corners SPD[0,4]=SPD[4,0]=4.

Milestone 2 — build the phi table and spatial bias matrix B

Why: Use phi_weights = [0.0, -0.5, -1.2, -2.1, -3.0] (same as lesson). Construct a 5×5 tensor B where B[i,j]=phi(min(|i-j|,4)). Check B[0,4]=B[4,0]=-3.0.

Milestone 3 — run biased attention for node 0; print attn_weights[0] and verify attn[0,1] > attn[0,4]

Why: Node 0 has SPD=1 to node 1 and SPD=4 to node 4. phi(1)=-0.50 vs phi(4)=-3.00: a 2.50 penalty differential guarantees more weight on node 1 (barring extreme raw score differences).

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
N, d_model, d_k = 5, 8, 4
# 5-node path graph: SPD[i,j] = |i-j|
spd = torch.tensor([[abs(i-j) for j in range(N)] for i in range(N)])
phi_w = torch.tensor([0.0, -0.5, -1.2, -2.1, -3.0])
B = phi_w[spd.clamp(0,4)]               # (5,5) spatial bias
X = torch.eye(N)                         # one-hot node features
proj = nn.Linear(N, d_model, bias=False)
X_emb = proj(X)
Wq = nn.Linear(d_model, d_k, bias=False)
Wk = nn.Linear(d_model, d_k, bias=False)
Wv = nn.Linear(d_model, d_k, bias=False)
for m in [Wq,Wk,Wv]: nn.init.normal_(m.weight, std=0.3)
Q, K, V = Wq(X_emb), Wk(X_emb), Wv(X_emb)
scores = Q @ K.T / d_k**0.5
attn = F.softmax(scores + B, dim=-1)
print('attn[0]:', attn[0].detach().numpy().round(4))
print('node1 > node4:', attn[0,1].item() > attn[0,4].item())
attn[0,j]j=0 (self)j=1j=2j=3j=4
Expected orderhighest or 2ndhigh (SPD=1)mediumlowlowest (SPD=4)
phi bias0.00-0.50-1.20-2.10-3.00

Show it off: print the full 5×5 attention weight matrix and verify the diagonal dominance pattern along rows

Why: Each row i should show decreasing attention as SPD increases from 0 outward. This is the graph-aware inductive bias: closer nodes, more attention — without any positional integer index.

52. Inspect it line by line: Your turn — Graphormer attention from scratch

Error analysis

Annotate

Walk the callouts on Your turn — Graphormer attention from scratch. Each one is a place this is easy to get subtly wrong.

  • For a path 0–1–2–3–4, SPD[i,j] = |i−j|. Write the 5×5 SPD matrix and confirm the diagonal is 0 and the corners SPD[0,4]=SPD[4,0]=4.
  • Use phi_weights = [0.0, -0.5, -1.2, -2.1, -3.0] (same as lesson). Construct a 5×5 tensor B where B[i,j]=phi(min(|i-j|,4)). Check B[0,4]=B[4,0]=-3.0.
  • Node 0 has SPD=1 to node 1 and SPD=4 to node 4. phi(1)=-0.50 vs phi(4)=-3.00: a 2.50 penalty differential guarantees more weight on node 1 (barring extreme raw score differences).

53. Connect it up: Lesson 96: Graph Transformer & Graphormer

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why graphs? Transformers already are graphs · Graphormer's three structural encodings · Soft bias vs hard mask — two Graph Transformer flavours · Complexity, scaling, applications. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

54. Recap: Graph Transformer & Graphormer

Recap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 96 — Graph Transformers — Barron · USAAIO Round 2 Preparation, 2026
  2. Ying et al. 'Do Transformers Really Perform Bad for Graph Representation?' (NeurIPS 2021) — Graphormer — arXiv:2106.05234
  3. Dwivedi & Bresson, 'A Generalization of Transformers to Graphs' (2020) — arXiv:2012.09699
  4. All attention weights, SPD, spatial bias, and edge encoding verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108