USAAIO Lesson 83, from Phase 3. It explains why self-attention is permutation-equivariant and therefore needs positional information, then covers sinusoidal positional encoding, with its formula and the proof that relative position is linear in it; learned positional encoding via nn.Embedding; RoPE, which rotates Q and K by an angle set by position and so gives a dot product that depends on relative position; and ALiBi, which adds a linear bias to the attention scores and extrapolates to longer sequences. All the values were computed with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 83 · Phase 3 (Transformers & NLP)
Self-attention is order-blind without help. Today: sinusoidal PE, learned PE, RoPE (LLaMA), and ALiBi — and the precise math that makes each one tick.
Objectives
ddot(Q_rot[m], K_rot[n]) encodes relative positionWarm-up
Discussion prompt
Before we open Lesson 83: Positional Encoding: without looking back, what was the main idea of LayerNorm, FFN with GELU, and the Full Encoder Block, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
LayerNorm from scratch (normalize over the feature dimension, verified against nn.LayerNorm), pre-norm vs post-norm stability, the FFN sublayer (two Linear layers + GELU, d_ff = 4·d_model), residual connections for gradient flow, and the complete pre-norm EncoderBlock as an nn.Module.
Section
Part 1 of 4
Concept
In Lesson 82 you saw that attention computes softmax(QK^T/√d)V. The score for token i against token j is purely their dot product — it does not know whether j is two slots left or fifty.
Permutation equivariance: shuffle all token embeddings in the same way and every output embedding shuffles identically. The attention pattern is the same regardless of order.
Without extra position information, "The cat sat on the mat" and "The mat sat on the cat" produce the same set of output vectors (just in different row order). Positional encoding breaks this symmetry.
Counterexample
Discussion prompt
In Lesson 82 you saw that attention computes softmax(QK^T/√d)V. The score for token i against token j is purely their dot product — it does not know whether j is two slots left or fifty.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
| approach | how it works | where in the model |
|---|---|---|
| Add to input embedding | PE vector added to token embedding before the first layer | sinusoidal PE, learned PE |
| Rotate Q and K | Multiply Q and K by position-dependent rotation matrices | RoPE (LLaMA, GPT-NeoX) |
| Bias attention scores | Add −slope × |i−j| to raw QK scores before softmax | ALiBi |
All three inject the same information — relative or absolute position — but with different extrapolation and computational properties.
Comparison
Comparison matrix
From Three places to inject position: refill the how it works column from what you know. The rest of the table is as it appeared.
| approach | how it works | where in the model |
|---|---|---|
| Add to input embedding | PE vector added to token embedding before the first layer | sinusoidal PE, learned PE |
| Rotate Q and K | Multiply Q and K by position-dependent rotation matrices | RoPE (LLaMA, GPT-NeoX) |
| Bias attention scores | Add −slope × |i−j| to raw QK scores before softmax | ALiBi |
Section
Part 2 of 4
Concept
For position pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:
\[ \text{PE}(\text{pos},\,2i) = \sin\!\left(\frac{\text{pos}}{10000^{2i/d}}\right), \quad \text{PE}(\text{pos},\,2i+1) = \cos\!\left(\frac{\text{pos}}{10000^{2i/d}}\right) \]
Low-index dimensions oscillate fast (period ≈ 2π); high-index dimensions oscillate extremely slowly — the rightmost dims barely move across a typical sequence length. Every position gets a unique fingerprint.
Analogy
Discussion prompt
Explain The sinusoidal PE formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For position pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:
Estimation
Predict first
With d=4 we have two frequency pairs (i=0 and i=1). The denominators are 10000^0=1 and 10000^0.5=100. Compute the first 6 rows.
Commit before you compute: what does Computing sinusoidal PE for d=4 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: pos=0 → all-sin dims = 0; all-cos dims = 1. pos=1: fast dim (i=0) already at sin(1)=0.8415.
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. At pos=0, sin(0)=0 and cos(0)=1 for every frequency.
Worked example
With d=4 we have two frequency pairs (i=0 and i=1). The denominators are 10000^0=1 and 10000^0.5=100. Compute the first 6 rows.
import numpy as np, math
def sinusoidal_pe(max_len, d):
pe = np.zeros((max_len, d))
for pos in range(max_len):
for i in range(d // 2):
w = 10000 ** (2*i / d)
pe[pos, 2*i] = math.sin(pos / w)
pe[pos, 2*i+1] = math.cos(pos / w)
return pe
pe = sinusoidal_pe(6, 4)
print(pe.round(4))pos=0 → all-sin dims = 0; all-cos dims = 1. pos=1: fast dim (i=0) already at sin(1)=0.8415.
Why: At pos=0, sin(0)=0 and cos(0)=1 for every frequency. The first even/odd pair oscillates at period 2π≈6.28 tokens; the second pair at period 200π≈628 tokens.
| pos | dim0 (sin,fast) | dim1 (cos,fast) | dim2 (sin,slow) | dim3 (cos,slow) |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| 1 | 0.8415 | 0.5403 | 0.0100 | 1.0000 |
| 2 | 0.9093 | -0.4161 | 0.0200 | 0.9998 |
| 3 | 0.1411 | -0.9900 | 0.0300 | 0.9996 |
| 4 | -0.7568 | -0.6536 | 0.0400 | 0.9992 |
| 5 | -0.9589 | 0.2837 | 0.0500 | 0.9988 |
Trade off
Comparison matrix
From Computing sinusoidal PE for d=4: every row here is a choice with a cost. Fill the dim2 (sin,slow) column, then say which row you would actually pick and what you give up for it.
| pos | dim0 (sin,fast) | dim1 (cos,fast) | dim2 (sin,slow) | dim3 (cos,slow) |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| 1 | 0.8415 | 0.5403 | 0.0100 | 1.0000 |
| 2 | 0.9093 | -0.4161 | 0.0200 | 0.9998 |
| 3 | 0.1411 | -0.9900 | 0.0300 | 0.9996 |
| 4 | -0.7568 | -0.6536 | 0.0400 | 0.9992 |
| 5 | -0.9589 | 0.2837 | 0.0500 | 0.9988 |
Concept
Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.
\[ \text{PE}(\text{pos}+k,\,2i) = \sin\!\left(\frac{\text{pos}}{w_i} + \frac{k}{w_i}\right) = \text{PE}(\text{pos},2i)\cos\!\frac{k}{w_i} + \text{PE}(\text{pos},2i+1)\sin\!\frac{k}{w_i} \]
The coefficients cos(k/w_i) and sin(k/w_i) depend only on k and i, not on the absolute position. This means a linear attention head can extract the k-offset neighbor from any position.
| LHS (direct) | RHS (linear combo) | match? |
|---|---|---|
| PE(5, 0) = sin(5) = -0.9589 | PE(3,0)·cos(2) + PE(3,1)·sin(2) = 0.1411·(-0.4161) + (-0.9900)·0.9093 = -0.9589 | True |
Explain it
Discussion prompt
Explain Relative position via linear combination to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.
Missing information
Discussion prompt
The standard implementation pre-computes the full PE buffer using torch.exp(log-space) for numerical stability, then adds it to the input in forward.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
register_buffer stores the PE as a non-trainable tensor that moves to the same device as the model. Adding to x broadcasts over the batch dimension.
Worked example
The standard implementation pre-computes the full PE buffer using torch.exp(log-space) for numerical stability, then adds it to the input in forward.
import torch, torch.nn as nn, math
class SinusoidalPE(nn.Module):
def __init__(self, d_model, max_len=512):
super().__init__()
pe = torch.zeros(max_len, d_model)
pos = torch.arange(max_len).float().unsqueeze(1)
div = torch.exp(torch.arange(0, d_model, 2).float()
* (-math.log(10000.0) / d_model))
pe[:, 0::2] = torch.sin(pos * div)
pe[:, 1::2] = torch.cos(pos * div)
self.register_buffer('pe', pe) # not a parameter
def forward(self, x): # x: (B, T, d)
return x + self.pe[:x.shape[1]] # broadcast over batch
pe_mod = SinusoidalPE(16, max_len=50)
x = torch.zeros(2, 6, 16)
print('out shape:', pe_mod(x).shape)
print('PE[0,:4]:', [round(v,4) for v in pe_mod.pe[0,:4].tolist()])
print('PE[1,:4]:', [round(v,4) for v in pe_mod.pe[1,:4].tolist()])out shape = (2, 6, 16); PE[0,:4] = [0.0, 1.0, 0.0, 1.0]; PE[1,:4] = [0.8415, 0.5403, 0.311, 0.9504]
Why: register_buffer stores the PE as a non-trainable tensor that moves to the same device as the model. Adding to x broadcasts over the batch dimension.
| PE row | dim 0 (sin) | dim 1 (cos) | dim 2 (sin) | dim 3 (cos) |
|---|---|---|---|---|
| pos=0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| pos=1 | 0.8415 | 0.5403 | 0.3110 | 0.9504 |
| pos=2 | 0.9093 | -0.4161 | 0.5878 | 0.8090 |
Pattern
Step through it
Step through PyTorch sinusoidal PE module one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Store the sinusoidal PE with self.pe = pe (plain Python attribute) or with nn.Parameter(pe).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.
Use self.register_buffer('pe', pe) for fixed, device-aware, non-trainable storage.
Why: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.
Trap
Store the sinusoidal PE with self.pe = pe (plain Python attribute) or with nn.Parameter(pe).
Use nn.Parameter(pe) so the PE can be fine-tuned
Why: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.
Use self.register_buffer('pe', pe) for fixed, device-aware, non-trainable storage.
register_buffer: moves with .to(device), saves/loads with state_dict, but never appears in model.parameters()
Why: Sinusoidal PE is closed-form — there is nothing to learn. register_buffer is the canonical PyTorch idiom for non-parameter tensors that belong to the module.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Sinusoidal PE is closed-form — there is nothing to learn. register_buffer is the canonical PyTorch idiom for non-parameter tensors that belong to the module.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.
Concept
Replace the formula with a simple nn.Embedding(max_len, d_model) — one learned vector per position, trained jointly with the rest of the model. Used in BERT, GPT-2, ViT.
| property | sinusoidal | learned |
|---|---|---|
| extra parameters | 0 (fixed formula) | max_len × d_model (e.g. 512×768 = 393,216 for BERT) |
| relative-position structure | guaranteed by construction | must be learned from data |
| generalizes to longer sequences | yes — formula works for any pos | no — OOV beyond max_train_len |
| used in | original Transformer, T5 | BERT, GPT-2, ViT |
Empirically, learned PE often matches or beats sinusoidal on tasks within the training length, but fails silently when inference sequences exceed max_len.
Section
Part 3 of 4
Concept
RoPE (Su et al., 2021; used in LLaMA, GPT-NeoX, Mistral) injects position by rotating query and key vectors by a position-dependent angle before the dot product.
\[ f_q(\mathbf{x}_m, m) = \mathbf{R}_m\, \mathbf{q}_m, \quad f_k(\mathbf{x}_n, n) = \mathbf{R}_n\, \mathbf{k}_n \]
\[ (\mathbf{R}_m \mathbf{q}_m)^\top (\mathbf{R}_n \mathbf{k}_n) = \mathbf{q}_m^\top \mathbf{R}_m^\top \mathbf{R}_n\, \mathbf{k}_n = \mathbf{q}_m^\top \mathbf{R}_{n-m}\, \mathbf{k}_n \]
Because R_m^T R_n = R_{n-m} (rotation matrices compose by subtraction), the dot product Q_rot[m]·K_rot[n] is a function of the relative position n-m only, not the absolute positions.
Estimation
Predict first
Build the cos/sin cache and apply the rotation pair-wise to (x1, x2) sub-vectors. Verify shape and relative-position property.
Commit before you compute: what does RoPE implementation from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: cos_cache (6,4), Q_rot (6,8). Dot products for offset=2: vary across absolute positions (content-dependent), but different offsets give structurally different scores.
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. RoPE encodes relative position in the rotation geometry.
Worked example
Build the cos/sin cache and apply the rotation pair-wise to (x1, x2) sub-vectors. Verify shape and relative-position property.
import torch, math
torch.manual_seed(42)
seq, d = 6, 8
def rope_cache(seq, d):
theta = 10000.0 ** (-torch.arange(0, d, 2).float() / d)
pos = torch.arange(seq).float()
freqs = torch.outer(pos, theta) # (seq, d/2)
return torch.cos(freqs), torch.sin(freqs)
def apply_rope(x, cos, sin): # x: (seq, d)
x1, x2 = x[..., :d//2], x[..., d//2:]
return torch.cat([x1*cos - x2*sin,
x1*sin + x2*cos], dim=-1)
cos_c, sin_c = rope_cache(seq, d)
Q = torch.randn(seq, d)
K = torch.randn(seq, d)
Q_r = apply_rope(Q, cos_c, sin_c)
K_r = apply_rope(K, cos_c, sin_c)
print('cos_cache shape:', cos_c.shape)
print('Q_rot shape: ', Q_r.shape)
for m in range(2, seq):
n = m - 2
print(f' offset=2 Q[{m}]·K[{n}]:', round((Q_r[m]*K_r[n]).sum().item(), 4))cos_cache (6,4), Q_rot (6,8). Dot products for offset=2: vary across absolute positions (content-dependent), but different offsets give structurally different scores.
Why: RoPE encodes relative position in the rotation geometry. Dot product magnitude also depends on Q and K content — what RoPE guarantees is that the positional BIAS term factors out as R_{n-m}, not that all same-offset pairs have the same value.
| pair (m,n) | offset | Q_rot[m]·K_rot[n] |
|---|---|---|
| (2,0) | 2 | -1.1148 |
| (3,1) | 2 | 3.2626 |
| (3,2) | 1 | 6.0944 |
| (3,1) | 2 | 3.2626 |
Pattern
Step through it
Step through RoPE implementation from scratch one row at a time. What is driving the change, and what would the row after the last one be?
Concept
For d=8, the rotation frequencies are theta_i = 10000^(-2i/d) for i=0..3. Low-index pairs rotate fast; high-index pairs rotate slowly — the same frequency hierarchy as sinusoidal PE.
| pair index i | theta_i | angle at pos=1 | angle at pos=4 |
|---|---|---|---|
| 0 | 1.0000 | 1.0000 rad (57.3°) | 4.0000 rad (229°) |
| 1 | 0.1000 | 0.1000 rad (5.7°) | 0.4000 rad (22.9°) |
| 2 | 0.0100 | 0.0100 rad (0.57°) | 0.0400 rad (2.3°) |
| 3 | 0.0010 | 0.0010 rad (0.057°) | 0.0040 rad (0.23°) |
Fast-rotating pairs carry short-range position signal; slow-rotating pairs carry long-range signal. Unlike additive PE, RoPE is applied inside the attention computation — only to Q and K, never to V.
Comparison
Comparison matrix
From RoPE: theta values and frequency structure: refill the theta_i column from what you know. The rest of the table is as it appeared.
| pair index i | theta_i | angle at pos=1 | angle at pos=4 |
|---|---|---|---|
| 0 | 1.0000 | 1.0000 rad (57.3°) | 4.0000 rad (229°) |
| 1 | 0.1000 | 0.1000 rad (5.7°) | 0.4000 rad (22.9°) |
| 2 | 0.0100 | 0.0100 rad (0.57°) | 0.0400 rad (2.3°) |
| 3 | 0.0010 | 0.0010 rad (0.057°) | 0.0040 rad (0.23°) |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.
Apply RoPE only to Q and K — the vectors used to compute similarity scores.
Why: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.
Trap
Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.
Rotate V by position before the weighted sum
Why: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.
Apply RoPE only to Q and K — the vectors used to compute similarity scores.
Rotate Q and K; leave V untouched
Why: The positional encoding purpose is to make the similarity score between token m and token n depend on their relative offset. V carries content; the attention weights already encode position via the rotated Q·K score.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:; Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.self.pe = pe (plain Python attribute) or with nn.Parameter(pe).; Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.Section
Part 4 of 4
Concept
ALiBi (Press et al., 2021) skips embedding-level position encoding entirely. Instead, it adds a linear negative bias to each raw attention score proportional to distance, before softmax.
\[ \text{Attention score}_{m,n} = \frac{\mathbf{q}_m \cdot \mathbf{k}_n}{\sqrt{d}} + m_{\text{head}} \cdot (n - m) \]
The slope m_head is head-specific and fixed (not learned): m_h = 2^{-8h/H} for head h = 1, …, H. For H=4: slopes = 0.25, 0.0625, 0.0156, 0.0039.
Explain it
Discussion prompt
Explain ALiBi: linear penalty on distance to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
ALiBi (Press et al., 2021) skips embedding-level position encoding entirely. Instead, it adds a linear negative bias to each raw attention score proportional to distance, before softmax.
Estimation
Predict first
Compute the bias matrix for head 0 (slope=0.25) on a 5-token sequence. Rows = query positions, cols = key positions.
Commit before you compute: what does ALiBi bias matrix (4 heads, seq=5) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: slopes = [0.25, 0.0625, 0.0156, 0.0039]. bias_h0[3,1] = 0.25*(1-3) = -0.5. Nearer keys get a smaller penalty.
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Farther keys receive a larger negative adjustment, softly encouraging the model to attend to nearby tokens.
Worked example
Compute the bias matrix for head 0 (slope=0.25) on a 5-token sequence. Rows = query positions, cols = key positions.
import torch
H, seq = 4, 5
slopes = torch.tensor([2**(-8.0/H * h) for h in range(1, H+1)])
print('slopes:', [round(s,4) for s in slopes.tolist()])
# head 0 bias matrix
pos = torch.arange(seq).float()
rel = pos.unsqueeze(1) - pos.unsqueeze(0) # key_pos - query_pos
bias_h0 = slopes[0] * rel # positive = closer key; causal: mask future
print(bias_h0.numpy().round(4))slopes = [0.25, 0.0625, 0.0156, 0.0039]. bias_h0[3,1] = 0.25*(1-3) = -0.5. Nearer keys get a smaller penalty.
Why: Farther keys receive a larger negative adjustment, softly encouraging the model to attend to nearby tokens. Head 0 (steepest slope) is the most local; head 3 (flattest) is the most global.
| distance |m-n| | slope=0.25 (h=0) | slope=0.0625 (h=1) | slope=0.0156 (h=2) | slope=0.0039 (h=3) |
|---|---|---|---|---|
| 0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 1 | -0.25 | -0.0625 | -0.0156 | -0.0039 |
| 2 | -0.50 | -0.1250 | -0.0313 | -0.0078 |
| 4 | -1.00 | -0.2500 | -0.0625 | -0.0156 |
Pattern
Step through it
Step through ALiBi bias matrix (4 heads, seq=5) one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Because ALiBi's bias is a simple formula (not a lookup table), it generates a well-defined penalty for any distance — including distances beyond the training length. The model was never shown position 2048 during training at length 1024, but the bias at distance 2000 is simply −slope × 2000.
| scheme | generalizes past train length? | extra params |
|---|---|---|
| sinusoidal PE | yes (formula, no OOV) | 0 |
| learned PE | no (OOV beyond max_len) | max_len × d_model |
| RoPE | yes (rotation formula) | 0 (cache recomputed) |
| ALiBi | yes (linear formula) | 0 (slopes fixed by head) |
ALiBi also removes PE entirely from the input, slightly simplifying the architecture. The trade-off: it bakes in a locality inductive bias (nearby tokens preferred) that may not suit all tasks.
Analogy
Discussion prompt
Explain ALiBi extrapolation advantage by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
ALiBi also removes PE entirely from the input, slightly simplifying the architecture. The trade-off: it bakes in a locality inductive bias (nearby tokens preferred) that may not suit all tasks.
Constraint
Discussion prompt
Run Positional encoding decision recipe with this step confiscated:
Need relative-position geometry inside attention: use RoPE (apply to Q and K only; leave V untouched)
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
max_lenregister_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance…Pattern
max_lenregister_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance added to raw scoresEdge cases
Discussion prompt
Positional encoding decision recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
max_lenregister_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance…Elimination
Eliminate the wrong options
A transformer encoder with no positional encoding processes the sequence [A, B, C]. You shuffle it to [C, A, B]. What is true of the output?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Self-attention is permutation-equivariant: each output vector changes because the query that produced it changed position, but the full set of output vectors is just a reordering. The attention pattern (softmax of QK^T) is symmetric in the same way the input was shuffled.
Check
Think about what changes and what doesn't when you permute the inputs.
Check your understanding
A transformer encoder with no positional encoding processes the sequence [A, B, C]. You shuffle it to [C, A, B]. What is true of the output?
Answer: A
Why: Self-attention is permutation-equivariant: each output vector changes because the query that produced it changed position, but the full set of output vectors is just a reordering. The attention pattern (softmax of QK^T) is symmetric in the same way the input was shuffled.
Prediction
Predict first
For d=4, what is PE(1, 2) — i.e., position=1, dimension index=2?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: sin(1/100) ≈ 0.01
Why: Dimension index 2 is even, so it uses sin. The pair index is i = 2/2 = 1. The denominator is 10000^(2·1/4) = 10000^0.5 = 100. So PE(1, 2) = sin(1/100) ≈ 0.0100. Verified in the d=4 trace table.
Check
Use the formula: PE(pos, 2i) = sin(pos / 10000^(2i/d)). Work it out for the specific case before clicking.
Check your understanding
For d=4, what is PE(1, 2) — i.e., position=1, dimension index=2?
Answer: A
Why: Dimension index 2 is even, so it uses sin. The pair index is i = 2/2 = 1. The denominator is 10000^(2·1/4) = 10000^0.5 = 100. So PE(1, 2) = sin(1/100) ≈ 0.0100. Verified in the d=4 trace table.
Elimination
Eliminate the wrong options
Which statement correctly distinguishes RoPE from ALiBi?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: RoPE rotates Q and K by position-dependent angles before computing QK^T, so the dot product itself encodes relative position. ALiBi skips all positional vectors and instead adds −slope×|i−j| to the raw QK scores before softmax. Both require zero extra parameters.
Check
One key operational difference separates RoPE from ALiBi.
Check your understanding
Which statement correctly distinguishes RoPE from ALiBi?
Answer: A
Why: RoPE rotates Q and K by position-dependent angles before computing QK^T, so the dot product itself encodes relative position. ALiBi skips all positional vectors and instead adds −slope×|i−j| to the raw QK scores before softmax. Both require zero extra parameters.
Section
Project
Concept
Build and verify all three encode methods: sinusoidal PE, a simple learned PE, and RoPE. Each milestone has a concrete test you can run to confirm correctness.
| # | milestone | correctness test |
|---|---|---|
| 1 | Sinusoidal PE (numpy) | PE[0] = [0,1,0,1,...]; verify linear-combo identity for k=2 |
| 2 | Learned PE (nn.Embedding) | Shape (T, d); confirm no PE in model.parameters() when frozen |
| 3 | RoPE (cos/sin cache + apply_rope) | Confirm Q_rot shape, verify different offsets give different dot products |
Build rules: type every line, use register_buffer for sinusoidal PE, and never rotate V in RoPE.
Counterexample
Discussion prompt
Build and verify all three encode methods: sinusoidal PE, a simple learned PE, and RoPE. Each milestone has a concrete test you can run to confirm correctness.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line, use register_buffer for sinusoidal PE, and never rotate V in RoPE.
Worked example
Your turn: implement sinusoidal PE for d=4, compute PE[3] and PE[5], then verify that PE[5,0] = PE[3,0]·cos(2) + PE[3,1]·sin(2).
Hint: denom = 10000 ** (2*i/d), even dims → sin, odd → cos. For i=0, d=4: denom=1.0. cos(2) = -0.4161, sin(2) = 0.9093.
import numpy as np, math
def sinusoidal_pe(L, d):
pe = np.zeros((L, d))
for pos in range(L):
for i in range(d//2):
w = 10000 ** (2*i/d)
pe[pos,2*i] = math.sin(pos/w)
pe[pos,2*i+1] = math.cos(pos/w)
return pe
pe = sinusoidal_pe(6, 4)
print('PE[3]:', pe[3].round(4))
print('PE[5]:', pe[5].round(4))
lhs = pe[5, 0]
rhs = pe[3,0]*math.cos(2) + pe[3,1]*math.sin(2)
print(f'identity: {lhs:.6f} == {rhs:.6f}, match={abs(lhs-rhs)<1e-9}')| quantity | value |
|---|---|
| PE[3, 0] | 0.1411 (sin(3/1)) |
| PE[3, 1] | -0.9900 (cos(3/1)) |
| PE[5, 0] | -0.9589 (sin(5/1)) |
| 0.1411·cos(2) + (-0.9900)·sin(2) | -0.9589 (match: True) |
Worked example
Your turn: create an nn.Embedding(10, 8) learned PE, look up positions 0-4, and predict the weight matrix shape.
Hint: nn.Embedding(num_embeddings, embedding_dim). The weight matrix is (num_embeddings, embedding_dim). Accessed with learned_pe(torch.arange(T)).
import torch, torch.nn as nn
torch.manual_seed(42)
learned_pe = nn.Embedding(10, 8)
seq_pos = torch.arange(5)
emb = learned_pe(seq_pos)
print('lookup shape:', emb.shape) # (5, 8)
print('weight shape:', learned_pe.weight.shape) # (10, 8)| tensor | shape | note |
|---|---|---|
| embedding lookup for T=5 | (5, 8) | one 8-dim vector per position |
| weight matrix | (10, 8) | all 10 position vectors; requires grad |
| params count | 80 | 10 × 8 = 80 learned values |
Trade off
Comparison matrix
From Milestone 2 — learned PE shape check: every row here is a choice with a cost. Fill the shape column, then say which row you would actually pick and what you give up for it.
| tensor | shape | note |
|---|---|---|
| embedding lookup for T=5 | (5, 8) | one 8-dim vector per position |
| weight matrix | (10, 8) | all 10 position vectors; requires grad |
| params count | 80 | 10 × 8 = 80 learned values |
Worked example
Your turn: implement rope_cache and apply_rope, apply to Q and K (not V), and confirm shapes. Predict: will two pairs with the same offset always produce the same dot product?
Hint: theta = 10000^(-arange(0,d,2)/d); freqs = outer(pos, theta); rotation: [x1*cos - x2*sin, x1*sin + x2*cos]. The dot product depends on Q/K content too, not just position.
import torch
torch.manual_seed(42)
seq, d = 6, 8
def rope_cache(seq, d):
theta = 10000.0**(-torch.arange(0,d,2).float()/d)
freqs = torch.outer(torch.arange(seq).float(), theta)
return torch.cos(freqs), torch.sin(freqs)
def apply_rope(x, cos, sin):
x1, x2 = x[...,:d//2], x[...,d//2:]
return torch.cat([x1*cos-x2*sin, x1*sin+x2*cos], dim=-1)
cos_c, sin_c = rope_cache(seq, d)
Q = torch.randn(seq, d); K = torch.randn(seq, d)
Q_r = apply_rope(Q, cos_c, sin_c)
K_r = apply_rope(K, cos_c, sin_c)
print('Q_rot:', Q_r.shape)
for m in range(2, 5):
dot = (Q_r[m]*K_r[m-2]).sum().item()
print(f' Q[{m}]·K[{m-2}] (offset=2):', round(dot, 4))| pair | offset | dot product |
|---|---|---|
| Q[2]·K[0] | 2 | -1.1148 |
| Q[3]·K[1] | 2 | 3.2626 |
| Q[4]·K[2] | 2 | -1.5243 |
Comparison
Comparison matrix
From Milestone 3 — RoPE full program: refill the offset column from what you know. The rest of the table is as it appeared.
| pair | offset | dot product |
|---|---|---|
| Q[2]·K[0] | 2 | -1.1148 |
| Q[3]·K[1] | 2 | 3.2626 |
| Q[4]·K[2] | 2 | -1.5243 |
Concept
Out loud, slides closed: (1) explain why a pure self-attention layer treats a shuffled sequence identically, (2) state the sinusoidal PE formula and what the denominator 10000^(2i/d) controls, (3) explain why RoPE applies only to Q and K, and (4) state the ALiBi slope formula and why it generalizes to longer sequences.
Stretch (homework): implement the sinusoidal PE as a PyTorch nn.Module with register_buffer; implement ALiBi by adding the bias tensor to raw attention logits in a toy multi-head attention block; show mathematically that PE(pos+k,2i) is linear in PE(pos,2i) and PE(pos,2i+1). Next up: Lesson 84 — transformer architecture deep dive (encoder, decoder, cross-attention, layer norm placement).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why position matters at all · Sinusoidal positional encoding · RoPE — rotate Q and K · ALiBi — attention with linear biases · Your turn: implement sinusoidal PE + RoPE. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
PE(pos, 2i) = sin(pos / 10000^(2i/d)) and verify the relative-position linear identityregister_buffer and a learned PE with nn.Embeddingrope_cache + apply_rope) and apply it only to Q and K| scheme | where applied | generalizes beyond train length? | params |
|---|---|---|---|
| sinusoidal PE | add to input | yes | 0 |
| learned PE | add to input | no (OOV) | max_len × d |
| RoPE | rotate Q & K | yes | 0 |
| ALiBi | bias QK scores | yes | 0 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.