Lesson 83: Positional Encoding

USAAIO Lesson 83, from Phase 3. It explains why self-attention is permutation-equivariant and therefore needs positional information, then covers sinusoidal positional encoding, with its formula and the proof that relative position is linear in it; learned positional encoding via nn.Embedding; RoPE, which rotates Q and K by an angle set by position and so gives a dot product that depends on relative position; and ALiBi, which adds a linear bias to the attention scores and extrapolates to longer sequences. All the values were computed with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Positional Encoding Four Ways to Tell the Transformer Where You Are

Title

USAAIO · Lesson 83 · Phase 3 (Transformers & NLP)

Self-attention is order-blind without help. Today: sinusoidal PE, learned PE, RoPE (LLaMA), and ALiBi — and the precise math that makes each one tick.

2. By the end of this lesson you can

Objectives

  1. Explain why self-attention is permutation-equivariant and why it needs positional information
  2. Apply the sinusoidal PE formula and verify the relative-position linearity property
  3. Implement sinusoidal PE in PyTorch and trace exact output values for small d
  4. Implement RoPE from scratch and explain why dot(Q_rot[m], K_rot[n]) encodes relative position
  5. Describe ALiBi (linear bias added to attention scores) and its length-extrapolation advantage over learned PE

3. What survived from LayerNorm, FFN with GELU, and the Full Encoder Block?

Warm-up

Discussion prompt

Before we open Lesson 83: Positional Encoding: without looking back, what was the main idea of LayerNorm, FFN with GELU, and the Full Encoder Block, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

LayerNorm from scratch (normalize over the feature dimension, verified against nn.LayerNorm), pre-norm vs post-norm stability, the FFN sublayer (two Linear layers + GELU, d_ff = 4·d_model), residual connections for gradient flow, and the complete pre-norm EncoderBlock as an nn.Module.

4. Why position matters at all

Section

Part 1 of 4

5. Self-attention is permutation-equivariant

Concept

In Lesson 82 you saw that attention computes softmax(QK^T/√d)V. The score for token i against token j is purely their dot product — it does not know whether j is two slots left or fifty.

Permutation equivariance: shuffle all token embeddings in the same way and every output embedding shuffles identically. The attention pattern is the same regardless of order.

Without extra position information, "The cat sat on the mat" and "The mat sat on the cat" produce the same set of output vectors (just in different row order). Positional encoding breaks this symmetry.

6. Break it if you can: Self-attention is permutation-equivariant

Counterexample

Discussion prompt

In Lesson 82 you saw that attention computes softmax(QK^T/√d)V. The score for token i against token j is purely their dot product — it does not know whether j is two slots left or fifty.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Three places to inject position

Concept

approachhow it workswhere in the model
Add to input embeddingPE vector added to token embedding before the first layersinusoidal PE, learned PE
Rotate Q and KMultiply Q and K by position-dependent rotation matricesRoPE (LLaMA, GPT-NeoX)
Bias attention scoresAdd −slope × |i−j| to raw QK scores before softmaxALiBi

All three inject the same information — relative or absolute position — but with different extrapolation and computational properties.

8. Fill in: how it works for Three places to inject position

Comparison

Comparison matrix

From Three places to inject position: refill the how it works column from what you know. The rest of the table is as it appeared.

approachhow it workswhere in the model
Add to input embeddingPE vector added to token embedding before the first layersinusoidal PE, learned PE
Rotate Q and KMultiply Q and K by position-dependent rotation matricesRoPE (LLaMA, GPT-NeoX)
Bias attention scoresAdd −slope × |i−j| to raw QK scores before softmaxALiBi

9. Sinusoidal positional encoding

Section

Part 2 of 4

10. The sinusoidal PE formula

Concept

For position pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:

\[ \text{PE}(\text{pos},\,2i) = \sin\!\left(\frac{\text{pos}}{10000^{2i/d}}\right), \quad \text{PE}(\text{pos},\,2i+1) = \cos\!\left(\frac{\text{pos}}{10000^{2i/d}}\right) \]

Low-index dimensions oscillate fast (period ≈ 2π); high-index dimensions oscillate extremely slowly — the rightmost dims barely move across a typical sequence length. Every position gets a unique fingerprint.

11. By analogy: The sinusoidal PE formula

Analogy

Discussion prompt

Explain The sinusoidal PE formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

For position pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:

12. Guess the shape of the answer: Computing sinusoidal PE for d=4

Estimation

Predict first

With d=4 we have two frequency pairs (i=0 and i=1). The denominators are 10000^0=1 and 10000^0.5=100. Compute the first 6 rows.

Commit before you compute: what does Computing sinusoidal PE for d=4 come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: pos=0 → all-sin dims = 0; all-cos dims = 1. pos=1: fast dim (i=0) already at sin(1)=0.8415.

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. At pos=0, sin(0)=0 and cos(0)=1 for every frequency.

13. Computing sinusoidal PE for d=4

Worked example

With d=4 we have two frequency pairs (i=0 and i=1). The denominators are 10000^0=1 and 10000^0.5=100. Compute the first 6 rows.

import numpy as np, math
def sinusoidal_pe(max_len, d):
    pe = np.zeros((max_len, d))
    for pos in range(max_len):
        for i in range(d // 2):
            w = 10000 ** (2*i / d)
            pe[pos, 2*i]   = math.sin(pos / w)
            pe[pos, 2*i+1] = math.cos(pos / w)
    return pe
pe = sinusoidal_pe(6, 4)
print(pe.round(4))

pos=0 → all-sin dims = 0; all-cos dims = 1. pos=1: fast dim (i=0) already at sin(1)=0.8415.

Why: At pos=0, sin(0)=0 and cos(0)=1 for every frequency. The first even/odd pair oscillates at period 2π≈6.28 tokens; the second pair at period 200π≈628 tokens.

posdim0 (sin,fast)dim1 (cos,fast)dim2 (sin,slow)dim3 (cos,slow)
00.00001.00000.00001.0000
10.84150.54030.01001.0000
20.9093-0.41610.02000.9998
30.1411-0.99000.03000.9996
4-0.7568-0.65360.04000.9992
5-0.95890.28370.05000.9988

14. What each one costs: Computing sinusoidal PE for d=4

Trade off

Comparison matrix

From Computing sinusoidal PE for d=4: every row here is a choice with a cost. Fill the dim2 (sin,slow) column, then say which row you would actually pick and what you give up for it.

posdim0 (sin,fast)dim1 (cos,fast)dim2 (sin,slow)dim3 (cos,slow)
00.00001.00000.00001.0000
10.84150.54030.01001.0000
20.9093-0.41610.02000.9998
30.1411-0.99000.03000.9996
4-0.7568-0.65360.04000.9992
5-0.95890.28370.05000.9988

15. Relative position via linear combination

Concept

Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.

\[ \text{PE}(\text{pos}+k,\,2i) = \sin\!\left(\frac{\text{pos}}{w_i} + \frac{k}{w_i}\right) = \text{PE}(\text{pos},2i)\cos\!\frac{k}{w_i} + \text{PE}(\text{pos},2i+1)\sin\!\frac{k}{w_i} \]

The coefficients cos(k/w_i) and sin(k/w_i) depend only on k and i, not on the absolute position. This means a linear attention head can extract the k-offset neighbor from any position.

LHS (direct)RHS (linear combo)match?
PE(5, 0) = sin(5) = -0.9589PE(3,0)·cos(2) + PE(3,1)·sin(2) = 0.1411·(-0.4161) + (-0.9900)·0.9093 = -0.9589True

16. Teach it back: Relative position via linear combination

Explain it

Discussion prompt

Explain Relative position via linear combination to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.

17. What has to be given first: PyTorch sinusoidal PE module

Missing information

Discussion prompt

The standard implementation pre-computes the full PE buffer using torch.exp(log-space) for numerical stability, then adds it to the input in forward.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

register_buffer stores the PE as a non-trainable tensor that moves to the same device as the model. Adding to x broadcasts over the batch dimension.

18. PyTorch sinusoidal PE module

Worked example

The standard implementation pre-computes the full PE buffer using torch.exp(log-space) for numerical stability, then adds it to the input in forward.

import torch, torch.nn as nn, math
class SinusoidalPE(nn.Module):
    def __init__(self, d_model, max_len=512):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        pos = torch.arange(max_len).float().unsqueeze(1)
        div = torch.exp(torch.arange(0, d_model, 2).float()
                        * (-math.log(10000.0) / d_model))
        pe[:, 0::2] = torch.sin(pos * div)
        pe[:, 1::2] = torch.cos(pos * div)
        self.register_buffer('pe', pe)         # not a parameter
    def forward(self, x):                      # x: (B, T, d)
        return x + self.pe[:x.shape[1]]        # broadcast over batch

pe_mod = SinusoidalPE(16, max_len=50)
x = torch.zeros(2, 6, 16)
print('out shape:', pe_mod(x).shape)
print('PE[0,:4]:', [round(v,4) for v in pe_mod.pe[0,:4].tolist()])
print('PE[1,:4]:', [round(v,4) for v in pe_mod.pe[1,:4].tolist()])

out shape = (2, 6, 16); PE[0,:4] = [0.0, 1.0, 0.0, 1.0]; PE[1,:4] = [0.8415, 0.5403, 0.311, 0.9504]

Why: register_buffer stores the PE as a non-trainable tensor that moves to the same device as the model. Adding to x broadcasts over the batch dimension.

PE rowdim 0 (sin)dim 1 (cos)dim 2 (sin)dim 3 (cos)
pos=00.00001.00000.00001.0000
pos=10.84150.54030.31100.9504
pos=20.9093-0.41610.58780.8090

19. Watch it run: PyTorch sinusoidal PE module

Pattern

Step through it

Step through PyTorch sinusoidal PE module one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: PE row is pos=0
  2. Step 2: PE row is pos=1
  3. Step 3: PE row is pos=2

20. Something is wrong here: treating PE as a learned parameter

Anomaly

Predict first

A student writes this, and it looks reasonable:

Store the sinusoidal PE with self.pe = pe (plain Python attribute) or with nn.Parameter(pe).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.

Use self.register_buffer('pe', pe) for fixed, device-aware, non-trainable storage.

Why: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.

21. Trap: treating PE as a learned parameter

Trap

The trap

Store the sinusoidal PE with self.pe = pe (plain Python attribute) or with nn.Parameter(pe).

Use nn.Parameter(pe) so the PE can be fine-tuned

Why: This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.

The fix

Use self.register_buffer('pe', pe) for fixed, device-aware, non-trainable storage.

register_buffer: moves with .to(device), saves/loads with state_dict, but never appears in model.parameters()

Why: Sinusoidal PE is closed-form — there is nothing to learn. register_buffer is the canonical PyTorch idiom for non-parameter tensors that belong to the module.

22. Break it on purpose: treating PE as a learned parameter

Break the constraint

Discussion prompt

The rule this trap just fixed:

Sinusoidal PE is closed-form — there is nothing to learn. register_buffer is the canonical PyTorch idiom for non-parameter tensors that belong to the module.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This turns the fixed sinusoidal table into a gradient-consuming tensor — wasting computation and potentially allowing the PE to drift away from its analytical relative-position properties.

23. Learned PE — the lookup table

Concept

Replace the formula with a simple nn.Embedding(max_len, d_model) — one learned vector per position, trained jointly with the rest of the model. Used in BERT, GPT-2, ViT.

propertysinusoidallearned
extra parameters0 (fixed formula)max_len × d_model (e.g. 512×768 = 393,216 for BERT)
relative-position structureguaranteed by constructionmust be learned from data
generalizes to longer sequencesyes — formula works for any posno — OOV beyond max_train_len
used inoriginal Transformer, T5BERT, GPT-2, ViT

Empirically, learned PE often matches or beats sinusoidal on tasks within the training length, but fails silently when inference sequences exceed max_len.

24. RoPE — rotate Q and K

Section

Part 3 of 4

25. RoPE: the core idea

Concept

RoPE (Su et al., 2021; used in LLaMA, GPT-NeoX, Mistral) injects position by rotating query and key vectors by a position-dependent angle before the dot product.

\[ f_q(\mathbf{x}_m, m) = \mathbf{R}_m\, \mathbf{q}_m, \quad f_k(\mathbf{x}_n, n) = \mathbf{R}_n\, \mathbf{k}_n \]

\[ (\mathbf{R}_m \mathbf{q}_m)^\top (\mathbf{R}_n \mathbf{k}_n) = \mathbf{q}_m^\top \mathbf{R}_m^\top \mathbf{R}_n\, \mathbf{k}_n = \mathbf{q}_m^\top \mathbf{R}_{n-m}\, \mathbf{k}_n \]

Because R_m^T R_n = R_{n-m} (rotation matrices compose by subtraction), the dot product Q_rot[m]·K_rot[n] is a function of the relative position n-m only, not the absolute positions.

26. Guess the shape of the answer: RoPE implementation from scratch

Estimation

Predict first

Build the cos/sin cache and apply the rotation pair-wise to (x1, x2) sub-vectors. Verify shape and relative-position property.

Commit before you compute: what does RoPE implementation from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: cos_cache (6,4), Q_rot (6,8). Dot products for offset=2: vary across absolute positions (content-dependent), but different offsets give structurally different scores.

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. RoPE encodes relative position in the rotation geometry.

27. RoPE implementation from scratch

Worked example

Build the cos/sin cache and apply the rotation pair-wise to (x1, x2) sub-vectors. Verify shape and relative-position property.

import torch, math
torch.manual_seed(42)
seq, d = 6, 8

def rope_cache(seq, d):
    theta = 10000.0 ** (-torch.arange(0, d, 2).float() / d)
    pos   = torch.arange(seq).float()
    freqs = torch.outer(pos, theta)      # (seq, d/2)
    return torch.cos(freqs), torch.sin(freqs)

def apply_rope(x, cos, sin):             # x: (seq, d)
    x1, x2 = x[..., :d//2], x[..., d//2:]
    return torch.cat([x1*cos - x2*sin,
                      x1*sin + x2*cos], dim=-1)

cos_c, sin_c = rope_cache(seq, d)
Q = torch.randn(seq, d)
K = torch.randn(seq, d)
Q_r = apply_rope(Q, cos_c, sin_c)
K_r = apply_rope(K, cos_c, sin_c)
print('cos_cache shape:', cos_c.shape)
print('Q_rot shape:    ', Q_r.shape)
for m in range(2, seq):
    n = m - 2
    print(f'  offset=2 Q[{m}]·K[{n}]:', round((Q_r[m]*K_r[n]).sum().item(), 4))

cos_cache (6,4), Q_rot (6,8). Dot products for offset=2: vary across absolute positions (content-dependent), but different offsets give structurally different scores.

Why: RoPE encodes relative position in the rotation geometry. Dot product magnitude also depends on Q and K content — what RoPE guarantees is that the positional BIAS term factors out as R_{n-m}, not that all same-offset pairs have the same value.

pair (m,n)offsetQ_rot[m]·K_rot[n]
(2,0)2-1.1148
(3,1)23.2626
(3,2)16.0944
(3,1)23.2626

28. Watch it run: RoPE implementation from scratch

Pattern

Step through it

Step through RoPE implementation from scratch one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: pair (m,n) is (2,0)
  2. Step 2: pair (m,n) is (3,1)
  3. Step 3: pair (m,n) is (3,2)
  4. Step 4: pair (m,n) is (3,1)

29. RoPE: theta values and frequency structure

Concept

For d=8, the rotation frequencies are theta_i = 10000^(-2i/d) for i=0..3. Low-index pairs rotate fast; high-index pairs rotate slowly — the same frequency hierarchy as sinusoidal PE.

pair index itheta_iangle at pos=1angle at pos=4
01.00001.0000 rad (57.3°)4.0000 rad (229°)
10.10000.1000 rad (5.7°)0.4000 rad (22.9°)
20.01000.0100 rad (0.57°)0.0400 rad (2.3°)
30.00100.0010 rad (0.057°)0.0040 rad (0.23°)

Fast-rotating pairs carry short-range position signal; slow-rotating pairs carry long-range signal. Unlike additive PE, RoPE is applied inside the attention computation — only to Q and K, never to V.

30. Fill in: theta_i for RoPE: theta values and frequency structure

Comparison

Comparison matrix

From RoPE: theta values and frequency structure: refill the theta_i column from what you know. The rest of the table is as it appeared.

pair index itheta_iangle at pos=1angle at pos=4
01.00001.0000 rad (57.3°)4.0000 rad (229°)
10.10000.1000 rad (5.7°)0.4000 rad (22.9°)
20.01000.0100 rad (0.57°)0.0400 rad (2.3°)
30.00100.0010 rad (0.057°)0.0040 rad (0.23°)

31. Something is wrong here: applying RoPE to V as well

Anomaly

Predict first

A student writes this, and it looks reasonable:

Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.

Apply RoPE only to Q and K — the vectors used to compute similarity scores.

Why: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.

32. Trap: applying RoPE to V as well

Trap

The trap

Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.

Rotate V by position before the weighted sum

Why: V stores the content to be aggregated — rotating it by position would corrupt the values gathered by the attention pattern, injecting spurious position bias into what the model reads, not where it looks.

The fix

Apply RoPE only to Q and K — the vectors used to compute similarity scores.

Rotate Q and K; leave V untouched

Why: The positional encoding purpose is to make the similarity score between token m and token n depend on their relative offset. V carries content; the attention weights already encode position via the rotated Q·K score.

33. Which of these survive contact with Lesson 83: Positional Encoding?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
All three inject the same information — relative or absolute position — but with different extrapolation and computational properties.; For position pos in the sequence and dimension index i (0-indexed), the original Transformer (Vaswani et al., 2017) defines:; Why sinusoidal? Because the model can compute any fixed relative offset k as a linear transformation of the current PE — using the angle-addition identity.
Breaks
Store the sinusoidal PE with self.pe = pe (plain Python attribute) or with nn.Parameter(pe).; Apply the rotation to Q, K, and V — they are all projections of the same token, so position should affect all of them.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 83: Positional Encoding puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. ALiBi — attention with linear biases

Section

Part 4 of 4

35. ALiBi: linear penalty on distance

Concept

ALiBi (Press et al., 2021) skips embedding-level position encoding entirely. Instead, it adds a linear negative bias to each raw attention score proportional to distance, before softmax.

\[ \text{Attention score}_{m,n} = \frac{\mathbf{q}_m \cdot \mathbf{k}_n}{\sqrt{d}} + m_{\text{head}} \cdot (n - m) \]

The slope m_head is head-specific and fixed (not learned): m_h = 2^{-8h/H} for head h = 1, …, H. For H=4: slopes = 0.25, 0.0625, 0.0156, 0.0039.

36. Teach it back: ALiBi: linear penalty on distance

Explain it

Discussion prompt

Explain ALiBi: linear penalty on distance to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

ALiBi (Press et al., 2021) skips embedding-level position encoding entirely. Instead, it adds a linear negative bias to each raw attention score proportional to distance, before softmax.

37. Guess the shape of the answer: ALiBi bias matrix (4 heads, seq=5)

Estimation

Predict first

Compute the bias matrix for head 0 (slope=0.25) on a 5-token sequence. Rows = query positions, cols = key positions.

Commit before you compute: what does ALiBi bias matrix (4 heads, seq=5) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: slopes = [0.25, 0.0625, 0.0156, 0.0039]. bias_h0[3,1] = 0.25*(1-3) = -0.5. Nearer keys get a smaller penalty.

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Farther keys receive a larger negative adjustment, softly encouraging the model to attend to nearby tokens.

38. ALiBi bias matrix (4 heads, seq=5)

Worked example

Compute the bias matrix for head 0 (slope=0.25) on a 5-token sequence. Rows = query positions, cols = key positions.

import torch
H, seq = 4, 5
slopes = torch.tensor([2**(-8.0/H * h) for h in range(1, H+1)])
print('slopes:', [round(s,4) for s in slopes.tolist()])

# head 0 bias matrix
pos = torch.arange(seq).float()
rel = pos.unsqueeze(1) - pos.unsqueeze(0)   # key_pos - query_pos
bias_h0 = slopes[0] * rel                   # positive = closer key; causal: mask future
print(bias_h0.numpy().round(4))

slopes = [0.25, 0.0625, 0.0156, 0.0039]. bias_h0[3,1] = 0.25*(1-3) = -0.5. Nearer keys get a smaller penalty.

Why: Farther keys receive a larger negative adjustment, softly encouraging the model to attend to nearby tokens. Head 0 (steepest slope) is the most local; head 3 (flattest) is the most global.

distance |m-n|slope=0.25 (h=0)slope=0.0625 (h=1)slope=0.0156 (h=2)slope=0.0039 (h=3)
00.00.00.00.0
1-0.25-0.0625-0.0156-0.0039
2-0.50-0.1250-0.0313-0.0078
4-1.00-0.2500-0.0625-0.0156

39. Watch it run: ALiBi bias matrix (4 heads, seq=5)

Pattern

Step through it

Step through ALiBi bias matrix (4 heads, seq=5) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: distance |m-n| is 0
  2. Step 2: distance |m-n| is 1
  3. Step 3: distance |m-n| is 2
  4. Step 4: distance |m-n| is 4

40. ALiBi extrapolation advantage

Concept

Because ALiBi's bias is a simple formula (not a lookup table), it generates a well-defined penalty for any distance — including distances beyond the training length. The model was never shown position 2048 during training at length 1024, but the bias at distance 2000 is simply −slope × 2000.

schemegeneralizes past train length?extra params
sinusoidal PEyes (formula, no OOV)0
learned PEno (OOV beyond max_len)max_len × d_model
RoPEyes (rotation formula)0 (cache recomputed)
ALiBiyes (linear formula)0 (slopes fixed by head)

ALiBi also removes PE entirely from the input, slightly simplifying the architecture. The trade-off: it bakes in a locality inductive bias (nearby tokens preferred) that may not suit all tasks.

41. By analogy: ALiBi extrapolation advantage

Analogy

Discussion prompt

Explain ALiBi extrapolation advantage by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

ALiBi also removes PE entirely from the input, slightly simplifying the architecture. The trade-off: it bakes in a locality inductive bias (nearby tokens preferred) that may not suit all tasks.

42. Without one step: Positional encoding decision recipe

Constraint

Discussion prompt

Run Positional encoding decision recipe with this step confiscated:

Need relative-position geometry inside attention: use RoPE (apply to Q and K only; leave V untouched)

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Check if order matters: if the task is bag-of-words or set-based, skip PE entirely
  2. Fixed-length training: sinusoidal or learned PE work; learned PE is easy but OOV beyond max_len
  3. Need relative-position geometry inside attention: use RoPE (apply to Q and K only; leave V untouched)
  4. Need length extrapolation with minimal overhead: use ALiBi (no position vectors at all; add bias to scores)
  5. Implementation check: sinusoidal → register_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance…

43. Positional encoding decision recipe

Pattern

  1. Check if order matters: if the task is bag-of-words or set-based, skip PE entirely
  2. Fixed-length training: sinusoidal or learned PE work; learned PE is easy but OOV beyond max_len
  3. Need relative-position geometry inside attention: use RoPE (apply to Q and K only; leave V untouched)
  4. Need length extrapolation with minimal overhead: use ALiBi (no position vectors at all; add bias to scores)
  5. Implementation check: sinusoidal → register_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance added to raw scores

44. Where does it stop working: Positional encoding decision recipe

Edge cases

Discussion prompt

Positional encoding decision recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Check if order matters: if the task is bag-of-words or set-based, skip PE entirely
  2. Fixed-length training: sinusoidal or learned PE work; learned PE is easy but OOV beyond max_len
  3. Need relative-position geometry inside attention: use RoPE (apply to Q and K only; leave V untouched)
  4. Need length extrapolation with minimal overhead: use ALiBi (no position vectors at all; add bias to scores)
  5. Implementation check: sinusoidal → register_buffer; learned → nn.Embedding; RoPE → cos/sin cache + pair-wise rotation; ALiBi → fixed slope × distance…

45. Rule out three: Check yourself — permutation equivariance

Elimination

Eliminate the wrong options

A transformer encoder with no positional encoding processes the sequence [A, B, C]. You shuffle it to [C, A, B]. What is true of the output?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The output vectors are the same set, reordered to match the shuffled input
  • B. The output is completely unchanged (same order, same values)
  • C. The output differs because attention scores change with input order
  • D. Only the first output vector is affected, the others are unchanged

Survives elimination: A

Why: Self-attention is permutation-equivariant: each output vector changes because the query that produced it changed position, but the full set of output vectors is just a reordering. The attention pattern (softmax of QK^T) is symmetric in the same way the input was shuffled.

46. Check yourself — permutation equivariance

Check

Think about what changes and what doesn't when you permute the inputs.

Check your understanding

A transformer encoder with no positional encoding processes the sequence [A, B, C]. You shuffle it to [C, A, B]. What is true of the output?

  • A. The output vectors are the same set, reordered to match the shuffled input (correct)
  • B. The output is completely unchanged (same order, same values)
  • C. The output differs because attention scores change with input order
  • D. Only the first output vector is affected, the others are unchanged

Answer: A

Why: Self-attention is permutation-equivariant: each output vector changes because the query that produced it changed position, but the full set of output vectors is just a reordering. The attention pattern (softmax of QK^T) is symmetric in the same way the input was shuffled.

Why B tempts people
The output ORDER changes with the input order, even though the SET of output vectors is the same. Completely unchanged would mean permutation invariance (like a pooling operation), not equivariance.
Why C tempts people
Attention scores do change (the row ordering of the query matrix changed), but they change in exactly the same permuted pattern — the output is a consistent reordering, not a different result.
Why D tempts people
Equivariance applies uniformly: every output vector is the reordered counterpart of the original output. There is no privileged first-position effect in a permutation-equivariant operation.

47. Answer it before you see the options: Check yourself — sinusoidal PE

Prediction

Predict first

For d=4, what is PE(1, 2) — i.e., position=1, dimension index=2?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: sin(1/100) ≈ 0.01

Why: Dimension index 2 is even, so it uses sin. The pair index is i = 2/2 = 1. The denominator is 10000^(2·1/4) = 10000^0.5 = 100. So PE(1, 2) = sin(1/100) ≈ 0.0100. Verified in the d=4 trace table.

48. Check yourself — sinusoidal PE

Check

Use the formula: PE(pos, 2i) = sin(pos / 10000^(2i/d)). Work it out for the specific case before clicking.

Check your understanding

For d=4, what is PE(1, 2) — i.e., position=1, dimension index=2?

  • A. sin(1/100) ≈ 0.01 (correct)
  • B. cos(1/10) ≈ 0.9950
  • C. sin(1/10) ≈ 0.0998
  • D. cos(1/100) ≈ 1.0000

Answer: A

Why: Dimension index 2 is even, so it uses sin. The pair index is i = 2/2 = 1. The denominator is 10000^(2·1/4) = 10000^0.5 = 100. So PE(1, 2) = sin(1/100) ≈ 0.0100. Verified in the d=4 trace table.

Why B tempts people
cos(1/10) is PE(1, 3) — dimension 3 (odd → cos) with i=1 and d=4 gives denominator 10000^(2·1/4)=100, so it should be cos(1/100)≈1.000. Actually cos(1/10) is for d larger; for d=4, i=1, the denominator is 100 not 10.
Why C tempts people
sin(1/10) would require denominator=10, i.e. 10000^(2i/d)=10 → 2i/d=0.25. That requires d=8 with i=1, not d=4 with i=1.
Why D tempts people
cos(1/100) is the correct formula result for dimension 3 (odd index → cos), not dimension 2. Dimension 2 is even → sin.

49. Rule out three: Check yourself — RoPE vs ALiBi

Elimination

Eliminate the wrong options

Which statement correctly distinguishes RoPE from ALiBi?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. RoPE modifies Q and K before the dot product; ALiBi adds a fixed bias to the dot-product result (before softmax)
  • B. RoPE adds position vectors to the input embeddings; ALiBi rotates Q and K
  • C. Both RoPE and ALiBi use a learned lookup table; they differ only in the number of parameters
  • D. RoPE applies only at training time; ALiBi applies only at inference time for long sequences

Survives elimination: A

Why: RoPE rotates Q and K by position-dependent angles before computing QK^T, so the dot product itself encodes relative position. ALiBi skips all positional vectors and instead adds −slope×|i−j| to the raw QK scores before softmax. Both require zero extra parameters.

50. Check yourself — RoPE vs ALiBi

Check

One key operational difference separates RoPE from ALiBi.

Check your understanding

Which statement correctly distinguishes RoPE from ALiBi?

  • A. RoPE modifies Q and K before the dot product; ALiBi adds a fixed bias to the dot-product result (before softmax) (correct)
  • B. RoPE adds position vectors to the input embeddings; ALiBi rotates Q and K
  • C. Both RoPE and ALiBi use a learned lookup table; they differ only in the number of parameters
  • D. RoPE applies only at training time; ALiBi applies only at inference time for long sequences

Answer: A

Why: RoPE rotates Q and K by position-dependent angles before computing QK^T, so the dot product itself encodes relative position. ALiBi skips all positional vectors and instead adds −slope×|i−j| to the raw QK scores before softmax. Both require zero extra parameters.

Why B tempts people
Adding position vectors to input embeddings is what sinusoidal PE and learned PE do — not RoPE. ALiBi does not rotate anything.
Why C tempts people
Neither RoPE nor ALiBi uses a learned lookup table. Sinusoidal PE and learned PE differ there; RoPE and ALiBi both use fixed, parameter-free formulas.
Why D tempts people
Both methods apply at every forward pass (training and inference). ALiBi's length-extrapolation advantage is that its formula works for any distance, but it is applied the same way at both training and inference time.

51. Your turn: implement sinusoidal PE + RoPE

Section

Project

52. Project: positional encoding from scratch

Concept

Build and verify all three encode methods: sinusoidal PE, a simple learned PE, and RoPE. Each milestone has a concrete test you can run to confirm correctness.

#milestonecorrectness test
1Sinusoidal PE (numpy)PE[0] = [0,1,0,1,...]; verify linear-combo identity for k=2
2Learned PE (nn.Embedding)Shape (T, d); confirm no PE in model.parameters() when frozen
3RoPE (cos/sin cache + apply_rope)Confirm Q_rot shape, verify different offsets give different dot products

Build rules: type every line, use register_buffer for sinusoidal PE, and never rotate V in RoPE.

53. Break it if you can: Project: positional encoding from scratch

Counterexample

Discussion prompt

Build and verify all three encode methods: sinusoidal PE, a simple learned PE, and RoPE. Each milestone has a concrete test you can run to confirm correctness.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line, use register_buffer for sinusoidal PE, and never rotate V in RoPE.

54. Milestone 1 — sinusoidal PE, verify identity

Worked example

Your turn: implement sinusoidal PE for d=4, compute PE[3] and PE[5], then verify that PE[5,0] = PE[3,0]·cos(2) + PE[3,1]·sin(2).

Hint: denom = 10000 ** (2*i/d), even dims → sin, odd → cos. For i=0, d=4: denom=1.0. cos(2) = -0.4161, sin(2) = 0.9093.

import numpy as np, math
def sinusoidal_pe(L, d):
    pe = np.zeros((L, d))
    for pos in range(L):
        for i in range(d//2):
            w = 10000 ** (2*i/d)
            pe[pos,2*i]   = math.sin(pos/w)
            pe[pos,2*i+1] = math.cos(pos/w)
    return pe
pe = sinusoidal_pe(6, 4)
print('PE[3]:', pe[3].round(4))
print('PE[5]:', pe[5].round(4))
lhs = pe[5, 0]
rhs = pe[3,0]*math.cos(2) + pe[3,1]*math.sin(2)
print(f'identity: {lhs:.6f} == {rhs:.6f}, match={abs(lhs-rhs)<1e-9}')
quantityvalue
PE[3, 0]0.1411 (sin(3/1))
PE[3, 1]-0.9900 (cos(3/1))
PE[5, 0]-0.9589 (sin(5/1))
0.1411·cos(2) + (-0.9900)·sin(2)-0.9589 (match: True)

55. Milestone 2 — learned PE shape check

Worked example

Your turn: create an nn.Embedding(10, 8) learned PE, look up positions 0-4, and predict the weight matrix shape.

Hint: nn.Embedding(num_embeddings, embedding_dim). The weight matrix is (num_embeddings, embedding_dim). Accessed with learned_pe(torch.arange(T)).

import torch, torch.nn as nn
torch.manual_seed(42)
learned_pe = nn.Embedding(10, 8)
seq_pos = torch.arange(5)
emb = learned_pe(seq_pos)
print('lookup shape:', emb.shape)          # (5, 8)
print('weight shape:', learned_pe.weight.shape)  # (10, 8)
tensorshapenote
embedding lookup for T=5(5, 8)one 8-dim vector per position
weight matrix(10, 8)all 10 position vectors; requires grad
params count8010 × 8 = 80 learned values

56. What each one costs: Milestone 2 — learned PE shape check

Trade off

Comparison matrix

From Milestone 2 — learned PE shape check: every row here is a choice with a cost. Fill the shape column, then say which row you would actually pick and what you give up for it.

tensorshapenote
embedding lookup for T=5(5, 8)one 8-dim vector per position
weight matrix(10, 8)all 10 position vectors; requires grad
params count8010 × 8 = 80 learned values

57. Milestone 3 — RoPE full program

Worked example

Your turn: implement rope_cache and apply_rope, apply to Q and K (not V), and confirm shapes. Predict: will two pairs with the same offset always produce the same dot product?

Hint: theta = 10000^(-arange(0,d,2)/d); freqs = outer(pos, theta); rotation: [x1*cos - x2*sin, x1*sin + x2*cos]. The dot product depends on Q/K content too, not just position.

import torch
torch.manual_seed(42)
seq, d = 6, 8
def rope_cache(seq, d):
    theta = 10000.0**(-torch.arange(0,d,2).float()/d)
    freqs = torch.outer(torch.arange(seq).float(), theta)
    return torch.cos(freqs), torch.sin(freqs)
def apply_rope(x, cos, sin):
    x1, x2 = x[...,:d//2], x[...,d//2:]
    return torch.cat([x1*cos-x2*sin, x1*sin+x2*cos], dim=-1)
cos_c, sin_c = rope_cache(seq, d)
Q = torch.randn(seq, d); K = torch.randn(seq, d)
Q_r = apply_rope(Q, cos_c, sin_c)
K_r = apply_rope(K, cos_c, sin_c)
print('Q_rot:', Q_r.shape)
for m in range(2, 5):
    dot = (Q_r[m]*K_r[m-2]).sum().item()
    print(f'  Q[{m}]·K[{m-2}] (offset=2):', round(dot, 4))
pairoffsetdot product
Q[2]·K[0]2-1.1148
Q[3]·K[1]23.2626
Q[4]·K[2]2-1.5243

58. Fill in: offset for Milestone 3 — RoPE full program

Comparison

Comparison matrix

From Milestone 3 — RoPE full program: refill the offset column from what you know. The rest of the table is as it appeared.

pairoffsetdot product
Q[2]·K[0]2-1.1148
Q[3]·K[1]23.2626
Q[4]·K[2]2-1.5243

59. Show it off

Concept

Out loud, slides closed: (1) explain why a pure self-attention layer treats a shuffled sequence identically, (2) state the sinusoidal PE formula and what the denominator 10000^(2i/d) controls, (3) explain why RoPE applies only to Q and K, and (4) state the ALiBi slope formula and why it generalizes to longer sequences.

Stretch (homework): implement the sinusoidal PE as a PyTorch nn.Module with register_buffer; implement ALiBi by adding the bias tensor to raw attention logits in a toy multi-head attention block; show mathematically that PE(pos+k,2i) is linear in PE(pos,2i) and PE(pos,2i+1). Next up: Lesson 84 — transformer architecture deep dive (encoder, decoder, cross-attention, layer norm placement).

60. Connect it up: Lesson 83: Positional Encoding

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why position matters at all · Sinusoidal positional encoding · RoPE — rotate Q and K · ALiBi — attention with linear biases · Your turn: implement sinusoidal PE + RoPE. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

schemewhere appliedgeneralizes beyond train length?params
sinusoidal PEadd to inputyes0
learned PEadd to inputno (OOV)max_len × d
RoPErotate Q & Kyes0
ALiBibias QK scoresyes0

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 83 — Positional Encoding — Barron · USAAIO Round 2 Preparation, 2026
  2. Attention Is All You Need (Vaswani et al., 2017) — sinusoidal PE formula — arXiv:1706.03762
  3. RoFormer: Enhanced Transformer with Rotary Position Embedding (Su et al., 2021) — arXiv:2104.09864
  4. Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation (Press et al., 2021) — arXiv:2108.12409
  5. Sinusoidal PE values, RoPE rotation correctness, ALiBi bias matrix, all verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108