Lesson 98: Word2Vec, GloVe & Static Word Embeddings

USAAIO Lesson 98, from Phase 3. It covers the Skip-gram objective and the binary classifier that negative sampling turns it into, CBOW's mean-context prediction, and GloVe's factorization of the co-occurrence matrix under a weighted MSE. It then explains the linguistic regularity king − man + woman ≈ queen through PMI geometry, and gives the equivalence between Skip-gram with negative sampling and SPPMI factorization, following Levy and Goldberg (2014). All the objectives, losses, and trace-table values were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Word2Vec, GloVe & Static Word Embeddings

Title

USAAIO · Lesson 98 · Phase 3 (NLP)

Predict context to learn meaning. Skip-gram, negative sampling, CBOW, and GloVe — four models that gave NLP its first dense, geometric word representations. All math and code verified.

2. By the end of this lesson you can

Objectives

  1. State the Skip-gram objective: maximize P(context | center) over a sliding window
  2. Replace full softmax with negative sampling (NS) and compute the binary cross-entropy loss for one (center, context, negatives) triple
  3. Distinguish CBOW from Skip-gram: mean-context predict center; faster, slightly worse on rare words
  4. Derive the GloVe objective: weighted MSE on log co-occurrence, show it factors the PMI matrix
  5. Explain linguistic regularities (king − man + woman ≈ queen) in terms of PMI geometry

3. What survived from Subword Tokenization — BPE, WordPiece, SentencePiece?

Warm-up

Discussion prompt

Before we open Lesson 98: Word2Vec, GloVe & Static Word Embeddings: without looking back, what was the main idea of Subword Tokenization — BPE, WordPiece, SentencePiece, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

character-level vs. word-level OOV problem; BPE algorithm — iterative merge of most-frequent adjacent pairs; WordPiece likelihood-based scoring; SentencePiece language-agnostic raw-byte treatment; vocabulary size tradeoffs (seq length vs.

4. Skip-gram — predict context from center

Section

Part 1 of 4

5. The distributional hypothesis

Concept

Words that appear in the same contexts tend to have similar meanings (Harris 1954). Word2Vec operationalizes this: given a center word w, predict its surrounding context words within a window of size c.

\[ \mathcal{L} = \frac{1}{T}\sum_{t=1}^{T}\sum_{\substack{-c\le j\le c \\ j\ne 0}} \log P(w_{t+j}\mid w_t) \]

Maximizing this log-likelihood forces the model to build embeddings where similar-context words occupy nearby regions of the embedding space — the geometry encodes semantic relationships.

6. Break it if you can: The distributional hypothesis

Counterexample

Discussion prompt

Maximizing this log-likelihood forces the model to build embeddings where similar-context words occupy nearby regions of the embedding space — the geometry encodes semantic relationships.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Two embedding matrices

Concept

Skip-gram maintains two separate V × D matrices: U (center word embeddings, u_w) and W (context word embeddings, v_c). This asymmetry is intentional — a word plays different statistical roles as center vs context.

matrixshaperoleused for
U (center)V × Du_w = U[w]dot with context candidates
W (context)V × Dv_c = W[c]scored against center embed
final embedV × DU or (U+W)/2downstream tasks

After training, the center embeddings U (or the average (U+W)/2) are used as the word vectors. The context matrix W is discarded or averaged in.

8. Fill in: shape for Two embedding matrices

Comparison

Comparison matrix

From Two embedding matrices: refill the shape column from what you know. The rest of the table is as it appeared.

matrixshaperoleused for
U (center)V × Du_w = U[w]dot with context candidates
W (context)V × Dv_c = W[c]scored against center embed
final embedV × DU or (U+W)/2downstream tasks

9. Full softmax bottleneck

Concept

The natural P(context | center) requires a full softmax over all V words — typically 50,000 to 1,000,000 vocabulary entries.

\[ P(c\mid w) = \frac{\exp(v_c^\top u_w)}{\sum_{c'=1}^{V}\exp(v_{c'}^\top u_w)} \]

methoddot products per stepV=50 000
full softmaxV50 000
hierarchical softmaxlog₂V≈ 16
negative sampling (k=5)k+16

10. What each one costs: Full softmax bottleneck

Trade off

Comparison matrix

From Full softmax bottleneck: every row here is a choice with a cost. Fill the dot products per step column, then say which row you would actually pick and what you give up for it.

methoddot products per stepV=50 000
full softmaxV50 000
hierarchical softmaxlog₂V≈ 16
negative sampling (k=5)k+16

11. Negative Sampling — binary classifier

Concept

Instead of normalizing over all V words, NS trains a binary classifier: label (w, c) pairs from the corpus as positive and k randomly drawn (w, noise) pairs as negative.

\[ \mathcal{L}_{\text{NS}} = -\log\sigma(v_c^\top u_w) - \sum_{i=1}^{k}\mathbb{E}_{n_i\sim P_n}\bigl[\log\sigma(-v_{n_i}^\top u_w)\bigr] \]

Noise distribution P_n(w) ∝ freq(w)^{3/4} upweights rare words slightly. Mikolov et al. found k=5–20 works well; k=2–5 suffices for large corpora.

12. By analogy: Negative Sampling — binary classifier

Analogy

Discussion prompt

Explain Negative Sampling — binary classifier by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Instead of normalizing over all V words, NS trains a binary classifier: label (w, c) pairs from the corpus as positive and k randomly drawn (w, noise) pairs as negative.

13. Guess the shape of the answer: NS loss — one step traced

Estimation

Predict first

Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.

Commit before you compute: what does NS loss — one step traced come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Near-zero scores from random init → σ(0)=0.5 for all → loss ≈ -log(0.5)×3 ≈ 2.08; verified at 2.0634.

14. NS loss — one step traced

Worked example

Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(99)
V, D = 8, 4
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)

c_idx = torch.tensor(0)          # king
o_idx = torch.tensor(1)          # man
n_idx = torch.tensor([7, 6])     # noble, prince

u_c = U(c_idx)                   # (D,)
v_o = W(o_idx)                   # (D,)
v_n = W(n_idx)                   # (K, D)

pos_score = torch.dot(u_c, v_o)
neg_scores = v_n @ u_c           # (K,)
loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum()
print(f'pos_score={pos_score.item():.4f}')
print(f'neg_scores={[round(x,4) for x in neg_scores.tolist()]}')
print(f'loss={loss.item():.4f}')

pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634

Why: Near-zero scores from random init → σ(0)=0.5 for all → loss ≈ -log(0.5)×3 ≈ 2.08; verified at 2.0634.

quantityvaluemeaning
pos_score u_king · v_man0.0001random init → near zero dot product
σ(pos_score)0.5000the classifier says 50/50 before training
neg_scores[0] (noble)-0.0156slightly negative random dot
neg_scores[1] (prince)-0.0166slightly negative random dot
loss2.0634max entropy binary classifier — consistent with -3·log(0.5)≈2.08

15. Work backwards from the answer: NS loss — one step traced

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.

16. Something is wrong here: treating negative sampling as approximate softmax

Anomaly

Predict first

A student writes this, and it looks reasonable:

Negative sampling approximates the softmax denominator by sampling k words, so it converges to the same P(c|w) as full softmax, just faster.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: NS does NOT approximate the softmax.

NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.

Why: NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.

17. Trap: treating negative sampling as approximate softmax

Trap

The trap

Negative sampling approximates the softmax denominator by sampling k words, so it converges to the same P(c|w) as full softmax, just faster.

Conclude NS and full softmax train the same objective

Why: Wrong. NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.

The fix

NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.

NS loss = -log σ(v_c·u_w) − Σ log σ(−v_n·u_w); full softmax = -log P(c|w)

Why: Levy & Goldberg (2014) showed NS implicitly factorizes the Shifted PMI matrix: w·c = SPPMI(w,c) = max(PMI(w,c)-log k, 0). The two methods produce similar embeddings empirically but are mathematically distinct objectives.

18. Break it on purpose: treating negative sampling as approximate…

Break the constraint

Discussion prompt

The rule this trap just fixed:

NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.

19. CBOW — predict center from context

Section

Part 2 of 4

20. CBOW architecture

Concept

Continuous Bag of Words (CBOW) inverts the Skip-gram task: given all context words in the window, predict the center word. This averages out positional information.

\[ \hat{u}_w = \frac{1}{2c}\sum_{\substack{-c\le j\le c \\ j\ne 0}} v_{w_{t+j}}, \quad P(w_t\mid \text{ctx}) = \text{softmax}(U\,\hat{u}_w) \]

propertySkip-gramCBOW
taskcenter → contextcontext → center
training signal per word2c separate predictionsone prediction
rare word qualitybetter (more examples)worse (averaged away)
training speedslower (more pairs)faster
practical defaultrare-word-heavy corporalarge uniform corpora

21. Teach it back: CBOW architecture

Explain it

Discussion prompt

Explain CBOW architecture to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Continuous Bag of Words (CBOW) inverts the Skip-gram task: given all context words in the window, predict the center word. This averages out positional information.

22. Guess the shape of the answer: CBOW forward pass

Estimation

Predict first

Context words [man, woman] (indices 1, 2) predict center. D=4, V=8. Trace the mean-context vector and the logit scores.

Commit before you compute: what does CBOW forward pass come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: ctx_embeds (2,4) → mean → (4,) → logits (8,) → uniform probs before training

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Random init → near-uniform softmax over V=8 words.

23. CBOW forward pass

Worked example

Context words [man, woman] (indices 1, 2) predict center. D=4, V=8. Trace the mean-context vector and the logit scores.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
vocab=['king','man','woman','queen','royal','princess','prince','noble']
V, D = 8, 4
cbow_in  = nn.Embedding(V, D)
cbow_out = nn.Linear(D, V, bias=False)
nn.init.uniform_(cbow_in.weight, -0.5/D, 0.5/D)

ctx_words  = torch.tensor([1, 2])     # man, woman
ctx_embeds = cbow_in(ctx_words)       # (2, D)
ctx_mean   = ctx_embeds.mean(0)       # (D,)
logits     = cbow_out(ctx_mean)       # (V,)
probs      = F.softmax(logits, dim=0)
print(f'ctx_embeds: {ctx_embeds.shape}')
print(f'ctx_mean:   {ctx_mean.shape}')
print(f'logits:     {logits.shape}')
print(f'probs: {[round(p,4) for p in probs.tolist()]}')

ctx_embeds (2,4) → mean → (4,) → logits (8,) → uniform probs before training

Why: Random init → near-uniform softmax over V=8 words. Training adjusts embeddings so the mean-context vector points toward the center-word embedding.

tensorshapecontent (untrained)
ctx_embeds (man, woman)(2, 4)random ±0.125 entries
ctx_mean(4,)element-wise mean of 2 rows
logits (scores vs all words)(8,)near-zero random values
probs(8,)[0.1259, 0.1216, 0.1274, 0.1265, 0.1274, 0.1276, 0.1224, 0.1212]

24. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at CBOW forward pass. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(2, 4)
ctx_embeds (man, woman)
(4,)
ctx_mean
(8,)
logits (scores vs all words); probs
g1
shape is "(2, 4)" for ctx_embeds (man, woman) — that is what the table on "CBOW forward pass" records, and it is the single property separating this group from the rest.
g2
shape is "(4,)" for ctx_mean — that is what the table on "CBOW forward pass" records, and it is the single property separating this group from the rest.
g3
shape is "(8,)" for logits (scores vs all words), probs — that is what the table on "CBOW forward pass" records, and it is the single property separating this group from the rest.

25. GloVe — factorize co-occurrences

Section

Part 3 of 4

26. GloVe objective

Concept

GloVe (Pennington 2014) starts from the global co-occurrence matrix X (built once from the whole corpus). It minimizes a weighted squared error on log co-occurrence counts — no mini-batches, no noise samples.

\[ J = \sum_{i,j=1}^{V} f(X_{ij})\bigl(w_i^\top\tilde{w}_j + b_i + \tilde{b}_j - \log X_{ij}\bigr)^2 \]

Weighting function f(x) = min(1, (x/x_max)^α) with x_max=100, α=0.75 downweights both rare pairs (low X_ij) and very frequent pairs (cap at 1.0) — neither dominates training.

27. Co-occurrence matrix and weighting

Concept

Toy sentence [king, man, woman, queen, royal] with window=2, harmonic weighting 1/distance. Entry X[king,man]=1.0 (adjacent), X[king,woman]=0.5 (2 apart).

pair (i,j)X_ijf(X_ij) α=0.75 x_max=100weight effect
king, man1.000.0316adjacent — gets some weight
king, woman0.500.01882 apart — less weight
man, woman1.000.0316adjacent
queen, royal1.000.0316adjacent

Because x_max=100, all toy counts fall well below the cap. On real Wikipedia, high-frequency function words (the, is, of) are capped at weight 1.0 to avoid dominating the loss.

28. Guess the shape of the answer: GloVe training — loss trace

Estimation

Predict first

Train GloVe on the toy co-occurrence matrix (V=8, D=4, Adagrad, lr=0.05, 200 epochs). All 14 non-zero (i,j) pairs contribute to the loss each step.

Commit before you compute: what does GloVe training — loss trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: epoch 0: 0.6736 → epoch 50: 0.0140 → epoch 100: 0.0039 → epoch 200: 0.0003

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. GloVe converges smoothly with Adagrad.

29. GloVe training — loss trace

Worked example

Train GloVe on the toy co-occurrence matrix (V=8, D=4, Adagrad, lr=0.05, 200 epochs). All 14 non-zero (i,j) pairs contribute to the loss each step.

import torch, torch.nn as nn, numpy as np
torch.manual_seed(3)
V, D = 8, 4
w_vecs = nn.Embedding(V, D); w_ctx = nn.Embedding(V, D)
b_w = nn.Embedding(V, 1);    b_c  = nn.Embedding(V, 1)
opt = torch.optim.Adagrad(
    list(w_vecs.parameters())+list(w_ctx.parameters())+
    list(b_w.parameters())+list(b_c.parameters()), lr=0.05)

non_zeros = [(0,1,1.0),(0,2,0.5),(1,0,1.0),(1,2,1.0),(1,3,0.5),
             (2,0,0.5),(2,1,1.0),(2,3,1.0),(2,4,0.5),(3,1,0.5),
             (3,2,1.0),(3,4,1.0),(4,2,0.5),(4,3,1.0)]

def f(x, x_max=100, a=0.75):
    return float(min(1.0, (x/x_max)**a))

for epoch in range(201):
    opt.zero_grad(); loss = torch.tensor(0.0)
    for i,j,x in non_zeros:
        wt = f(x)
        wi=w_vecs(torch.tensor(i)); wj=w_ctx(torch.tensor(j))
        bi=b_w(torch.tensor(i)).squeeze(); bj=b_c(torch.tensor(j)).squeeze()
        diff=(wi@wj+bi+bj-torch.log(torch.tensor(x)))**2
        loss = loss + wt*diff
    loss.backward(); opt.step()
    if epoch in [0,50,100,200]:
        print(f'epoch {epoch:3d}: loss={loss.item():.4f}')

epoch 0: 0.6736 → epoch 50: 0.0140 → epoch 100: 0.0039 → epoch 200: 0.0003

Why: GloVe converges smoothly with Adagrad. Each step adjusts w_i·w_j + b_i + b_j to match log(X_ij). Loss reaches near-zero because the tiny vocab has only 14 non-zero pairs — underdetermined system.

epochGloVe lossnote
00.6736random init
500.0140rapid early descent (Adagrad adaptive lr)
1000.0039fine-tuning phase
2000.0003near-perfect fit on toy corpus

30. Fill in: GloVe loss for GloVe training — loss trace

Comparison

Comparison matrix

From GloVe training — loss trace: refill the GloVe loss column from what you know. The rest of the table is as it appeared.

epochGloVe lossnote
00.6736random init
500.0140rapid early descent (Adagrad adaptive lr)
1000.0039fine-tuning phase
2000.0003near-perfect fit on toy corpus

31. Something is wrong here: confusing GloVe loss with an NLL objective

Anomaly

Predict first

A student writes this, and it looks reasonable:

GloVe minimizes the same cross-entropy loss as Skip-gram — it's equivalent to maximizing log P(context|center) over all word pairs.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood.

GloVe minimizes a weighted MSE: f(X_ij) * (w_i·w_j + b_i + b_j − log X_ij)². It is a matrix-factorization objective, not a probability model.

Why: GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood. There are no conditional probabilities, no softmax, and no noise samples. The weighting function f(X_ij) distinguishes it further — rare pairs get downweighted, not discarded.

32. Trap: confusing GloVe loss with an NLL objective

Trap

The trap

GloVe minimizes the same cross-entropy loss as Skip-gram — it's equivalent to maximizing log P(context|center) over all word pairs.

Treat GloVe as another language-modeling / log-likelihood objective

Why: Wrong. GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood. There are no conditional probabilities, no softmax, and no noise samples. The weighting function f(X_ij) distinguishes it further — rare pairs get downweighted, not discarded.

The fix

GloVe minimizes a weighted MSE: f(X_ij) * (w_i·w_j + b_i + b_j − log X_ij)². It is a matrix-factorization objective, not a probability model.

GloVe objective: regression; Skip-gram NS: binary classification; full softmax: log-likelihood

Why: All three factorize something close to the PMI matrix in the limit, but their loss functions, gradients, and training dynamics are completely different. Understanding which objective you are using matters for hyperparameter choices and convergence analysis.

33. Which of these survive contact with Lesson 98: Word2Vec, GloVe & Static Word…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
After training, the center embeddings U (or the average (U+W)/2) are used as the word vectors. The context matrix W is discarded or averaged in.; The natural P(context | center) requires a full softmax over all V words — typically 50,000 to 1,000,000 vocabulary entries.; Noise distribution P_n(w) ∝ freq(w)^{3/4} upweights rare words slightly. Mikolov et al. found k=5–20 works well; k=2–5 suffices for large corpora.
Breaks
Negative sampling approximates the softmax denominator by sampling k words, so it converges to the same P(c|w) as full softmax, just faster.; GloVe minimizes the same cross-entropy loss as Skip-gram — it's equivalent to maximizing log P(context|center) over all word pairs.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 98: Word2Vec, GloVe & Static Word Embeddings puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

34. PMI geometry — why analogies work

Section

Part 4 of 4

35. PMI and what Skip-gram NS really optimizes

Concept

Levy & Goldberg (2014) proved that Skip-gram NS is implicitly factorizing the Shifted PMI (SPPMI) matrix: at the global optimum, u_w · v_c = SPPMI(w, c).

\[ \text{PMI}(w,c) = \log\frac{X_{wc}\cdot N}{X_w\,X_c}, \quad \text{SPPMI}(w,c)=\max\!\left(\text{PMI}(w,c)-\log k,\;0\right) \]

Toy example: PMI(king,man) = log(1.0 × 9.5 / (2.5 × 3.5)) = 1.0761. Because log(k) with k=5 is 1.61, SPPMI(king,man) = max(1.08 − 1.61, 0) = 0.

36. Teach it back: PMI and what Skip-gram NS really optimizes

Explain it

Discussion prompt

Explain PMI and what Skip-gram NS really optimizes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Levy & Goldberg (2014) proved that Skip-gram NS is implicitly factorizing the Shifted PMI (SPPMI) matrix: at the global optimum, u_w · v_c = SPPMI(w, c).

37. Linguistic regularities — the geometry

Concept

The famous king − man + woman ≈ queen analogy emerges because PMI encodes differential co-occurrence. u_king − u_man captures the 'royalty' direction; adding u_woman lands near u_queen.

\[ \text{PMI}(\text{king},c) - \text{PMI}(\text{man},c) \approx \text{PMI}(\text{queen},c) - \text{PMI}(\text{woman},c) \]

This holds for any context word c that distinguishes royalty from gender — so the difference vector is stable across many contexts, making vector arithmetic meaningful.

38. By analogy: Linguistic regularities — the geometry

Analogy

Discussion prompt

Explain Linguistic regularities — the geometry by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The famous king − man + woman ≈ queen analogy emerges because PMI encodes differential co-occurrence. u_king − u_man captures the 'royalty' direction; adding u_woman lands near u_queen.

39. Without one step: The static-embedding recipe

Constraint

Discussion prompt

Run The static-embedding recipe with this step confiscated:

GloVe: build X (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagrad

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Choose model: Skip-gram NS for rare-word quality; CBOW for speed on large uniform corpora; GloVe for offline co-occurrence factorization
  2. Skip-gram NS loss: -log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}
  3. CBOW: mean context embeds → softmax (or NS) over center word — same PyTorch loop, one nn.Embedding + linear head
  4. GloVe: build X (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagrad
  5. Analogy task: argmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}
  6. PMI connection: Skip-gram NS implicitly factorizes SPPMI; GloVe factorizes a similar log-ratio — both produce linear geometric structure

40. The static-embedding recipe

Pattern

  1. Choose model: Skip-gram NS for rare-word quality; CBOW for speed on large uniform corpora; GloVe for offline co-occurrence factorization
  2. Skip-gram NS loss: -log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}
  3. CBOW: mean context embeds → softmax (or NS) over center word — same PyTorch loop, one nn.Embedding + linear head
  4. GloVe: build X (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagrad
  5. Analogy task: argmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}
  6. PMI connection: Skip-gram NS implicitly factorizes SPPMI; GloVe factorizes a similar log-ratio — both produce linear geometric structure

41. Where does it stop working: The static-embedding recipe

Edge cases

Discussion prompt

The static-embedding recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Choose model: Skip-gram NS for rare-word quality; CBOW for speed on large uniform corpora; GloVe for offline co-occurrence factorization
  2. Skip-gram NS loss: -log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}
  3. CBOW: mean context embeds → softmax (or NS) over center word — same PyTorch loop, one nn.Embedding + linear head
  4. GloVe: build X (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagrad
  5. Analogy task: argmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}
  6. PMI connection: Skip-gram NS implicitly factorizes SPPMI; GloVe factorizes a similar log-ratio — both produce linear geometric structure

42. Rule out three: Check yourself — negative sampling objective

Elimination

Eliminate the wrong options

In Skip-gram with negative sampling, what does the model learn to do for a positive pair (center w, context c)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Maximize σ(v_c · u_w) — push the dot product high so the binary classifier outputs probability ≈ 1
  • B. Minimize the full softmax log-probability log P(c|w) summed over the whole vocabulary
  • C. Maximize the cosine similarity between u_w and u_c (both center embeddings)
  • D. Minimize the L2 distance between u_w and v_c

Survives elimination: A

Why: NS converts the problem to binary classification: push σ(v_c·u_w) toward 1 for real pairs and σ(v_n·u_w) toward 0 for noise pairs. The loss is -log σ(pos) − Σ log σ(−neg). No softmax normalization over V occurs — that is the entire point.

43. Check yourself — negative sampling objective

Check

Work out the mechanics before clicking.

Check your understanding

In Skip-gram with negative sampling, what does the model learn to do for a positive pair (center w, context c)?

  • A. Maximize σ(v_c · u_w) — push the dot product high so the binary classifier outputs probability ≈ 1 (correct)
  • B. Minimize the full softmax log-probability log P(c|w) summed over the whole vocabulary
  • C. Maximize the cosine similarity between u_w and u_c (both center embeddings)
  • D. Minimize the L2 distance between u_w and v_c

Answer: A

Why: NS converts the problem to binary classification: push σ(v_c·u_w) toward 1 for real pairs and σ(v_n·u_w) toward 0 for noise pairs. The loss is -log σ(pos) − Σ log σ(−neg). No softmax normalization over V occurs — that is the entire point.

Why B tempts people
That is the full-softmax objective. NS specifically avoids the V-term denominator by replacing it with k noise samples and a binary BCE loss.
Why C tempts people
NS uses the asymmetric dot product v_c·u_w between context and center matrices, not cosine similarity between two center vectors. Using the same embedding matrix for both would conflate the two statistical roles.
Why D tempts people
GloVe minimizes a squared (log-count) error; NS uses log-sigmoid (BCE), not L2 distance in embedding space.

44. Answer it before you see the options: Check yourself — Skip-gram vs CBOW

Prediction

Predict first

You are training word embeddings on a corpus with a long-tail vocabulary (many rare technical terms). Which model/configuration should you prefer?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Skip-gram with negative sampling

Why: Skip-gram generates one (center, context) training pair per co-occurrence, giving rare words proportionally more gradient updates — each occurrence of a rare word contributes 2c independent predictions. CBOW averages context embeddings, washing out the signal for rare words that appear only a few times.

45. Check yourself — Skip-gram vs CBOW

Check

Apply the rare-word reasoning.

Check your understanding

You are training word embeddings on a corpus with a long-tail vocabulary (many rare technical terms). Which model/configuration should you prefer?

  • A. Skip-gram with negative sampling (correct)
  • B. CBOW with large window size
  • C. GloVe with x_max = 1 (no weight cap)
  • D. CBOW with window size 1

Answer: A

Why: Skip-gram generates one (center, context) training pair per co-occurrence, giving rare words proportionally more gradient updates — each occurrence of a rare word contributes 2c independent predictions. CBOW averages context embeddings, washing out the signal for rare words that appear only a few times.

Why B tempts people
CBOW averages all context vectors, which dilutes the signal for rare center words regardless of window size. A larger window makes this averaging worse for rare terms, not better.
Why C tempts people
Setting x_max=1 means f(X_ij) = min(1,(X/1)^0.75) ≥ 1 for all X≥1, removing the soft cap — high-frequency pairs dominate and rare pairs are relatively underweighted, the opposite of what you want.
Why D tempts people
CBOW with window=1 reduces the context to just one neighbor per side, giving fewer training signals per word and hurting both common and rare words.

46. Rule out three: Check yourself — GloVe objective

Elimination

Eliminate the wrong options

In the GloVe objective J = Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)², what is the purpose of f(X_ij) = min(1, (X_ij/x_max)^α)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Downweight very rare pairs (low X_ij) and cap the weight of very frequent pairs (X_ij ≥ x_max) so neither extreme dominates
  • B. Convert X_ij into a probability distribution so the objective becomes equivalent to cross-entropy
  • C. Apply L2 regularization to the word vectors w_i
  • D. Replace log X_ij with a smooth approximation that is defined even when X_ij = 0

Survives elimination: A

Why: f(X_ij) is a soft weight in [0,1]. For small X_ij it is small (rare co-occurrences get low weight, reducing noise). For X_ij ≥ x_max it is capped at 1 (common words like 'the' are not over-represented). The weighting is neither normalization nor regularization — it is an explicit importance weight on the squared residual.

47. Check yourself — GloVe objective

Check

Identify the role of the weighting function.

Check your understanding

In the GloVe objective J = Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)², what is the purpose of f(X_ij) = min(1, (X_ij/x_max)^α)?

  • A. Downweight very rare pairs (low X_ij) and cap the weight of very frequent pairs (X_ij ≥ x_max) so neither extreme dominates (correct)
  • B. Convert X_ij into a probability distribution so the objective becomes equivalent to cross-entropy
  • C. Apply L2 regularization to the word vectors w_i
  • D. Replace log X_ij with a smooth approximation that is defined even when X_ij = 0

Answer: A

Why: f(X_ij) is a soft weight in [0,1]. For small X_ij it is small (rare co-occurrences get low weight, reducing noise). For X_ij ≥ x_max it is capped at 1 (common words like 'the' are not over-represented). The weighting is neither normalization nor regularization — it is an explicit importance weight on the squared residual.

Why B tempts people
Dividing by f(X_ij) would not produce a valid probability. The objective remains a weighted squared loss — there is no softmax normalization anywhere in GloVe.
Why C tempts people
L2 regularization adds λ‖w‖² to the loss; f(X_ij) is a per-pair importance weight that multiplies the squared residual. These are structurally different terms.
Why D tempts people
GloVe only sums over pairs where X_ij > 0, so log X_ij is always defined. The function f handles importance weighting, not numerical stability at zero.

48. Your turn: Word2Vec from scratch

Section

Project

49. Project: Skip-gram with negative sampling

Concept

Build Skip-gram NS from scratch — two embedding matrices, a binary loss, a training loop — and watch the embeddings shift so that king-man+woman moves closer to queen.

#milestonekey tool
1Implement NS loss for one (center, context, negatives) triple; trace shapesnn.Embedding, torch.dot, F.logsigmoid
2Build the full training loop over skip-gram pairs; log loss at epoch 0/100/200SGD, manual_seed per epoch
3Compute king−man+woman analogy vector; print cosine sims to all vocab wordsF.cosine_similarity

Build rules: use torch.manual_seed(0) for reproducibility; verify u_c shape = (D,), v_n shape = (K,D), loss > 0 before training. Do NOT use a pre-trained embedding — the point is to watch it learn from scratch.

50. Break it if you can: Project: Skip-gram with negative sampling

Counterexample

Discussion prompt

Build Skip-gram NS from scratch — two embedding matrices, a binary loss, a training loop — and watch the embeddings shift so that king-man+woman moves closer to queen.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

51. Milestone 1 — NS loss one step

Worked example

Your turn: implement the NS loss for (king, man, [noble, prince]). Predict: is the loss positive or negative before training?

Hint: loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum(). Random init → scores ≈ 0 → σ(0) = 0.5 → loss ≈ -log(0.5) × (K+1) ≈ 2.08.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
V, D, K = 8, 4, 3
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)

c_t = torch.tensor(0)               # king
o_t = torch.tensor(1)               # man
n_t = torch.tensor([7, 6, 5])       # noble, prince, princess

u_c = U(c_t)         # (D,)
v_o = W(o_t)         # (D,)
v_n = W(n_t)         # (K, D)

pos_score  = torch.dot(u_c, v_o)
neg_scores = v_n @ u_c              # (K,)
loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum()
print(f'u_c shape: {u_c.shape}, v_n shape: {v_n.shape}')
print(f'pos_score: {pos_score.item():.4f}')
print(f'loss: {loss.item():.4f}')
quantityshapevalue (seed=42)
u_c (king center embed)(4,)4-dim random vector
v_o (man context embed)(4,)4-dim random vector
v_n (3 neg context embeds)(3, 4)3 rows of 4-dim random vecs
pos_score u_c·v_oscalar0.0015
NS lossscalar2.7777

52. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at Milestone 1 — NS loss one step. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(4,)
u_c (king center embed); v_o (man context embed)
(3, 4)
v_n (3 neg context embeds)
scalar
pos_score u_c·v_o; NS loss
g1
shape is "(4,)" for u_c (king center embed), v_o (man context embed) — that is what the table on "Milestone 1 — NS loss one step" records, and it is the single property separating this group from the rest.
g2
shape is "(3, 4)" for v_n (3 neg context embeds) — that is what the table on "Milestone 1 — NS loss one step" records, and it is the single property separating this group from the rest.
g3
shape is "scalar" for pos_score u_c·v_o, NS loss — that is what the table on "Milestone 1 — NS loss one step" records, and it is the single property separating this group from the rest.

53. Milestone 2 — training loop

Worked example

Your turn: loop 200 epochs over the 8 skip-gram pairs (window=1 on [king,man,woman,queen,royal]). Predict: does loss go up or down? How fast?

Hint: set torch.manual_seed(epoch) inside the epoch loop to get reproducible noise samples. SGD lr=0.05 works; Adam lr=1e-3 converges faster.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
V, D, K = 8, 4, 3
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
opt = torch.optim.SGD(list(U.parameters())+list(W.parameters()), lr=0.05)
pairs = [(0,1),(1,0),(1,2),(2,1),(2,3),(3,2),(3,4),(4,3)]
pair_t = [(torch.tensor(c),torch.tensor(o)) for c,o in pairs]
for epoch in range(201):
    torch.manual_seed(epoch)
    total = 0.0
    for c_t, o_t in pair_t:
        opt.zero_grad()
        noise = torch.randint(0, V, (K,))
        loss = (-F.logsigmoid(torch.dot(U(c_t), W(o_t)))
               -F.logsigmoid(-(W(noise)@U(c_t))).sum())
        loss.backward(); opt.step(); total += loss.item()
    if epoch in [0,10,50,100,200]:
        print(f'epoch {epoch:3d}: {total:.4f}')
epochtotal loss (8 pairs)
022.1619
1022.1292
5016.1344
10010.7495
20010.8946

54. Watch it run: Milestone 2 — training loop

Pattern

Step through it

Step through Milestone 2 — training loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 10
  3. Step 3: epoch is 50
  4. Step 4: epoch is 100
  5. Step 5: epoch is 200

55. Milestone 3 — analogy evaluation

Worked example

Your turn: compute the king − man + woman analogy vector and rank all 8 vocab words by cosine similarity. Predict which word is closest (the toy corpus is tiny — exact correctness is not guaranteed).

Hint: analogy = E[w2i['king']] - E[w2i['man']] + E[w2i['woman']]; then F.cosine_similarity(analogy.unsqueeze(0), E) gives scores for all V words.

# continues from Milestone 2 (same U, trained 200 epochs)
vocab = ['king','man','woman','queen','royal','princess','prince','noble']
E = U.weight.detach()             # (V, D)
analogy = E[0] - E[1] + E[2]     # king - man + woman
sims = F.cosine_similarity(analogy.unsqueeze(0), E)  # (V,)
ranked = sorted(zip(sims.tolist(), vocab), reverse=True)
for s, w in ranked:
    print(f'  {w:10s} {s:.4f}')
rankwordcosine simnote
1 (excluded)king—query word, excluded in real eval
2prince0.7066toy corpus; shares 'royal' context
3royal0.6626appears next to queen and king
4queen-0.0272tiny corpus — analogy geometry emerges at scale

56. What each one costs: Milestone 3 — analogy evaluation

Trade off

Comparison matrix

From Milestone 3 — analogy evaluation: every row here is a choice with a cost. Fill the note column, then say which row you would actually pick and what you give up for it.

rankwordcosine simnote
1 (excluded)king—query word, excluded in real eval
2prince0.7066toy corpus; shares 'royal' context
3royal0.6626appears next to queen and king
4queen-0.0272tiny corpus — analogy geometry emerges at scale

57. The full program

Concept

import torch, torch.nn as nn, torch.nn.functional as F

vocab  = ['king','man','woman','queen','royal','princess','prince','noble']
w2i    = {w:i for i,w in enumerate(vocab)}
V, D, K = 8, 4, 3

# --- Skip-gram pairs (window=1) ---
sentence = [0,1,2,3,4]
pairs = [(sentence[p],sentence[p+o])
         for p in range(len(sentence))
         for o in [-1,1] if 0<=p+o<len(sentence)]
pair_t = [(torch.tensor(c),torch.tensor(o)) for c,o in pairs]

# --- Model ---
torch.manual_seed(0)
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
opt = torch.optim.SGD(list(U.parameters())+list(W.parameters()), lr=0.05)

# --- Training ---
for epoch in range(201):
    torch.manual_seed(epoch); total=0.0
    for c_t,o_t in pair_t:
        opt.zero_grad()
        noise = torch.randint(0,V,(K,))
        loss  = (-F.logsigmoid(torch.dot(U(c_t),W(o_t)))
                 -F.logsigmoid(-(W(noise)@U(c_t))).sum())
        loss.backward(); opt.step(); total+=loss.item()
    if epoch in [0,100,200]: print(f'epoch {epoch}: {total:.4f}')

# --- Analogy ---
E = U.weight.detach()
analogy = E[w2i['king']]-E[w2i['man']]+E[w2i['woman']]
sims = F.cosine_similarity(analogy.unsqueeze(0), E)
print({vocab[i]:round(sims[i].item(),4) for i in range(V)})
output linevalue
epoch 0 loss22.1619
epoch 100 loss10.7495
epoch 200 loss10.8946
sim(queen)-0.0272 (analogy needs scale to emerge)
sim(prince)0.7066 (shares royal context)

At toy scale the analogy is noisy — the geometry requires hundreds of millions of training tokens to become reliable. What you've built is the exact mechanism Word2Vec uses; the only difference is corpus size.

58. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

output linevalue
epoch 0 loss22.1619
epoch 100 loss10.7495
epoch 200 loss10.8946
sim(queen)-0.0272 (analogy needs scale to emerge)
sim(prince)0.7066 (shares royal context)

59. Show it off

Concept

Out loud, slides closed: (1) explain the NS loss — what is maximized, what is minimized, why no softmax; (2) contrast Skip-gram and CBOW on rare words; (3) state the GloVe objective and name both terms that differ from Skip-gram NS.

Homework (lesson plan): implement on real data — train on the first 1M tokens of Wikipedia text (download from Hugging Face datasets), evaluate on Google's word-analogy benchmark (8,869 semantic + 10,675 syntactic pairs), visualize 2D t-SNE clusters. Next up: Lesson 99 — contextual embeddings (ELMo, BERT contextualized representations).

60. Connect it up: Lesson 98: Word2Vec, GloVe & Static Word Embeddings

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Skip-gram — predict context from center · CBOW — predict center from context · GloVe — factorize co-occurrences · PMI geometry — why analogies work · Your turn: Word2Vec from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

modelobjectivekey hyperparameterbest for
Skip-gram NSbinary BCE (SPPMI)k=5–20 negativesrare words, smaller corpora
CBOWsoftmax or NS (mean ctx)window size cspeed, large corpora
GloVeweighted MSE on log X_ijx_max=100, α=0.75offline matrix factorization
analogyargmax cosine(king−man+woman, ·)exclude query wordsintrinsic embedding eval

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 98 — Word2Vec Skip-gram, GloVe — Barron · USAAIO Round 2 Preparation, 2026
  2. Mikolov et al. 'Distributed Representations of Words and Phrases and their Compositionality' (NeurIPS 2013) — arXiv:1310.4546
  3. Pennington, Socher, Manning 'GloVe: Global Vectors for Word Representation' (EMNLP 2014) — https://nlp.stanford.edu/pubs/glove.pdf
  4. Levy & Goldberg 'Neural Word Embedding as Implicit Matrix Factorization' (NeurIPS 2014) — arXiv:1407.4543
  5. All objectives, loss values, trace tables, and embedding shapes verified with torch 2.7.1+cpu, numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108