USAAIO Lesson 98, from Phase 3. It covers the Skip-gram objective and the binary classifier that negative sampling turns it into, CBOW's mean-context prediction, and GloVe's factorization of the co-occurrence matrix under a weighted MSE. It then explains the linguistic regularity king − man + woman ≈ queen through PMI geometry, and gives the equivalence between Skip-gram with negative sampling and SPPMI factorization, following Levy and Goldberg (2014). All the objectives, losses, and trace-table values were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 32 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 98 · Phase 3 (NLP)
Predict context to learn meaning. Skip-gram, negative sampling, CBOW, and GloVe — four models that gave NLP its first dense, geometric word representations. All math and code verified.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 98: Word2Vec, GloVe & Static Word Embeddings: without looking back, what was the main idea of Subword Tokenization — BPE, WordPiece, SentencePiece, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
character-level vs. word-level OOV problem; BPE algorithm — iterative merge of most-frequent adjacent pairs; WordPiece likelihood-based scoring; SentencePiece language-agnostic raw-byte treatment; vocabulary size tradeoffs (seq length vs.
Section
Part 1 of 4
Concept
Words that appear in the same contexts tend to have similar meanings (Harris 1954). Word2Vec operationalizes this: given a center word w, predict its surrounding context words within a window of size c.
\[ \mathcal{L} = \frac{1}{T}\sum_{t=1}^{T}\sum_{\substack{-c\le j\le c \\ j\ne 0}} \log P(w_{t+j}\mid w_t) \]
Maximizing this log-likelihood forces the model to build embeddings where similar-context words occupy nearby regions of the embedding space — the geometry encodes semantic relationships.
Counterexample
Discussion prompt
Maximizing this log-likelihood forces the model to build embeddings where similar-context words occupy nearby regions of the embedding space — the geometry encodes semantic relationships.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Skip-gram maintains two separate V × D matrices: U (center word embeddings, u_w) and W (context word embeddings, v_c). This asymmetry is intentional — a word plays different statistical roles as center vs context.
| matrix | shape | role | used for |
|---|---|---|---|
| U (center) | V × D | u_w = U[w] | dot with context candidates |
| W (context) | V × D | v_c = W[c] | scored against center embed |
| final embed | V × D | U or (U+W)/2 | downstream tasks |
After training, the center embeddings U (or the average (U+W)/2) are used as the word vectors. The context matrix W is discarded or averaged in.
Comparison
Comparison matrix
From Two embedding matrices: refill the shape column from what you know. The rest of the table is as it appeared.
| matrix | shape | role | used for |
|---|---|---|---|
| U (center) | V × D | u_w = U[w] | dot with context candidates |
| W (context) | V × D | v_c = W[c] | scored against center embed |
| final embed | V × D | U or (U+W)/2 | downstream tasks |
Concept
The natural P(context | center) requires a full softmax over all V words — typically 50,000 to 1,000,000 vocabulary entries.
\[ P(c\mid w) = \frac{\exp(v_c^\top u_w)}{\sum_{c'=1}^{V}\exp(v_{c'}^\top u_w)} \]
| method | dot products per step | V=50 000 |
|---|---|---|
| full softmax | V | 50 000 |
| hierarchical softmax | log₂V | ≈ 16 |
| negative sampling (k=5) | k+1 | 6 |
Trade off
Comparison matrix
From Full softmax bottleneck: every row here is a choice with a cost. Fill the dot products per step column, then say which row you would actually pick and what you give up for it.
| method | dot products per step | V=50 000 |
|---|---|---|
| full softmax | V | 50 000 |
| hierarchical softmax | log₂V | ≈ 16 |
| negative sampling (k=5) | k+1 | 6 |
Concept
Instead of normalizing over all V words, NS trains a binary classifier: label (w, c) pairs from the corpus as positive and k randomly drawn (w, noise) pairs as negative.
\[ \mathcal{L}_{\text{NS}} = -\log\sigma(v_c^\top u_w) - \sum_{i=1}^{k}\mathbb{E}_{n_i\sim P_n}\bigl[\log\sigma(-v_{n_i}^\top u_w)\bigr] \]
Noise distribution P_n(w) ∝ freq(w)^{3/4} upweights rare words slightly. Mikolov et al. found k=5–20 works well; k=2–5 suffices for large corpora.
Analogy
Discussion prompt
Explain Negative Sampling — binary classifier by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Instead of normalizing over all V words, NS trains a binary classifier: label (w, c) pairs from the corpus as positive and k randomly drawn (w, noise) pairs as negative.
Estimation
Predict first
Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.
Commit before you compute: what does NS loss — one step traced come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Near-zero scores from random init → σ(0)=0.5 for all → loss ≈ -log(0.5)×3 ≈ 2.08; verified at 2.0634.
Worked example
Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(99)
V, D = 8, 4
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
c_idx = torch.tensor(0) # king
o_idx = torch.tensor(1) # man
n_idx = torch.tensor([7, 6]) # noble, prince
u_c = U(c_idx) # (D,)
v_o = W(o_idx) # (D,)
v_n = W(n_idx) # (K, D)
pos_score = torch.dot(u_c, v_o)
neg_scores = v_n @ u_c # (K,)
loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum()
print(f'pos_score={pos_score.item():.4f}')
print(f'neg_scores={[round(x,4) for x in neg_scores.tolist()]}')
print(f'loss={loss.item():.4f}')pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634
Why: Near-zero scores from random init → σ(0)=0.5 for all → loss ≈ -log(0.5)×3 ≈ 2.08; verified at 2.0634.
| quantity | value | meaning |
|---|---|---|
| pos_score u_king · v_man | 0.0001 | random init → near zero dot product |
| σ(pos_score) | 0.5000 | the classifier says 50/50 before training |
| neg_scores[0] (noble) | -0.0156 | slightly negative random dot |
| neg_scores[1] (prince) | -0.0166 | slightly negative random dot |
| loss | 2.0634 | max entropy binary classifier — consistent with -3·log(0.5)≈2.08 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
pos_score=0.0001 neg_scores=[-0.0156, -0.0166] loss=2.0634
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Trace one NS step: center=king, context=man, negatives=[noble, prince] with D=4, random init seed 99. Predict the sign of the loss before running.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Negative sampling approximates the softmax denominator by sampling k words, so it converges to the same P(c|w) as full softmax, just faster.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: NS does NOT approximate the softmax.
NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.
Why: NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.
Trap
Negative sampling approximates the softmax denominator by sampling k words, so it converges to the same P(c|w) as full softmax, just faster.
Conclude NS and full softmax train the same objective
Why: Wrong. NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.
NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.
NS loss = -log σ(v_c·u_w) − Σ log σ(−v_n·u_w); full softmax = -log P(c|w)
Why: Levy & Goldberg (2014) showed NS implicitly factorizes the Shifted PMI matrix: w·c = SPPMI(w,c) = max(PMI(w,c)-log k, 0). The two methods produce similar embeddings empirically but are mathematically distinct objectives.
Break the constraint
Discussion prompt
The rule this trap just fixed:
NS is a different objective — binary classification of real vs noise pairs. It is not an approximation of the softmax normalization.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
NS does NOT approximate the softmax. It is a completely different objective — a binary cross-entropy for (word, context) classification. The resulting embeddings optimize a different loss function and implicitly factorize the SPPMI matrix, not the log-probability matrix.
Section
Part 2 of 4
Concept
Continuous Bag of Words (CBOW) inverts the Skip-gram task: given all context words in the window, predict the center word. This averages out positional information.
\[ \hat{u}_w = \frac{1}{2c}\sum_{\substack{-c\le j\le c \\ j\ne 0}} v_{w_{t+j}}, \quad P(w_t\mid \text{ctx}) = \text{softmax}(U\,\hat{u}_w) \]
| property | Skip-gram | CBOW |
|---|---|---|
| task | center → context | context → center |
| training signal per word | 2c separate predictions | one prediction |
| rare word quality | better (more examples) | worse (averaged away) |
| training speed | slower (more pairs) | faster |
| practical default | rare-word-heavy corpora | large uniform corpora |
Explain it
Discussion prompt
Explain CBOW architecture to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Continuous Bag of Words (CBOW) inverts the Skip-gram task: given all context words in the window, predict the center word. This averages out positional information.
Estimation
Predict first
Context words [man, woman] (indices 1, 2) predict center. D=4, V=8. Trace the mean-context vector and the logit scores.
Commit before you compute: what does CBOW forward pass come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: ctx_embeds (2,4) → mean → (4,) → logits (8,) → uniform probs before training
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Random init → near-uniform softmax over V=8 words.
Worked example
Context words [man, woman] (indices 1, 2) predict center. D=4, V=8. Trace the mean-context vector and the logit scores.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(7)
vocab=['king','man','woman','queen','royal','princess','prince','noble']
V, D = 8, 4
cbow_in = nn.Embedding(V, D)
cbow_out = nn.Linear(D, V, bias=False)
nn.init.uniform_(cbow_in.weight, -0.5/D, 0.5/D)
ctx_words = torch.tensor([1, 2]) # man, woman
ctx_embeds = cbow_in(ctx_words) # (2, D)
ctx_mean = ctx_embeds.mean(0) # (D,)
logits = cbow_out(ctx_mean) # (V,)
probs = F.softmax(logits, dim=0)
print(f'ctx_embeds: {ctx_embeds.shape}')
print(f'ctx_mean: {ctx_mean.shape}')
print(f'logits: {logits.shape}')
print(f'probs: {[round(p,4) for p in probs.tolist()]}')ctx_embeds (2,4) → mean → (4,) → logits (8,) → uniform probs before training
Why: Random init → near-uniform softmax over V=8 words. Training adjusts embeddings so the mean-context vector points toward the center-word embedding.
| tensor | shape | content (untrained) |
|---|---|---|
| ctx_embeds (man, woman) | (2, 4) | random ±0.125 entries |
| ctx_mean | (4,) | element-wise mean of 2 rows |
| logits (scores vs all words) | (8,) | near-zero random values |
| probs | (8,) | [0.1259, 0.1216, 0.1274, 0.1265, 0.1274, 0.1276, 0.1224, 0.1212] |
Discrimination
Sort into buckets
Sort these by shape, from memory, without looking back at CBOW forward pass. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 3 of 4
Concept
GloVe (Pennington 2014) starts from the global co-occurrence matrix X (built once from the whole corpus). It minimizes a weighted squared error on log co-occurrence counts — no mini-batches, no noise samples.
\[ J = \sum_{i,j=1}^{V} f(X_{ij})\bigl(w_i^\top\tilde{w}_j + b_i + \tilde{b}_j - \log X_{ij}\bigr)^2 \]
Weighting function f(x) = min(1, (x/x_max)^α) with x_max=100, α=0.75 downweights both rare pairs (low X_ij) and very frequent pairs (cap at 1.0) — neither dominates training.
Concept
Toy sentence [king, man, woman, queen, royal] with window=2, harmonic weighting 1/distance. Entry X[king,man]=1.0 (adjacent), X[king,woman]=0.5 (2 apart).
| pair (i,j) | X_ij | f(X_ij) α=0.75 x_max=100 | weight effect |
|---|---|---|---|
| king, man | 1.00 | 0.0316 | adjacent — gets some weight |
| king, woman | 0.50 | 0.0188 | 2 apart — less weight |
| man, woman | 1.00 | 0.0316 | adjacent |
| queen, royal | 1.00 | 0.0316 | adjacent |
Because x_max=100, all toy counts fall well below the cap. On real Wikipedia, high-frequency function words (the, is, of) are capped at weight 1.0 to avoid dominating the loss.
Estimation
Predict first
Train GloVe on the toy co-occurrence matrix (V=8, D=4, Adagrad, lr=0.05, 200 epochs). All 14 non-zero (i,j) pairs contribute to the loss each step.
Commit before you compute: what does GloVe training — loss trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: epoch 0: 0.6736 → epoch 50: 0.0140 → epoch 100: 0.0039 → epoch 200: 0.0003
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. GloVe converges smoothly with Adagrad.
Worked example
Train GloVe on the toy co-occurrence matrix (V=8, D=4, Adagrad, lr=0.05, 200 epochs). All 14 non-zero (i,j) pairs contribute to the loss each step.
import torch, torch.nn as nn, numpy as np
torch.manual_seed(3)
V, D = 8, 4
w_vecs = nn.Embedding(V, D); w_ctx = nn.Embedding(V, D)
b_w = nn.Embedding(V, 1); b_c = nn.Embedding(V, 1)
opt = torch.optim.Adagrad(
list(w_vecs.parameters())+list(w_ctx.parameters())+
list(b_w.parameters())+list(b_c.parameters()), lr=0.05)
non_zeros = [(0,1,1.0),(0,2,0.5),(1,0,1.0),(1,2,1.0),(1,3,0.5),
(2,0,0.5),(2,1,1.0),(2,3,1.0),(2,4,0.5),(3,1,0.5),
(3,2,1.0),(3,4,1.0),(4,2,0.5),(4,3,1.0)]
def f(x, x_max=100, a=0.75):
return float(min(1.0, (x/x_max)**a))
for epoch in range(201):
opt.zero_grad(); loss = torch.tensor(0.0)
for i,j,x in non_zeros:
wt = f(x)
wi=w_vecs(torch.tensor(i)); wj=w_ctx(torch.tensor(j))
bi=b_w(torch.tensor(i)).squeeze(); bj=b_c(torch.tensor(j)).squeeze()
diff=(wi@wj+bi+bj-torch.log(torch.tensor(x)))**2
loss = loss + wt*diff
loss.backward(); opt.step()
if epoch in [0,50,100,200]:
print(f'epoch {epoch:3d}: loss={loss.item():.4f}')epoch 0: 0.6736 → epoch 50: 0.0140 → epoch 100: 0.0039 → epoch 200: 0.0003
Why: GloVe converges smoothly with Adagrad. Each step adjusts w_i·w_j + b_i + b_j to match log(X_ij). Loss reaches near-zero because the tiny vocab has only 14 non-zero pairs — underdetermined system.
| epoch | GloVe loss | note |
|---|---|---|
| 0 | 0.6736 | random init |
| 50 | 0.0140 | rapid early descent (Adagrad adaptive lr) |
| 100 | 0.0039 | fine-tuning phase |
| 200 | 0.0003 | near-perfect fit on toy corpus |
Comparison
Comparison matrix
From GloVe training — loss trace: refill the GloVe loss column from what you know. The rest of the table is as it appeared.
| epoch | GloVe loss | note |
|---|---|---|
| 0 | 0.6736 | random init |
| 50 | 0.0140 | rapid early descent (Adagrad adaptive lr) |
| 100 | 0.0039 | fine-tuning phase |
| 200 | 0.0003 | near-perfect fit on toy corpus |
Anomaly
Predict first
A student writes this, and it looks reasonable:
GloVe minimizes the same cross-entropy loss as Skip-gram — it's equivalent to maximizing log P(context|center) over all word pairs.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood.
GloVe minimizes a weighted MSE: f(X_ij) * (w_i·w_j + b_i + b_j − log X_ij)². It is a matrix-factorization objective, not a probability model.
Why: GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood. There are no conditional probabilities, no softmax, and no noise samples. The weighting function f(X_ij) distinguishes it further — rare pairs get downweighted, not discarded.
Trap
GloVe minimizes the same cross-entropy loss as Skip-gram — it's equivalent to maximizing log P(context|center) over all word pairs.
Treat GloVe as another language-modeling / log-likelihood objective
Why: Wrong. GloVe is a weighted least-squares regression objective on log co-occurrence, not a log-likelihood. There are no conditional probabilities, no softmax, and no noise samples. The weighting function f(X_ij) distinguishes it further — rare pairs get downweighted, not discarded.
GloVe minimizes a weighted MSE: f(X_ij) * (w_i·w_j + b_i + b_j − log X_ij)². It is a matrix-factorization objective, not a probability model.
GloVe objective: regression; Skip-gram NS: binary classification; full softmax: log-likelihood
Why: All three factorize something close to the PMI matrix in the limit, but their loss functions, gradients, and training dynamics are completely different. Understanding which objective you are using matters for hyperparameter choices and convergence analysis.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
U (or the average (U+W)/2) are used as the word vectors. The context matrix W is discarded or averaged in.; The natural P(context | center) requires a full softmax over all V words — typically 50,000 to 1,000,000 vocabulary entries.; Noise distribution P_n(w) ∝ freq(w)^{3/4} upweights rare words slightly. Mikolov et al. found k=5–20 works well; k=2–5 suffices for large corpora.Section
Part 4 of 4
Concept
Levy & Goldberg (2014) proved that Skip-gram NS is implicitly factorizing the Shifted PMI (SPPMI) matrix: at the global optimum, u_w · v_c = SPPMI(w, c).
\[ \text{PMI}(w,c) = \log\frac{X_{wc}\cdot N}{X_w\,X_c}, \quad \text{SPPMI}(w,c)=\max\!\left(\text{PMI}(w,c)-\log k,\;0\right) \]
Toy example: PMI(king,man) = log(1.0 × 9.5 / (2.5 × 3.5)) = 1.0761. Because log(k) with k=5 is 1.61, SPPMI(king,man) = max(1.08 − 1.61, 0) = 0.
Explain it
Discussion prompt
Explain PMI and what Skip-gram NS really optimizes to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Levy & Goldberg (2014) proved that Skip-gram NS is implicitly factorizing the Shifted PMI (SPPMI) matrix: at the global optimum, u_w · v_c = SPPMI(w, c).
Concept
The famous king − man + woman ≈ queen analogy emerges because PMI encodes differential co-occurrence. u_king − u_man captures the 'royalty' direction; adding u_woman lands near u_queen.
\[ \text{PMI}(\text{king},c) - \text{PMI}(\text{man},c) \approx \text{PMI}(\text{queen},c) - \text{PMI}(\text{woman},c) \]
This holds for any context word c that distinguishes royalty from gender — so the difference vector is stable across many contexts, making vector arithmetic meaningful.
Analogy
Discussion prompt
Explain Linguistic regularities — the geometry by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The famous king − man + woman ≈ queen analogy emerges because PMI encodes differential co-occurrence. u_king − u_man captures the 'royalty' direction; adding u_woman lands near u_queen.
Constraint
Discussion prompt
Run The static-embedding recipe with this step confiscated:
GloVe: build X (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagrad
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
-log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}nn.Embedding + linear headX (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagradargmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}Pattern
-log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}nn.Embedding + linear headX (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagradargmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}Edge cases
Discussion prompt
The static-embedding recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
-log σ(v_c·u_w) − Σ log σ(−v_n·u_w) with k=5–20 negative samples, noise dist P_n ∝ freq^{3/4}nn.Embedding + linear headX (window co-occurrence), minimize Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)² with Adagradargmax_v cos(u_king − u_man + u_woman, v) excluding {king, man, woman}Elimination
Eliminate the wrong options
In Skip-gram with negative sampling, what does the model learn to do for a positive pair (center w, context c)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: NS converts the problem to binary classification: push σ(v_c·u_w) toward 1 for real pairs and σ(v_n·u_w) toward 0 for noise pairs. The loss is -log σ(pos) − Σ log σ(−neg). No softmax normalization over V occurs — that is the entire point.
Check
Work out the mechanics before clicking.
Check your understanding
In Skip-gram with negative sampling, what does the model learn to do for a positive pair (center w, context c)?
Answer: A
Why: NS converts the problem to binary classification: push σ(v_c·u_w) toward 1 for real pairs and σ(v_n·u_w) toward 0 for noise pairs. The loss is -log σ(pos) − Σ log σ(−neg). No softmax normalization over V occurs — that is the entire point.
Prediction
Predict first
You are training word embeddings on a corpus with a long-tail vocabulary (many rare technical terms). Which model/configuration should you prefer?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Skip-gram with negative sampling
Why: Skip-gram generates one (center, context) training pair per co-occurrence, giving rare words proportionally more gradient updates — each occurrence of a rare word contributes 2c independent predictions. CBOW averages context embeddings, washing out the signal for rare words that appear only a few times.
Check
Apply the rare-word reasoning.
Check your understanding
You are training word embeddings on a corpus with a long-tail vocabulary (many rare technical terms). Which model/configuration should you prefer?
Answer: A
Why: Skip-gram generates one (center, context) training pair per co-occurrence, giving rare words proportionally more gradient updates — each occurrence of a rare word contributes 2c independent predictions. CBOW averages context embeddings, washing out the signal for rare words that appear only a few times.
Elimination
Eliminate the wrong options
In the GloVe objective J = Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)², what is the purpose of f(X_ij) = min(1, (X_ij/x_max)^α)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: f(X_ij) is a soft weight in [0,1]. For small X_ij it is small (rare co-occurrences get low weight, reducing noise). For X_ij ≥ x_max it is capped at 1 (common words like 'the' are not over-represented). The weighting is neither normalization nor regularization — it is an explicit importance weight on the squared residual.
Check
Identify the role of the weighting function.
Check your understanding
In the GloVe objective J = Σ f(X_ij)(w_i·w_j + b_i + b_j − log X_ij)², what is the purpose of f(X_ij) = min(1, (X_ij/x_max)^α)?
Answer: A
Why: f(X_ij) is a soft weight in [0,1]. For small X_ij it is small (rare co-occurrences get low weight, reducing noise). For X_ij ≥ x_max it is capped at 1 (common words like 'the' are not over-represented). The weighting is neither normalization nor regularization — it is an explicit importance weight on the squared residual.
Section
Project
Concept
Build Skip-gram NS from scratch — two embedding matrices, a binary loss, a training loop — and watch the embeddings shift so that king-man+woman moves closer to queen.
| # | milestone | key tool |
|---|---|---|
| 1 | Implement NS loss for one (center, context, negatives) triple; trace shapes | nn.Embedding, torch.dot, F.logsigmoid |
| 2 | Build the full training loop over skip-gram pairs; log loss at epoch 0/100/200 | SGD, manual_seed per epoch |
| 3 | Compute king−man+woman analogy vector; print cosine sims to all vocab words | F.cosine_similarity |
Build rules: use torch.manual_seed(0) for reproducibility; verify u_c shape = (D,), v_n shape = (K,D), loss > 0 before training. Do NOT use a pre-trained embedding — the point is to watch it learn from scratch.
Counterexample
Discussion prompt
Build Skip-gram NS from scratch — two embedding matrices, a binary loss, a training loop — and watch the embeddings shift so that king-man+woman moves closer to queen.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: implement the NS loss for (king, man, [noble, prince]). Predict: is the loss positive or negative before training?
Hint: loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum(). Random init → scores ≈ 0 → σ(0) = 0.5 → loss ≈ -log(0.5) × (K+1) ≈ 2.08.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
V, D, K = 8, 4, 3
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
c_t = torch.tensor(0) # king
o_t = torch.tensor(1) # man
n_t = torch.tensor([7, 6, 5]) # noble, prince, princess
u_c = U(c_t) # (D,)
v_o = W(o_t) # (D,)
v_n = W(n_t) # (K, D)
pos_score = torch.dot(u_c, v_o)
neg_scores = v_n @ u_c # (K,)
loss = -F.logsigmoid(pos_score) - F.logsigmoid(-neg_scores).sum()
print(f'u_c shape: {u_c.shape}, v_n shape: {v_n.shape}')
print(f'pos_score: {pos_score.item():.4f}')
print(f'loss: {loss.item():.4f}')| quantity | shape | value (seed=42) |
|---|---|---|
| u_c (king center embed) | (4,) | 4-dim random vector |
| v_o (man context embed) | (4,) | 4-dim random vector |
| v_n (3 neg context embeds) | (3, 4) | 3 rows of 4-dim random vecs |
| pos_score u_c·v_o | scalar | 0.0015 |
| NS loss | scalar | 2.7777 |
Discrimination
Sort into buckets
Sort these by shape, from memory, without looking back at Milestone 1 — NS loss one step. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Your turn: loop 200 epochs over the 8 skip-gram pairs (window=1 on [king,man,woman,queen,royal]). Predict: does loss go up or down? How fast?
Hint: set torch.manual_seed(epoch) inside the epoch loop to get reproducible noise samples. SGD lr=0.05 works; Adam lr=1e-3 converges faster.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(0)
V, D, K = 8, 4, 3
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
opt = torch.optim.SGD(list(U.parameters())+list(W.parameters()), lr=0.05)
pairs = [(0,1),(1,0),(1,2),(2,1),(2,3),(3,2),(3,4),(4,3)]
pair_t = [(torch.tensor(c),torch.tensor(o)) for c,o in pairs]
for epoch in range(201):
torch.manual_seed(epoch)
total = 0.0
for c_t, o_t in pair_t:
opt.zero_grad()
noise = torch.randint(0, V, (K,))
loss = (-F.logsigmoid(torch.dot(U(c_t), W(o_t)))
-F.logsigmoid(-(W(noise)@U(c_t))).sum())
loss.backward(); opt.step(); total += loss.item()
if epoch in [0,10,50,100,200]:
print(f'epoch {epoch:3d}: {total:.4f}')| epoch | total loss (8 pairs) |
|---|---|
| 0 | 22.1619 |
| 10 | 22.1292 |
| 50 | 16.1344 |
| 100 | 10.7495 |
| 200 | 10.8946 |
Pattern
Step through it
Step through Milestone 2 — training loop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: compute the king − man + woman analogy vector and rank all 8 vocab words by cosine similarity. Predict which word is closest (the toy corpus is tiny — exact correctness is not guaranteed).
Hint: analogy = E[w2i['king']] - E[w2i['man']] + E[w2i['woman']]; then F.cosine_similarity(analogy.unsqueeze(0), E) gives scores for all V words.
# continues from Milestone 2 (same U, trained 200 epochs)
vocab = ['king','man','woman','queen','royal','princess','prince','noble']
E = U.weight.detach() # (V, D)
analogy = E[0] - E[1] + E[2] # king - man + woman
sims = F.cosine_similarity(analogy.unsqueeze(0), E) # (V,)
ranked = sorted(zip(sims.tolist(), vocab), reverse=True)
for s, w in ranked:
print(f' {w:10s} {s:.4f}')| rank | word | cosine sim | note |
|---|---|---|---|
| 1 (excluded) | king | — | query word, excluded in real eval |
| 2 | prince | 0.7066 | toy corpus; shares 'royal' context |
| 3 | royal | 0.6626 | appears next to queen and king |
| 4 | queen | -0.0272 | tiny corpus — analogy geometry emerges at scale |
Trade off
Comparison matrix
From Milestone 3 — analogy evaluation: every row here is a choice with a cost. Fill the note column, then say which row you would actually pick and what you give up for it.
| rank | word | cosine sim | note |
|---|---|---|---|
| 1 (excluded) | king | — | query word, excluded in real eval |
| 2 | prince | 0.7066 | toy corpus; shares 'royal' context |
| 3 | royal | 0.6626 | appears next to queen and king |
| 4 | queen | -0.0272 | tiny corpus — analogy geometry emerges at scale |
Concept
import torch, torch.nn as nn, torch.nn.functional as F
vocab = ['king','man','woman','queen','royal','princess','prince','noble']
w2i = {w:i for i,w in enumerate(vocab)}
V, D, K = 8, 4, 3
# --- Skip-gram pairs (window=1) ---
sentence = [0,1,2,3,4]
pairs = [(sentence[p],sentence[p+o])
for p in range(len(sentence))
for o in [-1,1] if 0<=p+o<len(sentence)]
pair_t = [(torch.tensor(c),torch.tensor(o)) for c,o in pairs]
# --- Model ---
torch.manual_seed(0)
U = nn.Embedding(V, D); W = nn.Embedding(V, D)
nn.init.uniform_(U.weight, -0.5/D, 0.5/D)
nn.init.uniform_(W.weight, -0.5/D, 0.5/D)
opt = torch.optim.SGD(list(U.parameters())+list(W.parameters()), lr=0.05)
# --- Training ---
for epoch in range(201):
torch.manual_seed(epoch); total=0.0
for c_t,o_t in pair_t:
opt.zero_grad()
noise = torch.randint(0,V,(K,))
loss = (-F.logsigmoid(torch.dot(U(c_t),W(o_t)))
-F.logsigmoid(-(W(noise)@U(c_t))).sum())
loss.backward(); opt.step(); total+=loss.item()
if epoch in [0,100,200]: print(f'epoch {epoch}: {total:.4f}')
# --- Analogy ---
E = U.weight.detach()
analogy = E[w2i['king']]-E[w2i['man']]+E[w2i['woman']]
sims = F.cosine_similarity(analogy.unsqueeze(0), E)
print({vocab[i]:round(sims[i].item(),4) for i in range(V)})| output line | value |
|---|---|
| epoch 0 loss | 22.1619 |
| epoch 100 loss | 10.7495 |
| epoch 200 loss | 10.8946 |
| sim(queen) | -0.0272 (analogy needs scale to emerge) |
| sim(prince) | 0.7066 (shares royal context) |
At toy scale the analogy is noisy — the geometry requires hundreds of millions of training tokens to become reliable. What you've built is the exact mechanism Word2Vec uses; the only difference is corpus size.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output line | value |
|---|---|
| epoch 0 loss | 22.1619 |
| epoch 100 loss | 10.7495 |
| epoch 200 loss | 10.8946 |
| sim(queen) | -0.0272 (analogy needs scale to emerge) |
| sim(prince) | 0.7066 (shares royal context) |
Concept
Out loud, slides closed: (1) explain the NS loss — what is maximized, what is minimized, why no softmax; (2) contrast Skip-gram and CBOW on rare words; (3) state the GloVe objective and name both terms that differ from Skip-gram NS.
Homework (lesson plan): implement on real data — train on the first 1M tokens of Wikipedia text (download from Hugging Face datasets), evaluate on Google's word-analogy benchmark (8,869 semantic + 10,675 syntactic pairs), visualize 2D t-SNE clusters. Next up: Lesson 99 — contextual embeddings (ELMo, BERT contextualized representations).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Skip-gram — predict context from center · CBOW — predict center from context · GloVe — factorize co-occurrences · PMI geometry — why analogies work · Your turn: Word2Vec from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
-log σ(v_c·u_w) − Σ log σ(−v_n·u_w) — binary BCE, no softmaxf(X_ij)·(w_i·w_j + b_i + b_j − log X_ij)²| model | objective | key hyperparameter | best for |
|---|---|---|---|
| Skip-gram NS | binary BCE (SPPMI) | k=5–20 negatives | rare words, smaller corpora |
| CBOW | softmax or NS (mean ctx) | window size c | speed, large corpora |
| GloVe | weighted MSE on log X_ij | x_max=100, α=0.75 | offline matrix factorization |
| analogy | argmax cosine(king−man+woman, ·) | exclude query words | intrinsic embedding eval |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.