USAAIO Lesson 73, on learning embeddings in which distance reflects semantic similarity. It covers metric learning, Siamese networks, margin-based triplet loss, contrastive loss, online hard-negative mining, and the connection to attention. It implements a Siamese EmbeddingNet in PyTorch, where hard mining improves the ratio of inter-class to intra-class distance from 5.60× to 8.17× on a three-identity benchmark. The lesson runs to 30 slides.
Subject: Machine Learning · 55 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 73
Learn an embedding space where distance = semantic distance: triplet loss, contrastive loss, hard negative mining, and why it connects to the attention you built in the transformer.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 73: Metric Learning, Siamese Networks & Triplet Loss: without looking back, what was the main idea of Phase 2 Algorithm Review, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
a timed review of every classification and regression algorithm, ensemble methods (bagging/boosting/stacking), clustering (k-Means/GMM/DBSCAN/hierarchical), and a 30-minute PyTorch mastery check — all verified on sklearn datasets with real execution numbers.
Section
Part 1 of 3
Concept
Raw Euclidean distance in pixel space is meaningless: two photos of the same person under different lighting are far apart; two different people in matching poses may be close.
Metric learning trains a function f(x) such that ‖f(x_a) − f(x_p)‖ < ‖f(x_a) − f(x_n)‖ for same-class pairs (a, p) and different-class pairs (a, n).
| space | same-person distance | different-person distance |
|---|---|---|
| raw pixels | large (lighting change) | small (same pose) |
| learned embedding | small | large |
Comparison
Comparison matrix
From Learned distance, not raw distance: refill the same-person distance column from what you know. The rest of the table is as it appeared.
| space | same-person distance | different-person distance |
|---|---|---|
| raw pixels | large (lighting change) | small (same pose) |
| learned embedding | small | large |
Concept
A Siamese network is two copies of the same network (shared weights) applied to two inputs. The outputs are compared with a distance function — there is no softmax over fixed classes.
EmbeddingNet(x) with tied parametersShared weights are the key: each sample routes through identical f, so similarity is measured in a consistent space.
Counterexample
Discussion prompt
A Siamese network is two copies of the same network (shared weights) applied to two inputs. The outputs are compared with a distance function — there is no softmax over fixed classes.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Shared weights are the key: each sample routes through identical f, so similarity is measured in a consistent space.
Intuition
Project every embedding onto the unit hypersphere. Then ‖ea − eb‖² = 2 − 2·cos(θ) — Euclidean distance is a monotone function of angle. The network learns to spread identities across the sphere.
Without normalization the network can cheat by growing embedding magnitudes without moving the angular directions — you'd be measuring the scale of activations, not semantic similarity.
Analogy
Discussion prompt
Explain Why L2-normalize the output? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Project every embedding onto the unit hypersphere. Then ‖ea − eb‖² = 2 − 2·cos(θ) — Euclidean distance is a monotone function of angle. The network learns to spread identities across the sphere.
Section
Part 2 of 3
Concept
A triplet (a, p, n): anchor and positive are the same class; anchor and negative are different classes. The loss pushes d_ap below d_an by at least margin m.
\[ \mathcal{L}_{\text{trip}} = \max\!\bigl(0,\; d(a,p) - d(a,n) + m\bigr) \]
When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.
Explain it
Discussion prompt
Explain Triplet loss: anchor, positive, negative to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.
Estimation
Predict first
Use the trained EmbeddingNet (4-d, L2-normalized). Compute loss for 3 sampled triplets with margin m=0.5.
Commit before you compute: what does Triplet loss — step by step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Triplets 1–3 on the random-init model: triplets 1 and 3 are violated (negative too close), triplet 2 is satisfied
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Before any gradient steps the embedding is random.
Worked example
Use the trained EmbeddingNet (4-d, L2-normalized). Compute loss for 3 sampled triplets with margin m=0.5.
import torch, torch.nn as nn, torch.nn.functional as F
import numpy as np
from sklearn.datasets import make_blobs
np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150, 8, centers=3, cluster_std=4.0, random_state=42)
X = ((X - X.mean(0)) / (X.std(0)+1e-6)).astype(np.float32)
Xt = torch.tensor(X)
class EmbedNet(nn.Module):
def __init__(self): super().__init__(); self.net = nn.Sequential(nn.Linear(8,16), nn.ReLU(), nn.Linear(16,4))
def forward(self, x): return F.normalize(self.net(x), dim=-1)
model = EmbedNet()
# triplets (a, p, n) indices sampled with rng seed 10
triplets = [(118, 106, 63), (63, 141, 8), (67, 103, 15)]
print(f"{'triplet':>8} {'d_ap':>8} {'d_an':>8} {'loss':>8}")
for ai, pi, ni in triplets:
ea, ep, en = model(Xt[ai:ai+1]), model(Xt[pi:pi+1]), model(Xt[ni:ni+1])
dap = torch.norm(ea-ep).item(); dan = torch.norm(ea-en).item()
loss = max(0.0, dap - dan + 0.5)
print(f"{str((ai,pi,ni)):>8} {dap:.4f} {dan:.4f} {loss:.4f}")Triplets 1–3 on the random-init model: triplets 1 and 3 are violated (negative too close), triplet 2 is satisfied
Why: Before any gradient steps the embedding is random. Triplet (118,106,63): d_ap=0.68, d_an=0.79 — the negative lands barely farther than the positive, so loss=0.39. Triplet (63,141,8): negative is 1.06 away, far enough — satisfied. Triplet (67,103,15): negative is closer than positive, large loss=1.14.
| triplet | d(a,p) | d(a,n) | loss (m=0.5) |
|---|---|---|---|
| (118,106,63) | 0.6812 | 0.7907 | 0.3905 |
| (63,141,8) | 0.2588 | 1.0568 | 0.0000 |
| (67,103,15) | 1.4125 | 0.7681 | 1.1444 |
Trade off
Comparison matrix
From Triplet loss — step by step: every row here is a choice with a cost. Fill the loss (m=0.5) column, then say which row you would actually pick and what you give up for it.
| triplet | d(a,p) | d(a,n) | loss (m=0.5) |
|---|---|---|---|
| (118,106,63) | 0.6812 | 0.7907 | 0.3905 |
| (63,141,8) | 0.2588 | 1.0568 | 0.0000 |
| (67,103,15) | 1.4125 | 0.7681 | 1.1444 |
Concept
Contrastive loss operates on pairs (x_i, x_j) with a binary label y=1 (same class) or y=0 (different). Minimize distance for positives; push negatives past margin m.
\[ \mathcal{L}_{\text{contra}} = y\,d^2 + (1-y)\,\max(0,\, m - d)^2 \]
| pair type | label y | loss drives |
|---|---|---|
| same class (positive) | 1 | minimize d |
| diff class (negative) | 0 | push d above margin m |
Estimation
Predict first
Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).
Commit before you compute: what does Contrastive loss — two pairs come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On the random-init model, intra-class distance is large (0.99) — the model hasn't learned to cluster them yet, hence a high contrastive loss (0.97).
Worked example
Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).
# After 300-epoch training, embeddings from model_rand
# pair (0,1): same class (label=1)
# pair (0,2): different class (label=0)
# model_rand embeddings (L2-normalized 4-d)
with torch.no_grad():
embs = model(Xt[:3]) # shape (3, 4)
d01 = torch.norm(embs[0]-embs[1]).item()
d02 = torch.norm(embs[0]-embs[2]).item()
margin = 1.0
loss_same = d01**2 # y=1
loss_diff = max(0, margin-d02)**2 # y=0
print(f'd(0,1)={d01:.4f} L_same={loss_same:.4f}')
print(f'd(0,2)={d02:.4f} L_diff={loss_diff:.4f}')same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)
Why: On the random-init model, intra-class distance is large (0.99) — the model hasn't learned to cluster them yet, hence a high contrastive loss (0.97). The diff-class pair is already separated beyond the margin, so its loss is 0.
| pair | class rel. | d | contrastive loss |
|---|---|---|---|
| (0,1) | same (y=1) | 0.9857 | 0.9716 |
| (0,2) | different (y=0) | 1.2357 | 0.0000 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).
Anomaly
Predict first
A student writes this, and it looks reasonable:
To compute triplet loss, pass two inputs (anchor, positive) — just like contrastive loss on a same-class pair.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Triplet loss takes THREE embeddings (anchor, positive, negative).
Triplet loss always takes three embeddings: anchor, positive, negative.
Why: Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.
Trap
To compute triplet loss, pass two inputs (anchor, positive) — just like contrastive loss on a same-class pair.
loss = triplet_loss(f(a), f(p))
Why: Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.
Triplet loss always takes three embeddings: anchor, positive, negative.
loss = max(0, norm(f(a)-f(p)) - norm(f(a)-f(n)) + margin)
Why: All three play distinct roles: f(a)-f(p) is what we minimize, f(a)-f(n) is what we maximize. No binary label needed — the class membership is encoded in the triplet structure itself.
Break the constraint
Discussion prompt
The rule this trap just fixed:
All three play distinct roles: f(a)-f(p) is what we minimize, f(a)-f(n) is what we maximize. No binary label needed — the class membership is encoded in the triplet structure itself.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.
Section
Part 3 of 3
Concept
After a few epochs, the vast majority of randomly sampled triplets are already satisfied (d_ap < d_an − m). They contribute zero gradient and zero learning — you're burning compute on easy examples.
Hard negative mining selects the most informative negatives: the negative closest to the anchor in the current embedding space. These are the ones the model is most confused by.
| strategy | negative selection | epoch-1 loss | converged epoch |
|---|---|---|---|
| random negatives | random from diff-class | 0.26 | ~50 |
| hard negative mining | closest diff-class | 1.54 | ~150 |
Concept
Offline mining: run inference on all data before each epoch, compute all pairwise distances, pick the hardest triplets globally. Expensive but thorough.
Online mining: within each mini-batch, find the hardest positive (farthest same-class) and hardest negative (closest diff-class) on the fly. Standard in modern face recognition (FaceNet).
argmax_{p: y_p=y_a} ‖f(a)−f(p)‖ within batchargmin_{n: y_n≠y_a} ‖f(a)−f(n)‖ within batchEstimation
Predict first
Train two identical Siamese nets (8-d → 16 → 4-d, L2-norm) for 300 epochs on a 3-identity benchmark. Only the negative-sampling strategy differs.
Commit before you compute: what does Random vs hard negatives — training comparison come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Hard mining has 6× higher epoch-1 loss (1.54 vs 0.26) — it immediately selects violated triplets
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Random sampling finds mostly easy (satisfied) triplets in epoch 1; hard mining always picks the closest negative, which is always a violated triplet at the start.
Worked example
Train two identical Siamese nets (8-d → 16 → 4-d, L2-norm) for 300 epochs on a 3-identity benchmark. Only the negative-sampling strategy differs.
import torch, torch.nn as nn, torch.nn.functional as F, numpy as np
from sklearn.datasets import make_blobs
np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150,8,centers=3,cluster_std=4.0,random_state=42)
X = ((X-X.mean(0))/(X.std(0)+1e-6)).astype(np.float32); Xt=torch.tensor(X)
class EmbedNet(nn.Module):
def __init__(self): super().__init__(); self.net=nn.Sequential(nn.Linear(8,16),nn.ReLU(),nn.Linear(16,4))
def forward(self,x): return F.normalize(self.net(x),dim=-1)
def samp(rng,hard=False,embs=None):
ai=int(rng.integers(150)); lbl=y[ai]
pp=np.where(y==lbl)[0]; pp=pp[pp!=ai]; np2=np.where(y!=lbl)[0]
if hard and embs is not None:
pi=int(pp[np.argmax(np.linalg.norm(embs[pp]-embs[ai],axis=1))])
ni=int(np2[np.argmin(np.linalg.norm(embs[np2]-embs[ai],axis=1))])
else: pi=int(rng.choice(pp)); ni=int(rng.choice(np2))
return ai,pi,ni
def train(hard,seed=0):
torch.manual_seed(seed); rng=np.random.default_rng(seed)
m=EmbedNet(); opt=torch.optim.Adam(m.parameters(),lr=5e-3); losses=[]
for _ in range(300):
embs=m(Xt).detach().numpy() if hard else None
opt.zero_grad(); total=torch.tensor(0.0)
for __ in range(16):
ai,pi,ni=samp(rng,hard=hard,embs=embs)
ea,ep,en=m(Xt[ai:ai+1]),m(Xt[pi:pi+1]),m(Xt[ni:ni+1])
total+=torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5,min=0)
(total/16).backward(); opt.step(); losses.append((total/16).item())
return m,losses
model_rand,lr=train(False); model_hard,lh=train(True)
print('rand ep1:',round(lr[0],4),'ep50:',round(lr[49],4),'ep150:',round(lr[149],4))
print('hard ep1:',round(lh[0],4),'ep50:',round(lh[49],4),'ep150:',round(lh[149],4))Hard mining has 6× higher epoch-1 loss (1.54 vs 0.26) — it immediately selects violated triplets
Why: Random sampling finds mostly easy (satisfied) triplets in epoch 1; hard mining always picks the closest negative, which is always a violated triplet at the start. The signal is denser, so the embedding improves faster per meaningful gradient step.
| strategy | epoch 1 | epoch 10 | epoch 50 | epoch 150 |
|---|---|---|---|---|
| random negatives | 0.2616 | 0.2572 | 0.0000 | 0.0000 |
| hard negatives | 1.5390 | 0.9517 | 0.4862 | 0.0000 |
Comparison
Comparison matrix
From Random vs hard negatives — training comparison: refill the epoch 10 column from what you know. The rest of the table is as it appeared.
| strategy | epoch 1 | epoch 10 | epoch 50 | epoch 150 |
|---|---|---|---|---|
| random negatives | 0.2616 | 0.2572 | 0.0000 | 0.0000 |
| hard negatives | 1.5390 | 0.9517 | 0.4862 | 0.0000 |
Concept
After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.
| strategy | intra-class d | inter-class d | inter/intra ratio |
|---|---|---|---|
| random negatives | 0.3024 | 1.6938 | 5.60× |
| hard negatives | 0.1886 | 1.5414 | 8.17× |
Hard mining reduces intra-class spread more aggressively (0.30 → 0.19) while maintaining inter-class separation, improving the ratio by 1.5×.
Analogy
Discussion prompt
Explain Embedding quality: inter/intra distance ratio by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Always select the absolute closest diff-class sample as the negative — hardest means best signal.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The globally hardest negative is often an outlier or a mislabeled sample.
Use semi-hard negatives: farther than the positive but still within the margin.
Why: The globally hardest negative is often an outlier or a mislabeled sample. Gradients from a nearly-identical but wrong-class sample can be huge and destabilize training (mode collapse: all embeddings map to the same point).
Trap
Always select the absolute closest diff-class sample as the negative — hardest means best signal.
neg = argmin_{n: y_n≠y_a} ‖f(a)−f(n)‖ (global closest)
Why: The globally hardest negative is often an outlier or a mislabeled sample. Gradients from a nearly-identical but wrong-class sample can be huge and destabilize training (mode collapse: all embeddings map to the same point).
Use semi-hard negatives: farther than the positive but still within the margin.
neg = n where d(a,p) < d(a,n) < d(a,p) + margin
Why: Semi-hard negatives provide a gradient without being so extreme they destabilize learning. They're the FaceNet standard: hard enough to be informative, not so hard they cause collapse.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f, so similarity is measured in a consistent space.; When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.; After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.Concept
Both metric learning and attention compute similarity scores between vectors. In dot-product attention (Lesson 65): scores = QK^T / √d — the raw score is a cosine similarity between L2-normalized query and key.
\[ \text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V \]
If Q and K are L2-normalized embeddings, Q·K^T is exactly the cosine similarity matrix — attention weights are metric-learning similarity scores routed through a softmax.
Counterexample
Discussion prompt
If Q and K are L2-normalized embeddings, Q·K^T is exactly the cosine similarity matrix — attention weights are metric-learning similarity scores routed through a softmax.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Ranking
Put in order
These are the steps of The metric learning recipe, scrambled. Put them back in order before the next slide shows you.
loss = max(0, d_ap − d_an + m)loss = y·d² + (1−y)·max(0, m−d)²Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
loss = max(0, d_ap − d_an + m)loss = y·d² + (1−y)·max(0, m−d)²Edge cases
Discussion prompt
The metric learning recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
loss = max(0, d_ap − d_an + m)loss = y·d² + (1−y)·max(0, m−d)²Elimination
Eliminate the wrong options
A triplet (a, p, n) has d(a,p)=0.8 and d(a,n)=1.1. With margin m=0.5, what is the triplet loss?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: loss = max(0, 0.8 − 1.1 + 0.5) = max(0, 0.2) = 0.20. The triplet is violated because d_an − d_ap = 0.3, which is less than margin 0.5 — the negative is not far enough.
Check
Work it out before clicking.
Check your understanding
A triplet (a, p, n) has d(a,p)=0.8 and d(a,n)=1.1. With margin m=0.5, what is the triplet loss?
Answer: A
Why: loss = max(0, 0.8 − 1.1 + 0.5) = max(0, 0.2) = 0.20. The triplet is violated because d_an − d_ap = 0.3, which is less than margin 0.5 — the negative is not far enough.
Prediction
Predict first
Why does online hard negative mining produce a higher epoch-1 loss than random sampling?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: It always selects the closest diff-class sample, which is the most-violated triplet
Why: Hard mining picks the negative closest to the anchor, guaranteeing a violated triplet (d_an is minimized, so d_ap − d_an + m is maximized). Random sampling mostly finds easy triplets that are already satisfied, contributing zero loss.
Check
Think about what 'hard' means.
Check your understanding
Why does online hard negative mining produce a higher epoch-1 loss than random sampling?
Answer: A
Why: Hard mining picks the negative closest to the anchor, guaranteeing a violated triplet (d_an is minimized, so d_ap − d_an + m is maximized). Random sampling mostly finds easy triplets that are already satisfied, contributing zero loss.
Elimination
Eliminate the wrong options
With L2-normalized query Q and key K vectors, the raw attention score Q·K^T equals:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: For unit-norm vectors, Q·K = |Q||K|cos(θ) = cos(θ). The dot product of L2-normalized embeddings IS cosine similarity — so attention weights are softmax-normalized metric learning similarity scores.
Check
Connect the two ideas.
Check your understanding
With L2-normalized query Q and key K vectors, the raw attention score Q·K^T equals:
Answer: A
Why: For unit-norm vectors, Q·K = |Q||K|cos(θ) = cos(θ). The dot product of L2-normalized embeddings IS cosine similarity — so attention weights are softmax-normalized metric learning similarity scores.
Section
Project
Concept
Build a Siamese network that learns a 4-d embedding space for a 3-identity dataset. Train it twice — once with random negatives and once with hard negatives — and compare the inter/intra distance ratio.
| # | milestone | key step |
|---|---|---|
| 1 | EmbeddingNet + single triplet loss | forward pass, L2-norm, clamp |
| 2 | training loop (random negatives) | 300 epochs, Adam, 16 triplets/epoch |
| 3 | hard negative mining | per-epoch embedding snapshot, argmin/argmax |
Build rules: L2-normalize outputs with F.normalize(out, dim=-1). Never pass a satisfied triplet's gradient through — use torch.clamp(..., min=0). Verify inter/intra ratio improves with hard mining.
Worked example
Your turn: define EmbeddingNet (8→16→4, ReLU, L2-norm) and compute the triplet loss for one (a, p, n) tuple by hand.
Hint: F.normalize(self.net(x), dim=-1) returns unit-norm vectors. torch.norm(ea-ep) gives L2 distance. Loss = torch.clamp(dap - dan + margin, min=0).
import torch, torch.nn as nn, torch.nn.functional as F
import numpy as np
from sklearn.datasets import make_blobs
np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150, 8, centers=3, cluster_std=4.0, random_state=42)
X = ((X - X.mean(0))/(X.std(0)+1e-6)).astype(np.float32)
Xt = torch.tensor(X)
class EmbedNet(nn.Module):
def __init__(self): super().__init__(); self.net = nn.Sequential(nn.Linear(8,16), nn.ReLU(), nn.Linear(16,4))
def forward(self, x): return F.normalize(self.net(x), dim=-1)
model = EmbedNet()
ai, pi, ni = 118, 106, 63 # anchor, pos, neg
ea, ep, en = model(Xt[ai:ai+1]), model(Xt[pi:pi+1]), model(Xt[ni:ni+1])
dap = torch.norm(ea-ep).item(); dan = torch.norm(ea-en).item()
loss = torch.clamp(torch.tensor(dap - dan + 0.5), min=0).item()
print(f'dap={dap:.4f} dan={dan:.4f} loss={loss:.4f}')
print('L2 norm ea:', round(torch.norm(ea).item(), 4))| value | result |
|---|---|
| d(a,p) | 0.3478 |
| d(a,n) | 1.1799 |
| triplet loss | 0.0000 (satisfied) |
| ‖ea‖ | 1.0000 (unit norm) |
Worked example
Your turn: write the 300-epoch loop with 16 random triplets per step. Predict where the loss plateaus.
Hint: opt.zero_grad() once per epoch (outside the inner loop); accumulate losses across the 16 triplets, then divide and .backward(). Use Adam(lr=5e-3).
opt = torch.optim.Adam(model.parameters(), lr=5e-3)
rng = np.random.default_rng(0)
def sample_triplet(X, y, rng):
ai = int(rng.integers(len(X))); lbl = y[ai]
pos = np.where(y==lbl)[0]; pos = pos[pos!=ai]
neg = np.where(y!=lbl)[0]
return ai, int(rng.choice(pos)), int(rng.choice(neg))
for epoch in range(300):
opt.zero_grad(); total = torch.tensor(0.0)
for _ in range(16):
ai,pi,ni = sample_triplet(X, y, rng)
ea,ep,en = model(Xt[ai:ai+1]),model(Xt[pi:pi+1]),model(Xt[ni:ni+1])
total += torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5, min=0)
(total/16).backward(); opt.step()
print('final loss:', round((total/16).item(), 4))| epoch | loss (random neg) |
|---|---|
| 1 | 0.2616 |
| 10 | 0.2572 |
| 50 | 0.0000 |
| 300 | 0.0000 |
Trade off
Comparison matrix
From Milestone 2 — training loop, random negatives: every row here is a choice with a cost. Fill the loss (random neg) column, then say which row you would actually pick and what you give up for it.
| epoch | loss (random neg) |
|---|---|
| 1 | 0.2616 |
| 10 | 0.2572 |
| 50 | 0.0000 |
| 300 | 0.0000 |
Worked example
Your turn: re-train with hard negatives (snapshot embeddings each epoch, pick argmin diff-class distance). Predict whether hard mining reaches a higher or lower inter/intra ratio.
Hint: embs_np = model(Xt).detach().numpy() at the top of each epoch. For hard neg: neg_pool = where(y!=y_a), neg_idx = neg_pool[argmin(‖embs_np[neg_pool] − embs_np[a_i]‖)].
torch.manual_seed(0); model_h = EmbedNet()
opt_h = torch.optim.Adam(model_h.parameters(), lr=5e-3)
rng2 = np.random.default_rng(0)
for epoch in range(300):
with torch.no_grad(): embs_np = model_h(Xt).numpy()
opt_h.zero_grad(); total = torch.tensor(0.0)
for _ in range(16):
ai = int(rng2.integers(len(X))); lbl = y[ai]
pos_pool = np.where(y==lbl)[0]; pos_pool = pos_pool[pos_pool!=ai]
neg_pool = np.where(y!=lbl)[0]
pi = int(pos_pool[np.argmax(np.linalg.norm(embs_np[pos_pool]-embs_np[ai], axis=1))])
ni = int(neg_pool[np.argmin(np.linalg.norm(embs_np[neg_pool]-embs_np[ai], axis=1))])
ea,ep,en = model_h(Xt[ai:ai+1]),model_h(Xt[pi:pi+1]),model_h(Xt[ni:ni+1])
total += torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5, min=0)
(total/16).backward(); opt_h.step()
print('hard final loss:', round((total/16).item(), 4))| strategy | intra d | inter d | inter/intra ratio |
|---|---|---|---|
| random neg | 0.3024 | 1.6938 | 5.60× |
| hard neg | 0.1886 | 1.5414 | 8.17× |
Comparison
Comparison matrix
From Milestone 3 — hard negative mining + ratio comparison: refill the inter d column from what you know. The rest of the table is as it appeared.
| strategy | intra d | inter d | inter/intra ratio |
|---|---|---|---|
| random neg | 0.3024 | 1.6938 | 5.60× |
| hard neg | 0.1886 | 1.5414 | 8.17× |
Concept
Out loud, slides closed: (1) explain the triplet loss formula and why satisfied triplets contribute zero gradient, (2) describe what hard negative mining selects and why it can cause collapse without a semi-hard guard, (3) state the connection between L2-normalized embeddings and attention scores.
Stretch (homework): implement face verification with a threshold on d(f(x1), f(x2)) and plot ROC at various thresholds; compare contrastive vs triplet loss on the same 3-identity data; analyze how metric learning relates to attention using the similarity matrix Q·K^T from Lesson 65.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What is metric learning? · Triplet loss & contrastive loss · Hard negative mining · Your turn: Siamese face verifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
f(x) so distance = semantic similaritymax(0, d_ap − d_an + m); zero for satisfied tripletsy·d² + (1−y)·max(0,m−d)²; pair-based alternativeQ·K^T = cosine similarity = metric learning score| idea | the one thing to remember |
|---|---|
| triplet loss | satisfied triplets give zero gradient — only violations train |
| Siamese net | shared weights + L2-norm; no fixed class head |
| hard mining | closest negative = most violated = fastest signal (but semi-hard to avoid collapse) |
| attention link | dot product of unit-norm embeddings = cosine similarity |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.