Lesson 73: Metric Learning, Siamese Networks & Triplet Loss

USAAIO Lesson 73, on learning embeddings in which distance reflects semantic similarity. It covers metric learning, Siamese networks, margin-based triplet loss, contrastive loss, online hard-negative mining, and the connection to attention. It implements a Siamese EmbeddingNet in PyTorch, where hard mining improves the ratio of inter-class to intra-class distance from 5.60× to 8.17× on a three-identity benchmark. The lesson runs to 30 slides.

Subject: Machine Learning · 55 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Metric Learning & Siamese Networks

Title

USAAIO · Lesson 73

Learn an embedding space where distance = semantic distance: triplet loss, contrastive loss, hard negative mining, and why it connects to the attention you built in the transformer.

2. By the end of this lesson you can

Objectives

  1. Explain metric learning: optimize an embedding so distance reflects similarity
  2. Build a Siamese network that maps two inputs through shared weights and compares embeddings
  3. Implement triplet loss: pull anchor-positive together, push anchor-negative apart with margin m
  4. Implement contrastive loss and identify when each loss is appropriate
  5. Apply online hard negative mining and explain why it accelerates convergence

3. What survived from Phase 2 Algorithm Review?

Warm-up

Discussion prompt

Before we open Lesson 73: Metric Learning, Siamese Networks & Triplet Loss: without looking back, what was the main idea of Phase 2 Algorithm Review, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a timed review of every classification and regression algorithm, ensemble methods (bagging/boosting/stacking), clustering (k-Means/GMM/DBSCAN/hierarchical), and a 30-minute PyTorch mastery check — all verified on sklearn datasets with real execution numbers.

4. What is metric learning?

Section

Part 1 of 3

5. Learned distance, not raw distance

Concept

Raw Euclidean distance in pixel space is meaningless: two photos of the same person under different lighting are far apart; two different people in matching poses may be close.

Metric learning trains a function f(x) such that ‖f(x_a) − f(x_p)‖ < ‖f(x_a) − f(x_n)‖ for same-class pairs (a, p) and different-class pairs (a, n).

spacesame-person distancedifferent-person distance
raw pixelslarge (lighting change)small (same pose)
learned embeddingsmalllarge

6. Fill in: same-person distance for Learned distance, not raw distance

Comparison

Comparison matrix

From Learned distance, not raw distance: refill the same-person distance column from what you know. The rest of the table is as it appeared.

spacesame-person distancedifferent-person distance
raw pixelslarge (lighting change)small (same pose)
learned embeddingsmalllarge

7. Siamese network architecture

Concept

A Siamese network is two copies of the same network (shared weights) applied to two inputs. The outputs are compared with a distance function — there is no softmax over fixed classes.

Shared weights are the key: each sample routes through identical f, so similarity is measured in a consistent space.

8. Break it if you can: Siamese network architecture

Counterexample

Discussion prompt

A Siamese network is two copies of the same network (shared weights) applied to two inputs. The outputs are compared with a distance function — there is no softmax over fixed classes.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Shared weights are the key: each sample routes through identical f, so similarity is measured in a consistent space.

9. Why L2-normalize the output?

Intuition

Project every embedding onto the unit hypersphere. Then ‖ea − eb‖² = 2 − 2·cos(θ) — Euclidean distance is a monotone function of angle. The network learns to spread identities across the sphere.

Without normalization the network can cheat by growing embedding magnitudes without moving the angular directions — you'd be measuring the scale of activations, not semantic similarity.

10. By analogy: Why L2-normalize the output?

Analogy

Discussion prompt

Explain Why L2-normalize the output? by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Project every embedding onto the unit hypersphere. Then ‖ea − eb‖² = 2 − 2·cos(θ) — Euclidean distance is a monotone function of angle. The network learns to spread identities across the sphere.

11. Triplet loss & contrastive loss

Section

Part 2 of 3

12. Triplet loss: anchor, positive, negative

Concept

A triplet (a, p, n): anchor and positive are the same class; anchor and negative are different classes. The loss pushes d_ap below d_an by at least margin m.

\[ \mathcal{L}_{\text{trip}} = \max\!\bigl(0,\; d(a,p) - d(a,n) + m\bigr) \]

When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.

13. Teach it back: Triplet loss: anchor, positive, negative

Explain it

Discussion prompt

Explain Triplet loss: anchor, positive, negative to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.

14. Guess the shape of the answer: Triplet loss — step by step

Estimation

Predict first

Use the trained EmbeddingNet (4-d, L2-normalized). Compute loss for 3 sampled triplets with margin m=0.5.

Commit before you compute: what does Triplet loss — step by step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Triplets 1–3 on the random-init model: triplets 1 and 3 are violated (negative too close), triplet 2 is satisfied

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Before any gradient steps the embedding is random.

15. Triplet loss — step by step

Worked example

Use the trained EmbeddingNet (4-d, L2-normalized). Compute loss for 3 sampled triplets with margin m=0.5.

import torch, torch.nn as nn, torch.nn.functional as F
import numpy as np
from sklearn.datasets import make_blobs

np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150, 8, centers=3, cluster_std=4.0, random_state=42)
X = ((X - X.mean(0)) / (X.std(0)+1e-6)).astype(np.float32)
Xt = torch.tensor(X)

class EmbedNet(nn.Module):
    def __init__(self): super().__init__(); self.net = nn.Sequential(nn.Linear(8,16), nn.ReLU(), nn.Linear(16,4))
    def forward(self, x): return F.normalize(self.net(x), dim=-1)

model = EmbedNet()
# triplets (a, p, n) indices sampled with rng seed 10
triplets = [(118, 106, 63), (63, 141, 8), (67, 103, 15)]
print(f"{'triplet':>8}  {'d_ap':>8}  {'d_an':>8}  {'loss':>8}")
for ai, pi, ni in triplets:
    ea, ep, en = model(Xt[ai:ai+1]), model(Xt[pi:pi+1]), model(Xt[ni:ni+1])
    dap = torch.norm(ea-ep).item(); dan = torch.norm(ea-en).item()
    loss = max(0.0, dap - dan + 0.5)
    print(f"{str((ai,pi,ni)):>8}  {dap:.4f}    {dan:.4f}    {loss:.4f}")

Triplets 1–3 on the random-init model: triplets 1 and 3 are violated (negative too close), triplet 2 is satisfied

Why: Before any gradient steps the embedding is random. Triplet (118,106,63): d_ap=0.68, d_an=0.79 — the negative lands barely farther than the positive, so loss=0.39. Triplet (63,141,8): negative is 1.06 away, far enough — satisfied. Triplet (67,103,15): negative is closer than positive, large loss=1.14.

tripletd(a,p)d(a,n)loss (m=0.5)
(118,106,63)0.68120.79070.3905
(63,141,8)0.25881.05680.0000
(67,103,15)1.41250.76811.1444

16. What each one costs: Triplet loss — step by step

Trade off

Comparison matrix

From Triplet loss — step by step: every row here is a choice with a cost. Fill the loss (m=0.5) column, then say which row you would actually pick and what you give up for it.

tripletd(a,p)d(a,n)loss (m=0.5)
(118,106,63)0.68120.79070.3905
(63,141,8)0.25881.05680.0000
(67,103,15)1.41250.76811.1444

17. Contrastive loss: pair-based alternative

Concept

Contrastive loss operates on pairs (x_i, x_j) with a binary label y=1 (same class) or y=0 (different). Minimize distance for positives; push negatives past margin m.

\[ \mathcal{L}_{\text{contra}} = y\,d^2 + (1-y)\,\max(0,\, m - d)^2 \]

pair typelabel yloss drives
same class (positive)1minimize d
diff class (negative)0push d above margin m

18. Guess the shape of the answer: Contrastive loss — two pairs

Estimation

Predict first

Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).

Commit before you compute: what does Contrastive loss — two pairs come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On the random-init model, intra-class distance is large (0.99) — the model hasn't learned to cluster them yet, hence a high contrastive loss (0.97).

19. Contrastive loss — two pairs

Worked example

Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).

# After 300-epoch training, embeddings from model_rand
# pair (0,1): same class (label=1)
# pair (0,2): different class (label=0)
# model_rand embeddings (L2-normalized 4-d)
with torch.no_grad():
    embs = model(Xt[:3])         # shape (3, 4)
    d01 = torch.norm(embs[0]-embs[1]).item()
    d02 = torch.norm(embs[0]-embs[2]).item()

margin = 1.0
loss_same = d01**2               # y=1
loss_diff = max(0, margin-d02)**2  # y=0
print(f'd(0,1)={d01:.4f}  L_same={loss_same:.4f}')
print(f'd(0,2)={d02:.4f}  L_diff={loss_diff:.4f}')

same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)

Why: On the random-init model, intra-class distance is large (0.99) — the model hasn't learned to cluster them yet, hence a high contrastive loss (0.97). The diff-class pair is already separated beyond the margin, so its loss is 0.

pairclass rel.dcontrastive loss
(0,1)same (y=1)0.98570.9716
(0,2)different (y=0)1.23570.0000

20. Work backwards from the answer: Contrastive loss — two pairs

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

same-class pair (0,1): d=0.9857, loss=0.9716 (random-init model not yet collapsed); diff-class pair (0,2): d=1.2357, loss=0.0 (already beyond margin=1.0)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Compute contrastive loss for a same-class pair and a different-class pair after 300 epochs of training (margin m=1.0).

21. Something is wrong here: confusing triplet and contrastive loss inputs

Anomaly

Predict first

A student writes this, and it looks reasonable:

To compute triplet loss, pass two inputs (anchor, positive) — just like contrastive loss on a same-class pair.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Triplet loss takes THREE embeddings (anchor, positive, negative).

Triplet loss always takes three embeddings: anchor, positive, negative.

Why: Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.

22. Trap: confusing triplet and contrastive loss inputs

Trap

The trap

To compute triplet loss, pass two inputs (anchor, positive) — just like contrastive loss on a same-class pair.

loss = triplet_loss(f(a), f(p))

Why: Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.

The fix

Triplet loss always takes three embeddings: anchor, positive, negative.

loss = max(0, norm(f(a)-f(p)) - norm(f(a)-f(n)) + margin)

Why: All three play distinct roles: f(a)-f(p) is what we minimize, f(a)-f(n) is what we maximize. No binary label needed — the class membership is encoded in the triplet structure itself.

23. Break it on purpose: confusing triplet and contrastive loss inputs

Break the constraint

Discussion prompt

The rule this trap just fixed:

All three play distinct roles: f(a)-f(p) is what we minimize, f(a)-f(n) is what we maximize. No binary label needed — the class membership is encoded in the triplet structure itself.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Triplet loss takes THREE embeddings (anchor, positive, negative). Passing only two is a contrastive loss. The distinction matters: contrastive loss needs explicit pair labels; triplet loss is label-free and defines the contrast internally via the triplet structure.

24. Hard negative mining

Section

Part 3 of 3

25. Why most triplets are useless

Concept

After a few epochs, the vast majority of randomly sampled triplets are already satisfied (d_ap < d_an − m). They contribute zero gradient and zero learning — you're burning compute on easy examples.

Hard negative mining selects the most informative negatives: the negative closest to the anchor in the current embedding space. These are the ones the model is most confused by.

strategynegative selectionepoch-1 lossconverged epoch
random negativesrandom from diff-class0.26~50
hard negative miningclosest diff-class1.54~150

26. Hard mining strategy: online vs offline

Concept

Offline mining: run inference on all data before each epoch, compute all pairwise distances, pick the hardest triplets globally. Expensive but thorough.

Online mining: within each mini-batch, find the hardest positive (farthest same-class) and hardest negative (closest diff-class) on the fly. Standard in modern face recognition (FaceNet).

27. Guess the shape of the answer: Random vs hard negatives — training…

Estimation

Predict first

Train two identical Siamese nets (8-d → 16 → 4-d, L2-norm) for 300 epochs on a 3-identity benchmark. Only the negative-sampling strategy differs.

Commit before you compute: what does Random vs hard negatives — training comparison come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Hard mining has 6× higher epoch-1 loss (1.54 vs 0.26) — it immediately selects violated triplets

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Random sampling finds mostly easy (satisfied) triplets in epoch 1; hard mining always picks the closest negative, which is always a violated triplet at the start.

28. Random vs hard negatives — training comparison

Worked example

Train two identical Siamese nets (8-d → 16 → 4-d, L2-norm) for 300 epochs on a 3-identity benchmark. Only the negative-sampling strategy differs.

import torch, torch.nn as nn, torch.nn.functional as F, numpy as np
from sklearn.datasets import make_blobs
np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150,8,centers=3,cluster_std=4.0,random_state=42)
X = ((X-X.mean(0))/(X.std(0)+1e-6)).astype(np.float32); Xt=torch.tensor(X)
class EmbedNet(nn.Module):
    def __init__(self): super().__init__(); self.net=nn.Sequential(nn.Linear(8,16),nn.ReLU(),nn.Linear(16,4))
    def forward(self,x): return F.normalize(self.net(x),dim=-1)
def samp(rng,hard=False,embs=None):
    ai=int(rng.integers(150)); lbl=y[ai]
    pp=np.where(y==lbl)[0]; pp=pp[pp!=ai]; np2=np.where(y!=lbl)[0]
    if hard and embs is not None:
        pi=int(pp[np.argmax(np.linalg.norm(embs[pp]-embs[ai],axis=1))])
        ni=int(np2[np.argmin(np.linalg.norm(embs[np2]-embs[ai],axis=1))])
    else: pi=int(rng.choice(pp)); ni=int(rng.choice(np2))
    return ai,pi,ni
def train(hard,seed=0):
    torch.manual_seed(seed); rng=np.random.default_rng(seed)
    m=EmbedNet(); opt=torch.optim.Adam(m.parameters(),lr=5e-3); losses=[]
    for _ in range(300):
        embs=m(Xt).detach().numpy() if hard else None
        opt.zero_grad(); total=torch.tensor(0.0)
        for __ in range(16):
            ai,pi,ni=samp(rng,hard=hard,embs=embs)
            ea,ep,en=m(Xt[ai:ai+1]),m(Xt[pi:pi+1]),m(Xt[ni:ni+1])
            total+=torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5,min=0)
        (total/16).backward(); opt.step(); losses.append((total/16).item())
    return m,losses
model_rand,lr=train(False); model_hard,lh=train(True)
print('rand ep1:',round(lr[0],4),'ep50:',round(lr[49],4),'ep150:',round(lr[149],4))
print('hard ep1:',round(lh[0],4),'ep50:',round(lh[49],4),'ep150:',round(lh[149],4))

Hard mining has 6× higher epoch-1 loss (1.54 vs 0.26) — it immediately selects violated triplets

Why: Random sampling finds mostly easy (satisfied) triplets in epoch 1; hard mining always picks the closest negative, which is always a violated triplet at the start. The signal is denser, so the embedding improves faster per meaningful gradient step.

strategyepoch 1epoch 10epoch 50epoch 150
random negatives0.26160.25720.00000.0000
hard negatives1.53900.95170.48620.0000

29. Fill in: epoch 10 for Random vs hard negatives — training…

Comparison

Comparison matrix

From Random vs hard negatives — training comparison: refill the epoch 10 column from what you know. The rest of the table is as it appeared.

strategyepoch 1epoch 10epoch 50epoch 150
random negatives0.26160.25720.00000.0000
hard negatives1.53900.95170.48620.0000

30. Embedding quality: inter/intra distance ratio

Concept

After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.

strategyintra-class dinter-class dinter/intra ratio
random negatives0.30241.69385.60×
hard negatives0.18861.54148.17×

Hard mining reduces intra-class spread more aggressively (0.30 → 0.19) while maintaining inter-class separation, improving the ratio by 1.5×.

31. By analogy: Embedding quality: inter/intra distance ratio

Analogy

Discussion prompt

Explain Embedding quality: inter/intra distance ratio by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.

32. Something is wrong here: always using the absolute hardest negative

Anomaly

Predict first

A student writes this, and it looks reasonable:

Always select the absolute closest diff-class sample as the negative — hardest means best signal.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The globally hardest negative is often an outlier or a mislabeled sample.

Use semi-hard negatives: farther than the positive but still within the margin.

Why: The globally hardest negative is often an outlier or a mislabeled sample. Gradients from a nearly-identical but wrong-class sample can be huge and destabilize training (mode collapse: all embeddings map to the same point).

33. Trap: always using the absolute hardest negative

Trap

The trap

Always select the absolute closest diff-class sample as the negative — hardest means best signal.

neg = argmin_{n: y_n≠y_a} ‖f(a)−f(n)‖ (global closest)

Why: The globally hardest negative is often an outlier or a mislabeled sample. Gradients from a nearly-identical but wrong-class sample can be huge and destabilize training (mode collapse: all embeddings map to the same point).

The fix

Use semi-hard negatives: farther than the positive but still within the margin.

neg = n where d(a,p) < d(a,n) < d(a,p) + margin

Why: Semi-hard negatives provide a gradient without being so extreme they destabilize learning. They're the FaceNet standard: hard enough to be informative, not so hard they cause collapse.

34. Which of these survive contact with Lesson 73: Metric Learning, Siamese Networks…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Shared weights are the key: each sample routes through identical f, so similarity is measured in a consistent space.; When d_ap < d_an − m the triplet is satisfied and contributes zero loss. Only violated triplets (where the negative is too close) drive gradients.; After 300 epochs, measure the mean intra-class distance (same identity) and mean inter-class distance (different identity) across all pairs.
Breaks
To compute triplet loss, pass two inputs (anchor, positive) — just like contrastive loss on a same-class pair.; Always select the absolute closest diff-class sample as the negative — hardest means best signal.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 73: Metric Learning, Siamese Networks & Triplet Loss puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

35. Connection to attention (Lesson 65 callback)

Concept

Both metric learning and attention compute similarity scores between vectors. In dot-product attention (Lesson 65): scores = QK^T / √d — the raw score is a cosine similarity between L2-normalized query and key.

\[ \text{Attention}(Q,K,V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V \]

If Q and K are L2-normalized embeddings, Q·K^T is exactly the cosine similarity matrix — attention weights are metric-learning similarity scores routed through a softmax.

36. Break it if you can: Connection to attention (Lesson 65 callback)

Counterexample

Discussion prompt

If Q and K are L2-normalized embeddings, Q·K^T is exactly the cosine similarity matrix — attention weights are metric-learning similarity scores routed through a softmax.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

37. Rebuild the recipe: The metric learning recipe

Ranking

Put in order

These are the steps of The metric learning recipe, scrambled. Put them back in order before the next slide shows you.

  1. Architecture: Siamese EmbeddingNet with shared weights; L2-normalize output
  2. Triplets: sample (anchor a, positive p, negative n); loss = max(0, d_ap − d_an + m)
  3. Contrastive: pairs with label y; loss = y·d² + (1−y)·max(0, m−d)²
  4. Hard mining: closest same-class positive + closest diff-class negative per batch
  5. Evaluate: inter/intra distance ratio (higher = better separation)
  6. Connection: L2-normalized embedding dot product = cosine sim = attention score

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

38. The metric learning recipe

Pattern

  1. Architecture: Siamese EmbeddingNet with shared weights; L2-normalize output
  2. Triplets: sample (anchor a, positive p, negative n); loss = max(0, d_ap − d_an + m)
  3. Contrastive: pairs with label y; loss = y·d² + (1−y)·max(0, m−d)²
  4. Hard mining: closest same-class positive + closest diff-class negative per batch
  5. Evaluate: inter/intra distance ratio (higher = better separation)
  6. Connection: L2-normalized embedding dot product = cosine sim = attention score

39. Where does it stop working: The metric learning recipe

Edge cases

Discussion prompt

The metric learning recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Architecture: Siamese EmbeddingNet with shared weights; L2-normalize output
  2. Triplets: sample (anchor a, positive p, negative n); loss = max(0, d_ap − d_an + m)
  3. Contrastive: pairs with label y; loss = y·d² + (1−y)·max(0, m−d)²
  4. Hard mining: closest same-class positive + closest diff-class negative per batch
  5. Evaluate: inter/intra distance ratio (higher = better separation)
  6. Connection: L2-normalized embedding dot product = cosine sim = attention score

40. Rule out three: Check yourself — triplet loss

Elimination

Eliminate the wrong options

A triplet (a, p, n) has d(a,p)=0.8 and d(a,n)=1.1. With margin m=0.5, what is the triplet loss?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.20
  • B. 0.0
  • C. 0.50
  • D. 0.30

Survives elimination: A

Why: loss = max(0, 0.8 − 1.1 + 0.5) = max(0, 0.2) = 0.20. The triplet is violated because d_an − d_ap = 0.3, which is less than margin 0.5 — the negative is not far enough.

41. Check yourself — triplet loss

Check

Work it out before clicking.

Check your understanding

A triplet (a, p, n) has d(a,p)=0.8 and d(a,n)=1.1. With margin m=0.5, what is the triplet loss?

  • A. 0.20 (correct)
  • B. 0.0
  • C. 0.50
  • D. 0.30

Answer: A

Why: loss = max(0, 0.8 − 1.1 + 0.5) = max(0, 0.2) = 0.20. The triplet is violated because d_an − d_ap = 0.3, which is less than margin 0.5 — the negative is not far enough.

Why B tempts people
The triplet IS violated: d_an − d_ap = 0.3 < margin 0.5, so loss > 0. Zero would require d_an − d_ap ≥ 0.5.
Why C tempts people
0.50 would be the loss if d_ap = d_an (d_an − d_ap = 0): max(0, 0 + 0.5) = 0.5. Here d_an > d_ap, so the violation is partial.
Why D tempts people
0.30 = d_an − d_ap, which is the 'gap'. The loss is the DEFICIT from the margin: margin − gap = 0.5 − 0.3 = 0.2, not the gap itself.

42. Answer it before you see the options: Check yourself — hard negative mining

Prediction

Predict first

Why does online hard negative mining produce a higher epoch-1 loss than random sampling?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: It always selects the closest diff-class sample, which is the most-violated triplet

Why: Hard mining picks the negative closest to the anchor, guaranteeing a violated triplet (d_an is minimized, so d_ap − d_an + m is maximized). Random sampling mostly finds easy triplets that are already satisfied, contributing zero loss.

43. Check yourself — hard negative mining

Check

Think about what 'hard' means.

Check your understanding

Why does online hard negative mining produce a higher epoch-1 loss than random sampling?

  • A. It always selects the closest diff-class sample, which is the most-violated triplet (correct)
  • B. It uses a larger margin
  • C. The learning rate is higher
  • D. It uses more triplets per epoch

Answer: A

Why: Hard mining picks the negative closest to the anchor, guaranteeing a violated triplet (d_an is minimized, so d_ap − d_an + m is maximized). Random sampling mostly finds easy triplets that are already satisfied, contributing zero loss.

Why B tempts people
Both strategies use the same margin m=0.5 in the comparison; margin size is a hyperparameter choice, not a consequence of mining strategy.
Why C tempts people
The learning rate is identical (lr=5e-3) in both runs; the difference is entirely in which negatives are selected.
Why D tempts people
Both use the same 16 triplets per epoch; hard mining picks 16 violated triplets while random may select 16 mostly-satisfied ones.

44. Rule out three: Check yourself — metric learning & attention

Elimination

Eliminate the wrong options

With L2-normalized query Q and key K vectors, the raw attention score Q·K^T equals:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. cosine similarity — the same similarity metric used in metric learning embeddings
  • B. Euclidean distance between Q and K
  • C. a learned distance independent of the embedding norms
  • D. cross-entropy between Q and K

Survives elimination: A

Why: For unit-norm vectors, Q·K = |Q||K|cos(θ) = cos(θ). The dot product of L2-normalized embeddings IS cosine similarity — so attention weights are softmax-normalized metric learning similarity scores.

45. Check yourself — metric learning & attention

Check

Connect the two ideas.

Check your understanding

With L2-normalized query Q and key K vectors, the raw attention score Q·K^T equals:

  • A. cosine similarity — the same similarity metric used in metric learning embeddings (correct)
  • B. Euclidean distance between Q and K
  • C. a learned distance independent of the embedding norms
  • D. cross-entropy between Q and K

Answer: A

Why: For unit-norm vectors, Q·K = |Q||K|cos(θ) = cos(θ). The dot product of L2-normalized embeddings IS cosine similarity — so attention weights are softmax-normalized metric learning similarity scores.

Why B tempts people
‖Q−K‖² = 2 − 2Q·K for unit-norm vectors, so Euclidean distance is a monotone function of cosine similarity but not equal to it. The score itself is the dot product, not the distance.
Why C tempts people
The dot product is fully determined by the angle between the vectors. If both are L2-normalized, the norms are exactly 1 and 'learned' rescaling has no effect on the direction.
Why D tempts people
Cross-entropy is defined over probability distributions, not raw vectors. Attention uses cross-entropy only in the final loss (over output logits), not in the similarity score computation.

46. Your turn: Siamese face verifier

Section

Project

47. Project: Siamese EmbeddingNet with triplet loss

Concept

Build a Siamese network that learns a 4-d embedding space for a 3-identity dataset. Train it twice — once with random negatives and once with hard negatives — and compare the inter/intra distance ratio.

#milestonekey step
1EmbeddingNet + single triplet lossforward pass, L2-norm, clamp
2training loop (random negatives)300 epochs, Adam, 16 triplets/epoch
3hard negative miningper-epoch embedding snapshot, argmin/argmax

Build rules: L2-normalize outputs with F.normalize(out, dim=-1). Never pass a satisfied triplet's gradient through — use torch.clamp(..., min=0). Verify inter/intra ratio improves with hard mining.

48. Milestone 1 — EmbeddingNet + triplet forward pass

Worked example

Your turn: define EmbeddingNet (8→16→4, ReLU, L2-norm) and compute the triplet loss for one (a, p, n) tuple by hand.

Hint: F.normalize(self.net(x), dim=-1) returns unit-norm vectors. torch.norm(ea-ep) gives L2 distance. Loss = torch.clamp(dap - dan + margin, min=0).

import torch, torch.nn as nn, torch.nn.functional as F
import numpy as np
from sklearn.datasets import make_blobs

np.random.seed(42); torch.manual_seed(0)
X, y = make_blobs(150, 8, centers=3, cluster_std=4.0, random_state=42)
X = ((X - X.mean(0))/(X.std(0)+1e-6)).astype(np.float32)
Xt = torch.tensor(X)

class EmbedNet(nn.Module):
    def __init__(self): super().__init__(); self.net = nn.Sequential(nn.Linear(8,16), nn.ReLU(), nn.Linear(16,4))
    def forward(self, x): return F.normalize(self.net(x), dim=-1)

model = EmbedNet()
ai, pi, ni = 118, 106, 63   # anchor, pos, neg
ea, ep, en = model(Xt[ai:ai+1]), model(Xt[pi:pi+1]), model(Xt[ni:ni+1])
dap = torch.norm(ea-ep).item(); dan = torch.norm(ea-en).item()
loss = torch.clamp(torch.tensor(dap - dan + 0.5), min=0).item()
print(f'dap={dap:.4f}  dan={dan:.4f}  loss={loss:.4f}')
print('L2 norm ea:', round(torch.norm(ea).item(), 4))
valueresult
d(a,p)0.3478
d(a,n)1.1799
triplet loss0.0000 (satisfied)
‖ea‖1.0000 (unit norm)

49. Milestone 2 — training loop, random negatives

Worked example

Your turn: write the 300-epoch loop with 16 random triplets per step. Predict where the loss plateaus.

Hint: opt.zero_grad() once per epoch (outside the inner loop); accumulate losses across the 16 triplets, then divide and .backward(). Use Adam(lr=5e-3).

opt = torch.optim.Adam(model.parameters(), lr=5e-3)
rng = np.random.default_rng(0)

def sample_triplet(X, y, rng):
    ai = int(rng.integers(len(X))); lbl = y[ai]
    pos = np.where(y==lbl)[0]; pos = pos[pos!=ai]
    neg = np.where(y!=lbl)[0]
    return ai, int(rng.choice(pos)), int(rng.choice(neg))

for epoch in range(300):
    opt.zero_grad(); total = torch.tensor(0.0)
    for _ in range(16):
        ai,pi,ni = sample_triplet(X, y, rng)
        ea,ep,en = model(Xt[ai:ai+1]),model(Xt[pi:pi+1]),model(Xt[ni:ni+1])
        total += torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5, min=0)
    (total/16).backward(); opt.step()
print('final loss:', round((total/16).item(), 4))
epochloss (random neg)
10.2616
100.2572
500.0000
3000.0000

50. What each one costs: Milestone 2 — training loop, random negatives

Trade off

Comparison matrix

From Milestone 2 — training loop, random negatives: every row here is a choice with a cost. Fill the loss (random neg) column, then say which row you would actually pick and what you give up for it.

epochloss (random neg)
10.2616
100.2572
500.0000
3000.0000

51. Milestone 3 — hard negative mining + ratio comparison

Worked example

Your turn: re-train with hard negatives (snapshot embeddings each epoch, pick argmin diff-class distance). Predict whether hard mining reaches a higher or lower inter/intra ratio.

Hint: embs_np = model(Xt).detach().numpy() at the top of each epoch. For hard neg: neg_pool = where(y!=y_a), neg_idx = neg_pool[argmin(‖embs_np[neg_pool] − embs_np[a_i]‖)].

torch.manual_seed(0); model_h = EmbedNet()
opt_h = torch.optim.Adam(model_h.parameters(), lr=5e-3)
rng2 = np.random.default_rng(0)

for epoch in range(300):
    with torch.no_grad(): embs_np = model_h(Xt).numpy()
    opt_h.zero_grad(); total = torch.tensor(0.0)
    for _ in range(16):
        ai = int(rng2.integers(len(X))); lbl = y[ai]
        pos_pool = np.where(y==lbl)[0]; pos_pool = pos_pool[pos_pool!=ai]
        neg_pool = np.where(y!=lbl)[0]
        pi = int(pos_pool[np.argmax(np.linalg.norm(embs_np[pos_pool]-embs_np[ai], axis=1))])
        ni = int(neg_pool[np.argmin(np.linalg.norm(embs_np[neg_pool]-embs_np[ai], axis=1))])
        ea,ep,en = model_h(Xt[ai:ai+1]),model_h(Xt[pi:pi+1]),model_h(Xt[ni:ni+1])
        total += torch.clamp(torch.norm(ea-ep)-torch.norm(ea-en)+0.5, min=0)
    (total/16).backward(); opt_h.step()
print('hard final loss:', round((total/16).item(), 4))
strategyintra dinter dinter/intra ratio
random neg0.30241.69385.60×
hard neg0.18861.54148.17×

52. Fill in: inter d for Milestone 3 — hard negative mining + ratio…

Comparison

Comparison matrix

From Milestone 3 — hard negative mining + ratio comparison: refill the inter d column from what you know. The rest of the table is as it appeared.

strategyintra dinter dinter/intra ratio
random neg0.30241.69385.60×
hard neg0.18861.54148.17×

53. Show it off

Concept

Out loud, slides closed: (1) explain the triplet loss formula and why satisfied triplets contribute zero gradient, (2) describe what hard negative mining selects and why it can cause collapse without a semi-hard guard, (3) state the connection between L2-normalized embeddings and attention scores.

Stretch (homework): implement face verification with a threshold on d(f(x1), f(x2)) and plot ROC at various thresholds; compare contrastive vs triplet loss on the same 3-identity data; analyze how metric learning relates to attention using the similarity matrix Q·K^T from Lesson 65.

54. Connect it up: Lesson 73: Metric Learning, Siamese Networks & Triplet Loss

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — What is metric learning? · Triplet loss & contrastive loss · Hard negative mining · Your turn: Siamese face verifier. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

55. What you can do now

Recap

ideathe one thing to remember
triplet losssatisfied triplets give zero gradient — only violations train
Siamese netshared weights + L2-norm; no fixed class head
hard miningclosest negative = most violated = fastest signal (but semi-hard to avoid collapse)
attention linkdot product of unit-norm embeddings = cosine similarity

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 73 (Metric Learning, Siamese Networks, Triplet Loss) — Barron · USAAIO Round 2 Preparation, 2026
  2. Siamese EmbeddingNet with triplet loss and hard negative mining, inter/intra ratios verified — torch 2.7.1 + sklearn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108