USAAIO Lesson 112, from Phase 3. Mask R-CNN adds a per-proposal mask head to Faster R-CNN, running RoI features through a conv stack and a ConvTranspose2d to produce per-class binary masks; panoptic segmentation unifies the semantic and instance predictions; and self-supervised contrastive learning, in the form of SimCLR with InfoNCE, trains representations without labels by pulling augmented views of the same image together while pushing different images apart. The InfoNCE loss, NT-Xent, the temperature tau, the design of the projection head, and linear-probe evaluation were all verified with torch 2.7.1+cpu on synthetic data. The lesson runs to 29 slides.
Subject: Machine Learning · 54 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 112 · Phase 3
Mask R-CNN segments every object instance pixel-by-pixel. SimCLR learns visual representations without a single label by pulling augmented views together and pushing different images apart. Both topics verified in PyTorch with real computed shapes and losses.
Objectives
Section
Part 1 of 3
Concept
Three pixel-level vision tasks are often confused. The key axis: does the model count and separate individual objects, or assign each pixel a category label?
| task | assigns label per pixel | separates instances | example output |
|---|---|---|---|
| semantic seg. | yes | no | every road pixel = 'road' |
| instance seg. | yes (for objects) | yes | person #1 mask, person #2 mask |
| panoptic seg. | yes (all pixels) | yes (for things) | road + person #1 + person #2 |
Things (countable objects: cars, people) get per-instance masks. Stuff (amorphous regions: sky, road) only gets a semantic label. Panoptic = things + stuff.
Comparison
Comparison matrix
From Instance vs semantic vs panoptic segmentation: refill the assigns label per pixel column from what you know. The rest of the table is as it appeared.
| task | assigns label per pixel | separates instances | example output |
|---|---|---|---|
| semantic seg. | yes | no | every road pixel = 'road' |
| instance seg. | yes (for objects) | yes | person #1 mask, person #2 mask |
| panoptic seg. | yes (all pixels) | yes (for things) | road + person #1 + person #2 |
Concept
Mask R-CNN (He 2017) takes Faster R-CNN (Lesson 56) and adds a third head in parallel with the box and class heads. The mask head operates on RoI-aligned features — not RoI-pooled — which matters for pixel-level accuracy.
(N, C, 7, 7) for box/class head, (N, C, 14, 14) for mask head(N, K, 28, 28) binary mask per class KEstimation
Predict first
Trace a minimal mask head: 4 proposals, 256-channel RoI features at 14×14, 5 classes. Follow the shapes through each layer before reading on.
Commit before you compute: what does Mask head — shape trace in PyTorch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: ConvTranspose2d(64, 64, kernel=2, stride=2) doubles spatial resolution: 14 → 28
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Output size = (in_size - 1) * stride + kernel = (14-1)*2 + 2 = 28.
Worked example
Trace a minimal mask head: 4 proposals, 256-channel RoI features at 14×14, 5 classes. Follow the shapes through each layer before reading on.
import torch, torch.nn as nn
class TinyMaskHead(nn.Module):
def __init__(self, in_ch=256, n_cls=5):
super().__init__()
self.convs = nn.Sequential(
nn.Conv2d(in_ch, 64, 3, padding=1), nn.ReLU(),
nn.Conv2d(64, 64, 3, padding=1), nn.ReLU(),
)
self.deconv = nn.ConvTranspose2d(64, 64, 2, stride=2) # 14->28
self.pred = nn.Conv2d(64, n_cls, 1) # 1x1 per-class
def forward(self, x):
x = self.convs(x)
x = self.deconv(x)
return self.pred(x)
torch.manual_seed(42)
head = TinyMaskHead(in_ch=256, n_cls=5)
roi = torch.randn(4, 256, 14, 14) # 4 proposals
out = head(roi)
print(roi.shape) # torch.Size([4, 256, 14, 14])
print(out.shape) # torch.Size([4, 5, 28, 28])
print(sum(p.numel() for p in head.parameters())) # 201221ConvTranspose2d(64, 64, kernel=2, stride=2) doubles spatial resolution: 14 → 28
Why: Output size = (in_size - 1) * stride + kernel = (14-1)*2 + 2 = 28. The mask head up-samples back to 28×28 to give finer pixel detail than the 14×14 RoI feature grid.
| layer | input shape | output shape |
|---|---|---|
| RoIAlign output | (4, 256, 14, 14) | — |
| Conv2d 256→64 + Conv2d 64→64 | (4, 256, 14, 14) | (4, 64, 14, 14) |
| ConvTranspose2d stride=2 | (4, 64, 14, 14) | (4, 64, 28, 28) |
| Conv2d 1×1 (pred) | (4, 64, 28, 28) | (4, 5, 28, 28) |
| sigmoid (at inference) | (4, 5, 28, 28) | binary mask per class |
Trade off
Comparison matrix
From Mask head — shape trace in PyTorch: every row here is a choice with a cost. Fill the input shape column, then say which row you would actually pick and what you give up for it.
| layer | input shape | output shape |
|---|---|---|
| RoIAlign output | (4, 256, 14, 14) | — |
| Conv2d 256→64 + Conv2d 64→64 | (4, 256, 14, 14) | (4, 64, 14, 14) |
| ConvTranspose2d stride=2 | (4, 64, 14, 14) | (4, 64, 28, 28) |
| Conv2d 1×1 (pred) | (4, 64, 28, 28) | (4, 5, 28, 28) |
| sigmoid (at inference) | (4, 5, 28, 28) | binary mask per class |
Concept
Panoptic segmentation (Kirillov 2019) assigns every pixel either a thing label with an instance ID, or a stuff label with no instance ID. The Panoptic Quality (PQ) metric unifies detection and segmentation quality:
\[ \text{PQ} = \underbrace{\frac{\sum_{(p,g)\in TP} \text{IoU}(p,g)}{|TP|}}_{{\text{Segmentation Quality (SQ)}}} \times \underbrace{\frac{|TP|}{|TP| + \tfrac{1}{2}|FP| + \tfrac{1}{2}|FN|}}_{{\text{Recognition Quality (RQ)}}} \]
A modern approach is a single unified model (e.g. Panoptic FPN, Mask2Former) that shares a backbone and FPN between the semantic and instance branches, merging outputs in post-processing.
Counterexample
Discussion prompt
A modern approach is a single unified model (e.g. Panoptic FPN, Mask2Former) that shares a backbone and FPN between the semantic and instance branches, merging outputs in post-processing.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks.
Mask R-CNN uses RoIAlign for the mask head — it removes quantization by using bilinear interpolation at four sample points per bin.
Why: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.
Trap
Use RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.
RoIPool snaps the RoI boundary to the nearest grid cell before max-pooling
Why: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.
Mask R-CNN uses RoIAlign for the mask head — it removes quantization by using bilinear interpolation at four sample points per bin.
RoIAlign samples at sub-pixel positions; no rounding of the RoI boundary
Why: He 2017 shows that switching RoIPool → RoIAlign gives +10 to +50% relative mask AP improvement. The fractional coordinates are interpolated bilinearly from the nearest four feature-map cells.
Break the constraint
Discussion prompt
The rule this trap just fixed:
He 2017 shows that switching RoIPool → RoIAlign gives +10 to +50% relative mask AP improvement. The fractional coordinates are interpolated bilinearly from the nearest four feature-map cells.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.
Section
Part 2 of 3
Concept
Supervised pretraining needs millions of labeled images. Self-supervised pretraining creates a proxy task from the data itself — no labels required. The learned encoder can then be fine-tuned on a small labeled set.
Contrastive learning formulates the proxy task as: given two augmented views of the same image, their embeddings should be similar; embeddings of different images should be dissimilar. SimCLR (Chen 2020) is the canonical instantiation.
Concept
The InfoNCE loss (Oord 2018) treats contrastive learning as a 2N-1-way classification problem: given an anchor z_i, identify its positive z_j among all 2(N-1) negatives in the batch.
\[ \mathcal{L}_{i,j} = -\log \frac{\exp(\text{sim}(z_i, z_j) / \tau)}{\sum_{k=1, k \neq i}^{2N} \exp(\text{sim}(z_i, z_k) / \tau)} \]
sim(u, v) = uᵀv / (‖u‖ · ‖v‖) — cosine similarity after L2-normalizing both(i,j) and (j,i) directionsAnalogy
Discussion prompt
Explain InfoNCE loss — the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The InfoNCE loss (Oord 2018) treats contrastive learning as a 2N-1-way classification problem: given an anchor z_i, identify its positive z_j among all 2(N-1) negatives in the batch.
Missing information
Discussion prompt
Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Tight positives (cosine ~0.999) are easy to separate from negatives. Smaller tau sharpens the softmax, giving near-zero loss — but on hard negatives, small tau explodes gradients. SimCLR uses tau=0.1 in practice.
Worked example
Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.
import torch, torch.nn as nn
def nt_xent_loss(z1, z2, tau=0.5):
"""z1, z2: (B, D) embeddings from two augmented views."""
B = z1.shape[0]
z1 = nn.functional.normalize(z1, dim=-1)
z2 = nn.functional.normalize(z2, dim=-1)
z = torch.cat([z1, z2], dim=0) # (2B, D)
sim = z @ z.T / tau # (2B, 2B) cosine / tau
# Mask self-similarities (diagonal = same vector)
mask = torch.eye(2*B, dtype=torch.bool)
sim.masked_fill_(mask, -1e9)
# For index i in [0..B-1]: positive is i+B; for [B..2B-1]: positive is i-B
labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
return nn.CrossEntropyLoss()(sim, labels)
torch.manual_seed(42)
B, D = 6, 8
z1 = torch.randn(B, D)
z2 = z1 + 0.05 * torch.randn_like(z1) # tight positives
print(f'tau=0.07: {nt_xent_loss(z1, z2, 0.07):.4f}') # 0.0009
print(f'tau=0.50: {nt_xent_loss(z1, z2, 0.50):.4f}') # 0.8428
print(f'tau=1.00: {nt_xent_loss(z1, z2, 1.00):.4f}') # 1.4959tau=0.07 -> loss 0.0009; tau=0.50 -> 0.8428; tau=1.00 -> 1.4959
Why: Tight positives (cosine ~0.999) are easy to separate from negatives. Smaller tau sharpens the softmax, giving near-zero loss — but on hard negatives, small tau explodes gradients. SimCLR uses tau=0.1 in practice.
| tau | loss (tight pos) | loss (loose pos) | effect |
|---|---|---|---|
| 0.07 | 0.013 | 2.468 | very sharp; easy pairs score near 0, hard pairs punished hard |
| 0.10 | 0.029 | 2.131 | SimCLR default; good balance |
| 0.50 | 1.314 | 2.094 | softer; less discriminative, loss stays high even for easy pairs |
| 1.00 | 1.496 | ~2.1 | nearly uniform distribution; model gets almost no gradient signal |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
tau=0.07 -> loss 0.0009; tau=0.50 -> 0.8428; tau=1.00 -> 1.4959
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.
Concept
SimCLR has four components. Critically, the projection head g(·) is used only during pretraining — the linear probe and fine-tuning attach to the encoder output h, not z.
| component | role | discarded at fine-tune? |
|---|---|---|
| Augmentation (t, t') | Randomly crop, flip, color-jitter, grayscale — same image, two views | yes (used only at train time) |
| Encoder f(·) | ResNet-50 or small MLP; output h = f(x); this is the learned representation | no — fine-tune this |
| Projection head g(·) | MLP(h) → z; projects to space where InfoNCE is computed | yes — discard after pretraining |
| InfoNCE / NT-Xent | Contrastive loss over z pairs in the batch | yes — loss, not a module |
Why discard g? Chen 2020 Appendix B shows that representations in h transfer better than in z — the projection head learns to discard information (e.g. color, orientation) that the augmentations remove but that is useful for downstream tasks.
Analogy
Discussion prompt
Explain SimCLR pipeline: augment → encode → project → loss by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
SimCLR has four components. Critically, the projection head g(·) is used only during pretraining — the linear probe and fine-tuning attach to the encoder output h, not z.
Anomaly
Predict first
A student writes this, and it looks reasonable:
After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The projection head is trained to be invariant to augmentation-irrelevant information.
Fine-tune from the encoder output h = f(x), discarding the projection head entirely after pretraining.
Why: The projection head is trained to be invariant to augmentation-irrelevant information. Color, orientation, and texture details — dropped by augmentations — are discarded in z but are useful for, e.g., object classification. Fine-tuning from z gives significantly worse accuracy.
Trap
After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.
linear_head = nn.Linear(z_dim, n_classes); train on z
Why: Wrong. The projection head is trained to be invariant to augmentation-irrelevant information. Color, orientation, and texture details — dropped by augmentations — are discarded in z but are useful for, e.g., object classification. Fine-tuning from z gives significantly worse accuracy.
Fine-tune from the encoder output h = f(x), discarding the projection head entirely after pretraining.
linear_head = nn.Linear(h_dim, n_classes); train on h = encoder(x)
Why: The encoder f retains all information needed for downstream tasks; only the projection head learns to compress away augmentation-invariant details. Chen 2020 reports 10%+ accuracy gap when using z vs h for ImageNet linear probe.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.; After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.Estimation
Predict first
Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.
Commit before you compute: what does SimCLR architecture — encoder + projection head come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Encoder: Linear(64,32)=2112, Linear(32,16)=528 → 2640...
Worked example
Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.
import torch, torch.nn as nn
class TinyEncoder(nn.Module):
def __init__(self, in_dim=64, hidden=32, out_dim=16):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden), nn.ReLU(),
nn.Linear(hidden, out_dim)
)
def forward(self, x): return self.net(x)
class ProjectionHead(nn.Module):
def __init__(self, in_dim=16, hidden=16, out_dim=8):
super().__init__()
self.net = nn.Sequential(
nn.Linear(in_dim, hidden), nn.ReLU(),
nn.Linear(hidden, out_dim)
)
def forward(self, x): return self.net(x)
torch.manual_seed(42)
enc = TinyEncoder(64, 32, 16)
proj = ProjectionHead(16, 16, 8)
print(sum(p.numel() for p in enc.parameters()), # 2608
sum(p.numel() for p in proj.parameters())) # 408
x = torch.randn(8, 64)
h = enc(x) # (8, 16) - representation
z = proj(h) # (8, 8) - projection
print(h.shape, z.shape) # torch.Size([8, 16]) torch.Size([8, 8])Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)
Why: Encoder: Linear(64,32)=2112, Linear(32,16)=528 → 2640... wait, with biases: 6432+32=2080, 3216+16=528 → 2608. Proj: 1616+16=272, 168+8=136 → 408. Verified by running sum(p.numel()).
| module | layer | params | output shape (B=8) |
|---|---|---|---|
| TinyEncoder | Linear(64, 32) + ReLU | 64×32+32 = 2080 | (8, 32) |
| TinyEncoder | Linear(32, 16) | 32×16+16 = 528 | (8, 16) = h |
| ProjectionHead | Linear(16, 16) + ReLU | 16×16+16 = 272 | (8, 16) |
| ProjectionHead | Linear(16, 8) | 16×8+8 = 136 | (8, 8) = z |
| TOTAL | — | 2608 + 408 = 3016 | — |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.
Section
Part 3 of 3
Concept
A linear probe measures self-supervised representation quality: freeze the pretrained encoder, train only a linear classifier on top of h. If a single linear layer achieves high accuracy, the encoder has learned linearly separable representations.
Estimation
Predict first
Run SimCLR on 160 unlabeled synthetic 64-dim samples (two classes). Then attach a frozen-encoder linear probe and compare to a supervised baseline. Predict whether SSL or supervised wins.
Commit before you compute: what does SimCLR pretraining + linear probe on synthetic data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: SSL linear probe test accuracy: 62.5%; supervised baseline: 72.5%
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On this small synthetic dataset (200 samples, 60 SSL epochs), supervised wins — consistent with the data-scaling intuition: SSL learns better with more data and more epochs.
Worked example
Run SimCLR on 160 unlabeled synthetic 64-dim samples (two classes). Then attach a frozen-encoder linear probe and compare to a supervised baseline. Predict whether SSL or supervised wins.
import torch, torch.nn as nn
# --- Data: two clusters, no labels during pretraining ---
torch.manual_seed(42)
vec0 = torch.zeros(64); vec0[0] = 2.0
vec1 = torch.zeros(64); vec1[0] = -2.0
X_all = torch.cat([torch.randn(100,64)+vec0, torch.randn(100,64)+vec1])
y_all = torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm = torch.randperm(200, generator=torch.Generator().manual_seed(42))
X_tr, X_te = X_all[perm[:160]], X_all[perm[160:]]
y_tr, y_te = y_all[perm[:160]], y_all[perm[160:]]
# --- SSL pretraining (no y_tr used) ---
torch.manual_seed(42)
encoder = TinyEncoder(64, 32, 16) # defined earlier
projector= ProjectionHead(16, 16, 8)
opt_ssl = torch.optim.Adam(list(encoder.parameters())+list(projector.parameters()), lr=1e-3)
for epoch in range(60):
g1=torch.Generator().manual_seed(epoch)
g2=torch.Generator().manual_seed(epoch+100)
v1 = encoder(X_tr + 0.1*torch.randn(X_tr.shape, generator=g1))
v2 = encoder(X_tr + 0.1*torch.randn(X_tr.shape, generator=g2))
loss = nt_xent_loss(projector(v1), projector(v2), tau=0.5) # defined earlier
opt_ssl.zero_grad(); loss.backward(); opt_ssl.step()
# --- Linear probe ---
with torch.no_grad():
h_tr = encoder(X_tr); h_te = encoder(X_te)
lin = nn.Linear(16, 2)
opt_lin = torch.optim.Adam(lin.parameters(), lr=1e-2)
for _ in range(200):
opt_lin.zero_grad()
nn.CrossEntropyLoss()(lin(h_tr), y_tr).backward(); opt_lin.step()
with torch.no_grad():
acc_ssl = (lin(h_te).argmax(1)==y_te).float().mean().item()
print(f'SSL linear probe: {acc_ssl*100:.1f}%') # 62.5%SSL linear probe test accuracy: 62.5%; supervised baseline: 72.5%
Why: On this small synthetic dataset (200 samples, 60 SSL epochs), supervised wins — consistent with the data-scaling intuition: SSL learns better with more data and more epochs. At ImageNet scale, the gap closes and SSL matches supervised.
| epoch | SSL loss (NT-Xent) | note |
|---|---|---|
| 0 | 5.5636 | random init; loss ≈ log(2B) for B=160 |
| 10 | 5.1622 | encoder starts separating clusters |
| 20 | 4.6966 | representations improving |
| 30 | 4.4153 | — |
| 40 | 4.2773 | — |
| 59 | 4.2030 | converged; linear probe gives 62.5% |
Comparison
Comparison matrix
From SimCLR pretraining + linear probe on synthetic data: refill the note column from what you know. The rest of the table is as it appeared.
| epoch | SSL loss (NT-Xent) | note |
|---|---|---|
| 0 | 5.5636 | random init; loss ≈ log(2B) for B=160 |
| 10 | 5.1622 | encoder starts separating clusters |
| 20 | 4.6966 | representations improving |
| 30 | 4.4153 | — |
| 40 | 4.2773 | — |
| 59 | 4.2030 | converged; linear probe gives 62.5% |
Constraint
Discussion prompt
Run The contrastive learning recipe with this step confiscated:
NT-Xent loss: stack [z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
h_i = f(x_i_aug1), h_j = f(x_i_aug2) — both through the same encoder fz_i = g(h_i) via a 2-layer MLP projection head. L2-normalize z before loss.[z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]nn.Linear(h_dim, n_classes) — this is the linear probe. Do NOT attach to z.(N_proposals, K, 28, 28); use sigmoid per class, NOT softmax.Pattern
h_i = f(x_i_aug1), h_j = f(x_i_aug2) — both through the same encoder fz_i = g(h_i) via a 2-layer MLP projection head. L2-normalize z before loss.[z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]nn.Linear(h_dim, n_classes) — this is the linear probe. Do NOT attach to z.(N_proposals, K, 28, 28); use sigmoid per class, NOT softmax.Sorting
Sort into buckets
These are the pieces of Lesson 112: Instance Segmentation & Contrastive Learning, out of order. Put each one back under the part of the lesson it belongs to.
Elimination
Eliminate the wrong options
In NT-Xent loss, decreasing tau from 0.5 to 0.07 while keeping embeddings fixed will:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Smaller tau scales all logits up uniformly, concentrating the softmax on the highest-similarity entry. If the positive pair is only slightly more similar than the hardest negative, the model cannot score well — the sharpened distribution demands a large margin. Verified: tau=0.07 gives loss 0.0009 for tight positives (cosine ~0.999) but 2.468 for loose positives.
Check
Work through the temperature effect before clicking.
Check your understanding
In NT-Xent loss, decreasing tau from 0.5 to 0.07 while keeping embeddings fixed will:
Answer: A
Why: Smaller tau scales all logits up uniformly, concentrating the softmax on the highest-similarity entry. If the positive pair is only slightly more similar than the hardest negative, the model cannot score well — the sharpened distribution demands a large margin. Verified: tau=0.07 gives loss 0.0009 for tight positives (cosine ~0.999) but 2.468 for loose positives.
Prediction
Predict first
A Mask R-CNN mask head receives RoI-aligned features of shape (N, 256, 14, 14). After two Conv2d layers and a ConvTranspose2d(stride=2), then a 1×1 Conv2d to 80 output channels (COCO classes), what is the output shape?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (N, 80, 28, 28)
Why: ConvTranspose2d with kernel=2, stride=2 doubles spatial resolution: 14 → 28. The final 1×1 Conv2d changes channels from 64 to 80 (one mask per class) without changing spatial size. Result: (N, 80, 28, 28). At inference, sigmoid is applied per channel to get per-class binary masks.
Check
Trace the shapes before clicking.
Check your understanding
A Mask R-CNN mask head receives RoI-aligned features of shape (N, 256, 14, 14). After two Conv2d layers and a ConvTranspose2d(stride=2), then a 1×1 Conv2d to 80 output channels (COCO classes), what is the output shape?
Answer: A
Why: ConvTranspose2d with kernel=2, stride=2 doubles spatial resolution: 14 → 28. The final 1×1 Conv2d changes channels from 64 to 80 (one mask per class) without changing spatial size. Result: (N, 80, 28, 28). At inference, sigmoid is applied per channel to get per-class binary masks.
Elimination
Eliminate the wrong options
You have a SimCLR-pretrained ResNet-50 encoder f and projection head g. You want to classify 500 labeled images. Which approach gives the best accuracy?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The canonical SimCLR protocol attaches the linear probe to h = f(x), discarding the projection head g. The encoder h retains all information about the image; g is trained to be invariant to augmentation-irrelevant details, so z contains less information than h. Fine-tuning all of f+g on 500 labels (C) risks overfitting at this scale.
Check
Apply the representation quality argument.
Check your understanding
You have a SimCLR-pretrained ResNet-50 encoder f and projection head g. You want to classify 500 labeled images. Which approach gives the best accuracy?
Answer: A
Why: The canonical SimCLR protocol attaches the linear probe to h = f(x), discarding the projection head g. The encoder h retains all information about the image; g is trained to be invariant to augmentation-irrelevant details, so z contains less information than h. Fine-tuning all of f+g on 500 labels (C) risks overfitting at this scale.
Section
Project
Concept
Build a complete SimCLR training loop on synthetic 64-dim data. Three milestones: InfoNCE loss → full SimCLR pretrain → linear probe evaluation.
| # | milestone | key deliverable |
|---|---|---|
| 1 | Implement nt_xent_loss(z1, z2, tau) from scratch — no library | verify loss=0.0009 at tau=0.07 with tight positives |
| 2 | Pretrain TinyEncoder+ProjectionHead with NT-Xent for 60 epochs (no labels) | observe loss descent from ~5.56 to ~4.20 |
| 3 | Freeze encoder, train linear head on 160 labeled samples, report test accuracy | compare SSL (62.5%) vs supervised (72.5%) baseline |
Build rules: L2-normalize embeddings before any cosine similarity; mask the diagonal (self-similarity) from the (2B, 2B) matrix; attach the linear probe to h = encoder(x), NOT to z = projector(h).
Counterexample
Discussion prompt
Build a complete SimCLR training loop on synthetic 64-dim data. Three milestones: InfoNCE loss → full SimCLR pretrain → linear probe evaluation.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: implement nt_xent_loss. Key question: what are the labels in the cross-entropy call? For index i in [0..B-1], the positive is at index i+B in the 2B row; for index i in [B..2B-1], the positive is i-B.
Hint: torch.cat([torch.arange(B, 2*B), torch.arange(B)]) builds the label vector directly. After masking the diagonal with -1e9, nn.CrossEntropyLoss handles the rest.
import torch, torch.nn as nn
def nt_xent_loss(z1, z2, tau=0.5):
B = z1.shape[0]
z1 = nn.functional.normalize(z1, dim=-1)
z2 = nn.functional.normalize(z2, dim=-1)
z = torch.cat([z1, z2], dim=0) # (2B, D)
sim = z @ z.T / tau # (2B, 2B)
sim.masked_fill_(torch.eye(2*B, dtype=torch.bool), -1e9)
labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
return nn.CrossEntropyLoss()(sim, labels)
torch.manual_seed(42)
B, D = 6, 8
z1 = torch.randn(B, D)
z2 = z1 + 0.05 * torch.randn_like(z1) # tight positives
print(f'tau=0.07: {nt_xent_loss(z1, z2, 0.07):.4f}') # 0.0009
print(f'tau=0.50: {nt_xent_loss(z1, z2, 0.50):.4f}') # 0.8428
print(f'tau=1.00: {nt_xent_loss(z1, z2, 1.00):.4f}') # 1.4959| tau | loss | what it means |
|---|---|---|
| 0.07 | 0.0009 | tight positives dominate; nearly zero loss |
| 0.10 | 0.0083 | SimCLR default; slightly harder |
| 0.50 | 0.8428 | softer; non-trivial even for tight pairs |
| 1.00 | 1.4959 | almost uniform softmax; little gradient signal |
Worked example
Your turn: add data augmentation (Gaussian noise views), run 60 epochs of SSL, then attach a frozen-encoder linear probe. Predict the final test accuracy before running — is 60 epochs enough to beat supervised?
Hint: augment each batch with x + 0.1*torch.randn_like(x) twice (different seeds). After pretraining, call encoder.eval() and extract h_tr = encoder(X_tr) with torch.no_grad() before training the linear head.
import torch, torch.nn as nn
# (TinyEncoder, ProjectionHead, nt_xent_loss defined in earlier milestones)
# --- Data ---
torch.manual_seed(42)
vec0 = torch.zeros(64); vec0[0] = 2.0
vec1 = torch.zeros(64); vec1[0] = -2.0
X = torch.cat([torch.randn(100,64)+vec0, torch.randn(100,64)+vec1])
y = torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm = torch.randperm(200, generator=torch.Generator().manual_seed(42))
X_tr, X_te = X[perm[:160]], X[perm[160:]]
y_tr, y_te = y[perm[:160]], y[perm[160:]]
# --- SSL pretrain ---
torch.manual_seed(42)
enc = TinyEncoder(64, 32, 16); prj = ProjectionHead(16, 16, 8)
opt = torch.optim.Adam(list(enc.parameters())+list(prj.parameters()), lr=1e-3)
for epoch in range(60):
g1 = torch.Generator().manual_seed(epoch)
g2 = torch.Generator().manual_seed(epoch+100)
v1 = enc(X_tr + 0.1*torch.randn(X_tr.shape, generator=g1))
v2 = enc(X_tr + 0.1*torch.randn(X_tr.shape, generator=g2))
l = nt_xent_loss(prj(v1), prj(v2), tau=0.5)
opt.zero_grad(); l.backward(); opt.step()
# --- Linear probe ---
with torch.no_grad():
h_tr = enc(X_tr); h_te = enc(X_te) # attach to h, NOT z
lin = nn.Linear(16, 2)
opt2 = torch.optim.Adam(lin.parameters(), lr=1e-2)
for _ in range(200):
opt2.zero_grad()
nn.CrossEntropyLoss()(lin(h_tr),y_tr).backward(); opt2.step()
with torch.no_grad():
acc = (lin(h_te).argmax(1)==y_te).float().mean()
print(f'SSL linear probe: {acc*100:.1f}%') # 62.5%| method | test accuracy | labels used in training? |
|---|---|---|
| SSL pretrain + linear probe | 62.5% | no (pretraining); yes (linear head only) |
| Supervised MLP (same arch) | 72.5% | yes (full training) |
| SimCLR on ImageNet (ResNet-50) | 76.5% | no (pretraining); yes (linear probe) |
Trade off
Comparison matrix
From Milestones 2 & 3 — pretrain + linear probe: every row here is a choice with a cost. Fill the labels used in training? column, then say which row you would actually pick and what you give up for it.
| method | test accuracy | labels used in training? |
|---|---|---|
| SSL pretrain + linear probe | 62.5% | no (pretraining); yes (linear head only) |
| Supervised MLP (same arch) | 72.5% | yes (full training) |
| SimCLR on ImageNet (ResNet-50) | 76.5% | no (pretraining); yes (linear probe) |
Concept
import torch, torch.nn as nn
class TinyEncoder(nn.Module):
def __init__(self, in_d=64, hid=32, out_d=16):
super().__init__()
self.net = nn.Sequential(nn.Linear(in_d,hid), nn.ReLU(), nn.Linear(hid,out_d))
def forward(self, x): return self.net(x)
class ProjectionHead(nn.Module):
def __init__(self, in_d=16, hid=16, out_d=8):
super().__init__()
self.net = nn.Sequential(nn.Linear(in_d,hid), nn.ReLU(), nn.Linear(hid,out_d))
def forward(self, x): return self.net(x)
def nt_xent(z1, z2, tau=0.5):
B = z1.shape[0]
z1 = nn.functional.normalize(z1, dim=-1)
z2 = nn.functional.normalize(z2, dim=-1)
z = torch.cat([z1, z2])
sim = z @ z.T / tau
sim.masked_fill_(torch.eye(2*B, dtype=torch.bool), -1e9)
labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
return nn.CrossEntropyLoss()(sim, labels)
torch.manual_seed(42)
v0=torch.zeros(64); v0[0]=2.0; v1=torch.zeros(64); v1[0]=-2.0
X=torch.cat([torch.randn(100,64)+v0, torch.randn(100,64)+v1])
y=torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm=torch.randperm(200, generator=torch.Generator().manual_seed(42))
Xt,Xe,yt,ye=X[perm[:160]],X[perm[160:]],y[perm[:160]],y[perm[160:]]
torch.manual_seed(42)
enc=TinyEncoder(); prj=ProjectionHead()
opt=torch.optim.Adam(list(enc.parameters())+list(prj.parameters()), lr=1e-3)
for ep in range(60):
g1=torch.Generator().manual_seed(ep)
g2=torch.Generator().manual_seed(ep+100)
l=nt_xent(prj(enc(Xt+0.1*torch.randn(Xt.shape,generator=g1))),
prj(enc(Xt+0.1*torch.randn(Xt.shape,generator=g2))), tau=0.5)
opt.zero_grad(); l.backward(); opt.step()
if ep%10==0 or ep==59: print(f'ep {ep}: {l.item():.4f}')
with torch.no_grad(): ht=enc(Xt); he=enc(Xe)
lin=nn.Linear(16,2); opt2=torch.optim.Adam(lin.parameters(),lr=1e-2)
for _ in range(200):
opt2.zero_grad(); nn.CrossEntropyLoss()(lin(ht),yt).backward(); opt2.step()
with torch.no_grad():
print(f'Linear probe: {(lin(he).argmax(1)==ye).float().mean()*100:.1f}%') # 62.5%| design choice | value | rationale |
|---|---|---|
| tau (temperature) | 0.5 (training), 0.07 (tight-pair demo) | 0.5 trains stably; 0.07 is SimCLR paper default |
| augmentation | Gaussian noise std=0.1 | minimal; real SimCLR uses crop+flip+color jitter |
| projection head | 16→16→8 with ReLU | discard after pretraining; probe on h, not z |
| linear probe lr | 1e-2 (Adam) | much higher than SSL lr; only 2 parameters per class |
| SSL vs supervised gap | 62.5% vs 72.5% | closes at larger N and more epochs (scale is key) |
Comparison
Comparison matrix
From The full program: refill the rationale column from what you know. The rest of the table is as it appeared.
| design choice | value | rationale |
|---|---|---|
| tau (temperature) | 0.5 (training), 0.07 (tight-pair demo) | 0.5 trains stably; 0.07 is SimCLR paper default |
| augmentation | Gaussian noise std=0.1 | minimal; real SimCLR uses crop+flip+color jitter |
| projection head | 16→16→8 with ReLU | discard after pretraining; probe on h, not z |
| linear probe lr | 1e-2 (Adam) | much higher than SSL lr; only 2 parameters per class |
| SSL vs supervised gap | 62.5% vs 72.5% | closes at larger N and more epochs (scale is key) |
Concept
Out loud, slides closed: (1) describe the full SimCLR pipeline from a raw batch of images to a scalar NT-Xent loss, naming every tensor shape; (2) explain why you attach the linear probe to h = f(x) and not to z = g(f(x)); (3) explain what happens to NT-Xent loss as tau → 0 and as tau → ∞.
Stretch (from the lesson plan): implement the InfoNCE mutual information lower bound — show that minimizing InfoNCE is equivalent to maximizing a lower bound on I(h; x). Also: run the linear probe for multiple label fractions (10%, 25%, 50%, 100%) and plot SSL vs supervised accuracy as a function of label count.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Mask R-CNN — instance segmentation · Contrastive learning — SimCLR & InfoNCE · Linear probe & self-supervised evaluation · Your turn: SimCLR from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
(N, 256, 14, 14) → conv stack → ConvTranspose2d → (N, K, 28, 28) — and explain why RoIAlign is required instead of RoIPool(2B, D) → cosine sim matrix / tau → mask diagonal → cross-entropy with symmetric labels| concept | the one thing to remember |
|---|---|
| Mask R-CNN | Faster R-CNN + RoIAlign + conv→deconv mask head; K masks output |
| Panoptic seg. | things (instances) + stuff (semantic); metric = PQ = SQ × RQ |
| InfoNCE / NT-Xent | −log(exp(sim(zi,zj)/τ) / Σk exp(sim(zi,zk)/τ)); labels = positive index |
| tau | smaller = sharper softmax; SimCLR uses 0.1; too small → gradient explosion |
| Linear probe | freeze f, train nn.Linear on h (not z); measures representation quality |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.