Lesson 112: Instance Segmentation & Contrastive Learning

USAAIO Lesson 112, from Phase 3. Mask R-CNN adds a per-proposal mask head to Faster R-CNN, running RoI features through a conv stack and a ConvTranspose2d to produce per-class binary masks; panoptic segmentation unifies the semantic and instance predictions; and self-supervised contrastive learning, in the form of SimCLR with InfoNCE, trains representations without labels by pulling augmented views of the same image together while pushing different images apart. The InfoNCE loss, NT-Xent, the temperature tau, the design of the projection head, and linear-probe evaluation were all verified with torch 2.7.1+cpu on synthetic data. The lesson runs to 29 slides.

Subject: Machine Learning · 54 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Instance Segmentation & Contrastive Learning

Title

USAAIO · Lesson 112 · Phase 3

Mask R-CNN segments every object instance pixel-by-pixel. SimCLR learns visual representations without a single label by pulling augmented views together and pushing different images apart. Both topics verified in PyTorch with real computed shapes and losses.

2. By the end of this lesson you can

Objectives

  1. Explain how Mask R-CNN extends Faster R-CNN with a mask head and trace RoI feature shapes through it
  2. Distinguish instance segmentation from semantic segmentation and from panoptic segmentation
  3. Write the InfoNCE / NT-Xent loss from scratch and explain the role of temperature tau
  4. Describe the SimCLR pipeline: augmentation → encoder → projection head → contrastive loss
  5. Evaluate self-supervised representations via a linear probe and compare to supervised baselines

3. Mask R-CNN — instance segmentation

Section

Part 1 of 3

4. Instance vs semantic vs panoptic segmentation

Concept

Three pixel-level vision tasks are often confused. The key axis: does the model count and separate individual objects, or assign each pixel a category label?

taskassigns label per pixelseparates instancesexample output
semantic seg.yesnoevery road pixel = 'road'
instance seg.yes (for objects)yesperson #1 mask, person #2 mask
panoptic seg.yes (all pixels)yes (for things)road + person #1 + person #2

Things (countable objects: cars, people) get per-instance masks. Stuff (amorphous regions: sky, road) only gets a semantic label. Panoptic = things + stuff.

5. Fill in: assigns label per pixel for Instance vs semantic vs panoptic segmentation

Comparison

Comparison matrix

From Instance vs semantic vs panoptic segmentation: refill the assigns label per pixel column from what you know. The rest of the table is as it appeared.

taskassigns label per pixelseparates instancesexample output
semantic seg.yesnoevery road pixel = 'road'
instance seg.yes (for objects)yesperson #1 mask, person #2 mask
panoptic seg.yes (all pixels)yes (for things)road + person #1 + person #2

6. Mask R-CNN: Faster R-CNN + mask head

Concept

Mask R-CNN (He 2017) takes Faster R-CNN (Lesson 56) and adds a third head in parallel with the box and class heads. The mask head operates on RoI-aligned features — not RoI-pooled — which matters for pixel-level accuracy.

7. Guess the shape of the answer: Mask head — shape trace in PyTorch

Estimation

Predict first

Trace a minimal mask head: 4 proposals, 256-channel RoI features at 14×14, 5 classes. Follow the shapes through each layer before reading on.

Commit before you compute: what does Mask head — shape trace in PyTorch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: ConvTranspose2d(64, 64, kernel=2, stride=2) doubles spatial resolution: 14 → 28

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Output size = (in_size - 1) * stride + kernel = (14-1)*2 + 2 = 28.

8. Mask head — shape trace in PyTorch

Worked example

Trace a minimal mask head: 4 proposals, 256-channel RoI features at 14×14, 5 classes. Follow the shapes through each layer before reading on.

import torch, torch.nn as nn

class TinyMaskHead(nn.Module):
    def __init__(self, in_ch=256, n_cls=5):
        super().__init__()
        self.convs = nn.Sequential(
            nn.Conv2d(in_ch, 64, 3, padding=1), nn.ReLU(),
            nn.Conv2d(64,  64, 3, padding=1), nn.ReLU(),
        )
        self.deconv = nn.ConvTranspose2d(64, 64, 2, stride=2)  # 14->28
        self.pred   = nn.Conv2d(64, n_cls, 1)                   # 1x1 per-class
    def forward(self, x):
        x = self.convs(x)
        x = self.deconv(x)
        return self.pred(x)

torch.manual_seed(42)
head = TinyMaskHead(in_ch=256, n_cls=5)
roi  = torch.randn(4, 256, 14, 14)   # 4 proposals
out  = head(roi)
print(roi.shape)  # torch.Size([4, 256, 14, 14])
print(out.shape)  # torch.Size([4, 5, 28, 28])
print(sum(p.numel() for p in head.parameters()))  # 201221

ConvTranspose2d(64, 64, kernel=2, stride=2) doubles spatial resolution: 14 → 28

Why: Output size = (in_size - 1) * stride + kernel = (14-1)*2 + 2 = 28. The mask head up-samples back to 28×28 to give finer pixel detail than the 14×14 RoI feature grid.

layerinput shapeoutput shape
RoIAlign output(4, 256, 14, 14)—
Conv2d 256→64 + Conv2d 64→64(4, 256, 14, 14)(4, 64, 14, 14)
ConvTranspose2d stride=2(4, 64, 14, 14)(4, 64, 28, 28)
Conv2d 1×1 (pred)(4, 64, 28, 28)(4, 5, 28, 28)
sigmoid (at inference)(4, 5, 28, 28)binary mask per class

9. What each one costs: Mask head — shape trace in PyTorch

Trade off

Comparison matrix

From Mask head — shape trace in PyTorch: every row here is a choice with a cost. Fill the input shape column, then say which row you would actually pick and what you give up for it.

layerinput shapeoutput shape
RoIAlign output(4, 256, 14, 14)—
Conv2d 256→64 + Conv2d 64→64(4, 256, 14, 14)(4, 64, 14, 14)
ConvTranspose2d stride=2(4, 64, 14, 14)(4, 64, 28, 28)
Conv2d 1×1 (pred)(4, 64, 28, 28)(4, 5, 28, 28)
sigmoid (at inference)(4, 5, 28, 28)binary mask per class

10. Panoptic segmentation

Concept

Panoptic segmentation (Kirillov 2019) assigns every pixel either a thing label with an instance ID, or a stuff label with no instance ID. The Panoptic Quality (PQ) metric unifies detection and segmentation quality:

\[ \text{PQ} = \underbrace{\frac{\sum_{(p,g)\in TP} \text{IoU}(p,g)}{|TP|}}_{{\text{Segmentation Quality (SQ)}}} \times \underbrace{\frac{|TP|}{|TP| + \tfrac{1}{2}|FP| + \tfrac{1}{2}|FN|}}_{{\text{Recognition Quality (RQ)}}} \]

A modern approach is a single unified model (e.g. Panoptic FPN, Mask2Former) that shares a backbone and FPN between the semantic and instance branches, merging outputs in post-processing.

11. Break it if you can: Panoptic segmentation

Counterexample

Discussion prompt

A modern approach is a single unified model (e.g. Panoptic FPN, Mask2Former) that shares a backbone and FPN between the semantic and instance branches, merging outputs in post-processing.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

12. Something is wrong here: RoIPool vs RoIAlign in Mask R-CNN

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks.

Mask R-CNN uses RoIAlign for the mask head — it removes quantization by using bilinear interpolation at four sample points per bin.

Why: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.

13. Trap: RoIPool vs RoIAlign in Mask R-CNN

Trap

The trap

Use RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.

RoIPool snaps the RoI boundary to the nearest grid cell before max-pooling

Why: Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.

The fix

Mask R-CNN uses RoIAlign for the mask head — it removes quantization by using bilinear interpolation at four sample points per bin.

RoIAlign samples at sub-pixel positions; no rounding of the RoI boundary

Why: He 2017 shows that switching RoIPool → RoIAlign gives +10 to +50% relative mask AP improvement. The fractional coordinates are interpolated bilinearly from the nearest four feature-map cells.

14. Break it on purpose: RoIPool vs RoIAlign in Mask R-CNN

Break the constraint

Discussion prompt

The rule this trap just fixed:

He 2017 shows that switching RoIPool → RoIAlign gives +10 to +50% relative mask AP improvement. The fractional coordinates are interpolated bilinearly from the nearest four feature-map cells.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Wrong for masks. The quantization error (up to one full cell ~ 16px at stride 16) destroys the pixel-alignment needed for accurate masks. Detection boxes are tolerant of a few pixels of shift; binary masks are not.

15. Contrastive learning — SimCLR & InfoNCE

Section

Part 2 of 3

16. Self-supervised: learning without labels

Concept

Supervised pretraining needs millions of labeled images. Self-supervised pretraining creates a proxy task from the data itself — no labels required. The learned encoder can then be fine-tuned on a small labeled set.

Contrastive learning formulates the proxy task as: given two augmented views of the same image, their embeddings should be similar; embeddings of different images should be dissimilar. SimCLR (Chen 2020) is the canonical instantiation.

17. InfoNCE loss — the math

Concept

The InfoNCE loss (Oord 2018) treats contrastive learning as a 2N-1-way classification problem: given an anchor z_i, identify its positive z_j among all 2(N-1) negatives in the batch.

\[ \mathcal{L}_{i,j} = -\log \frac{\exp(\text{sim}(z_i, z_j) / \tau)}{\sum_{k=1, k \neq i}^{2N} \exp(\text{sim}(z_i, z_k) / \tau)} \]

18. By analogy: InfoNCE loss — the math

Analogy

Discussion prompt

Explain InfoNCE loss — the math by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The InfoNCE loss (Oord 2018) treats contrastive learning as a 2N-1-way classification problem: given an anchor z_i, identify its positive z_j among all 2(N-1) negatives in the batch.

19. What has to be given first: InfoNCE from scratch — code + trace

Missing information

Discussion prompt

Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Tight positives (cosine ~0.999) are easy to separate from negatives. Smaller tau sharpens the softmax, giving near-zero loss — but on hard negatives, small tau explodes gradients. SimCLR uses tau=0.1 in practice.

20. InfoNCE from scratch — code + trace

Worked example

Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.

import torch, torch.nn as nn

def nt_xent_loss(z1, z2, tau=0.5):
    """z1, z2: (B, D) embeddings from two augmented views."""
    B = z1.shape[0]
    z1 = nn.functional.normalize(z1, dim=-1)
    z2 = nn.functional.normalize(z2, dim=-1)
    z  = torch.cat([z1, z2], dim=0)      # (2B, D)
    sim = z @ z.T / tau                   # (2B, 2B) cosine / tau
    # Mask self-similarities (diagonal = same vector)
    mask = torch.eye(2*B, dtype=torch.bool)
    sim.masked_fill_(mask, -1e9)
    # For index i in [0..B-1]: positive is i+B; for [B..2B-1]: positive is i-B
    labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
    return nn.CrossEntropyLoss()(sim, labels)

torch.manual_seed(42)
B, D = 6, 8
z1 = torch.randn(B, D)
z2 = z1 + 0.05 * torch.randn_like(z1)  # tight positives
print(f'tau=0.07: {nt_xent_loss(z1, z2, 0.07):.4f}')   # 0.0009
print(f'tau=0.50: {nt_xent_loss(z1, z2, 0.50):.4f}')   # 0.8428
print(f'tau=1.00: {nt_xent_loss(z1, z2, 1.00):.4f}')   # 1.4959

tau=0.07 -> loss 0.0009; tau=0.50 -> 0.8428; tau=1.00 -> 1.4959

Why: Tight positives (cosine ~0.999) are easy to separate from negatives. Smaller tau sharpens the softmax, giving near-zero loss — but on hard negatives, small tau explodes gradients. SimCLR uses tau=0.1 in practice.

tauloss (tight pos)loss (loose pos)effect
0.070.0132.468very sharp; easy pairs score near 0, hard pairs punished hard
0.100.0292.131SimCLR default; good balance
0.501.3142.094softer; less discriminative, loss stays high even for easy pairs
1.001.496~2.1nearly uniform distribution; model gets almost no gradient signal

21. Work backwards from the answer: InfoNCE from scratch — code + trace

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

tau=0.07 -> loss 0.0009; tau=0.50 -> 0.8428; tau=1.00 -> 1.4959

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Implement NT-Xent (symmetric InfoNCE) for a batch of B=6 pairs. Both views are stacked into a 2B×D matrix; the similarity matrix is (2B, 2B). Verify the loss numerically.

22. SimCLR pipeline: augment → encode → project → loss

Concept

SimCLR has four components. Critically, the projection head g(·) is used only during pretraining — the linear probe and fine-tuning attach to the encoder output h, not z.

componentrolediscarded at fine-tune?
Augmentation (t, t')Randomly crop, flip, color-jitter, grayscale — same image, two viewsyes (used only at train time)
Encoder f(·)ResNet-50 or small MLP; output h = f(x); this is the learned representationno — fine-tune this
Projection head g(·)MLP(h) → z; projects to space where InfoNCE is computedyes — discard after pretraining
InfoNCE / NT-XentContrastive loss over z pairs in the batchyes — loss, not a module

Why discard g? Chen 2020 Appendix B shows that representations in h transfer better than in z — the projection head learns to discard information (e.g. color, orientation) that the augmentations remove but that is useful for downstream tasks.

23. By analogy: SimCLR pipeline: augment → encode → project → loss

Analogy

Discussion prompt

Explain SimCLR pipeline: augment → encode → project → loss by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

SimCLR has four components. Critically, the projection head g(·) is used only during pretraining — the linear probe and fine-tuning attach to the encoder output h, not z.

24. Something is wrong here: fine-tuning from the projection head output z

Anomaly

Predict first

A student writes this, and it looks reasonable:

After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The projection head is trained to be invariant to augmentation-irrelevant information.

Fine-tune from the encoder output h = f(x), discarding the projection head entirely after pretraining.

Why: The projection head is trained to be invariant to augmentation-irrelevant information. Color, orientation, and texture details — dropped by augmentations — are discarded in z but are useful for, e.g., object classification. Fine-tuning from z gives significantly worse accuracy.

25. Trap: fine-tuning from the projection head output z

Trap

The trap

After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.

linear_head = nn.Linear(z_dim, n_classes); train on z

Why: Wrong. The projection head is trained to be invariant to augmentation-irrelevant information. Color, orientation, and texture details — dropped by augmentations — are discarded in z but are useful for, e.g., object classification. Fine-tuning from z gives significantly worse accuracy.

The fix

Fine-tune from the encoder output h = f(x), discarding the projection head entirely after pretraining.

linear_head = nn.Linear(h_dim, n_classes); train on h = encoder(x)

Why: The encoder f retains all information needed for downstream tasks; only the projection head learns to compress away augmentation-invariant details. Chen 2020 reports 10%+ accuracy gap when using z vs h for ImageNet linear probe.

26. Which of these survive contact with Lesson 112: Instance Segmentation &…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Three pixel-level vision tasks are often confused. The key axis: does the model count and separate individual objects, or assign each pixel a category label?; Build a complete SimCLR training loop on synthetic 64-dim data. Three milestones: InfoNCE loss → full SimCLR pretrain → linear probe evaluation.
Breaks
Use RoIPool for the mask head — it is the standard component from Faster R-CNN and computes the same output.; After pretraining SimCLR, fine-tune a linear head on top of the projection head output z = g(f(x)) — the features that the contrastive loss was trained on.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 112: Instance Segmentation & Contrastive Learning puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

27. Guess the shape of the answer: SimCLR architecture — encoder + projection…

Estimation

Predict first

Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.

Commit before you compute: what does SimCLR architecture — encoder + projection head come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Encoder: Linear(64,32)=2112, Linear(32,16)=528 → 2640...

28. SimCLR architecture — encoder + projection head

Worked example

Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.

import torch, torch.nn as nn

class TinyEncoder(nn.Module):
    def __init__(self, in_dim=64, hidden=32, out_dim=16):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, hidden), nn.ReLU(),
            nn.Linear(hidden, out_dim)
        )
    def forward(self, x): return self.net(x)

class ProjectionHead(nn.Module):
    def __init__(self, in_dim=16, hidden=16, out_dim=8):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(in_dim, hidden), nn.ReLU(),
            nn.Linear(hidden, out_dim)
        )
    def forward(self, x): return self.net(x)

torch.manual_seed(42)
enc  = TinyEncoder(64, 32, 16)
proj = ProjectionHead(16, 16, 8)
print(sum(p.numel() for p in enc.parameters()),   # 2608
      sum(p.numel() for p in proj.parameters()))  # 408

x   = torch.randn(8, 64)
h   = enc(x)              # (8, 16) - representation
z   = proj(h)             # (8, 8)  - projection
print(h.shape, z.shape)   # torch.Size([8, 16]) torch.Size([8, 8])

Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)

Why: Encoder: Linear(64,32)=2112, Linear(32,16)=528 → 2640... wait, with biases: 6432+32=2080, 3216+16=528 → 2608. Proj: 1616+16=272, 168+8=136 → 408. Verified by running sum(p.numel()).

modulelayerparamsoutput shape (B=8)
TinyEncoderLinear(64, 32) + ReLU64×32+32 = 2080(8, 32)
TinyEncoderLinear(32, 16)32×16+16 = 528(8, 16) = h
ProjectionHeadLinear(16, 16) + ReLU16×16+16 = 272(8, 16)
ProjectionHeadLinear(16, 8)16×8+8 = 136(8, 8) = z
TOTAL—2608 + 408 = 3016—

29. Work backwards from the answer: SimCLR architecture — encoder + projection…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Encoder: 2608 params; projection head: 408 params; h shape (8,16), z shape (8,8)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Build a minimal SimCLR (encoder + projection head) on synthetic 64-dim inputs. Verify parameter counts and embedding shapes before running.

30. Linear probe & self-supervised evaluation

Section

Part 3 of 3

31. Linear probe: measuring representation quality

Concept

A linear probe measures self-supervised representation quality: freeze the pretrained encoder, train only a linear classifier on top of h. If a single linear layer achieves high accuracy, the encoder has learned linearly separable representations.

32. Guess the shape of the answer: SimCLR pretraining + linear probe on…

Estimation

Predict first

Run SimCLR on 160 unlabeled synthetic 64-dim samples (two classes). Then attach a frozen-encoder linear probe and compare to a supervised baseline. Predict whether SSL or supervised wins.

Commit before you compute: what does SimCLR pretraining + linear probe on synthetic data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: SSL linear probe test accuracy: 62.5%; supervised baseline: 72.5%

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. On this small synthetic dataset (200 samples, 60 SSL epochs), supervised wins — consistent with the data-scaling intuition: SSL learns better with more data and more epochs.

33. SimCLR pretraining + linear probe on synthetic data

Worked example

Run SimCLR on 160 unlabeled synthetic 64-dim samples (two classes). Then attach a frozen-encoder linear probe and compare to a supervised baseline. Predict whether SSL or supervised wins.

import torch, torch.nn as nn

# --- Data: two clusters, no labels during pretraining ---
torch.manual_seed(42)
vec0 = torch.zeros(64); vec0[0] = 2.0
vec1 = torch.zeros(64); vec1[0] = -2.0
X_all = torch.cat([torch.randn(100,64)+vec0, torch.randn(100,64)+vec1])
y_all = torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm  = torch.randperm(200, generator=torch.Generator().manual_seed(42))
X_tr, X_te = X_all[perm[:160]], X_all[perm[160:]]
y_tr, y_te = y_all[perm[:160]], y_all[perm[160:]]

# --- SSL pretraining (no y_tr used) ---
torch.manual_seed(42)
encoder  = TinyEncoder(64, 32, 16)   # defined earlier
projector= ProjectionHead(16, 16, 8)
opt_ssl  = torch.optim.Adam(list(encoder.parameters())+list(projector.parameters()), lr=1e-3)
for epoch in range(60):
    g1=torch.Generator().manual_seed(epoch)
    g2=torch.Generator().manual_seed(epoch+100)
    v1 = encoder(X_tr + 0.1*torch.randn(X_tr.shape, generator=g1))
    v2 = encoder(X_tr + 0.1*torch.randn(X_tr.shape, generator=g2))
    loss = nt_xent_loss(projector(v1), projector(v2), tau=0.5)  # defined earlier
    opt_ssl.zero_grad(); loss.backward(); opt_ssl.step()

# --- Linear probe ---
with torch.no_grad():
    h_tr = encoder(X_tr); h_te = encoder(X_te)
lin = nn.Linear(16, 2)
opt_lin = torch.optim.Adam(lin.parameters(), lr=1e-2)
for _ in range(200):
    opt_lin.zero_grad()
    nn.CrossEntropyLoss()(lin(h_tr), y_tr).backward(); opt_lin.step()
with torch.no_grad():
    acc_ssl = (lin(h_te).argmax(1)==y_te).float().mean().item()
print(f'SSL linear probe: {acc_ssl*100:.1f}%')   # 62.5%

SSL linear probe test accuracy: 62.5%; supervised baseline: 72.5%

Why: On this small synthetic dataset (200 samples, 60 SSL epochs), supervised wins — consistent with the data-scaling intuition: SSL learns better with more data and more epochs. At ImageNet scale, the gap closes and SSL matches supervised.

epochSSL loss (NT-Xent)note
05.5636random init; loss ≈ log(2B) for B=160
105.1622encoder starts separating clusters
204.6966representations improving
304.4153—
404.2773—
594.2030converged; linear probe gives 62.5%

34. Fill in: note for SimCLR pretraining + linear probe on…

Comparison

Comparison matrix

From SimCLR pretraining + linear probe on synthetic data: refill the note column from what you know. The rest of the table is as it appeared.

epochSSL loss (NT-Xent)note
05.5636random init; loss ≈ log(2B) for B=160
105.1622encoder starts separating clusters
204.6966representations improving
304.4153—
404.2773—
594.2030converged; linear probe gives 62.5%

35. Without one step: The contrastive learning recipe

Constraint

Discussion prompt

Run The contrastive learning recipe with this step confiscated:

NT-Xent loss: stack [z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Two augmented views of each image: random crop + flip + color-jitter + grayscale. Both views have the same label (unknown); different images are negatives.
  2. Encode: h_i = f(x_i_aug1), h_j = f(x_i_aug2) — both through the same encoder f
  3. Project: z_i = g(h_i) via a 2-layer MLP projection head. L2-normalize z before loss.
  4. NT-Xent loss: stack [z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]
  5. Temperature tau: use 0.1 (SimCLR default). Too small → exploding gradients on hard negatives; too large → uniform distribution, no gradient signal.
  6. Evaluate: freeze encoder, train nn.Linear(h_dim, n_classes) — this is the linear probe. Do NOT attach to z.
  7. Mask R-CNN checklist: RoIAlign (not RoIPool) for mask branch; mask head output shape (N_proposals, K, 28, 28); use sigmoid per class, NOT softmax.

36. The contrastive learning recipe

Pattern

  1. Two augmented views of each image: random crop + flip + color-jitter + grayscale. Both views have the same label (unknown); different images are negatives.
  2. Encode: h_i = f(x_i_aug1), h_j = f(x_i_aug2) — both through the same encoder f
  3. Project: z_i = g(h_i) via a 2-layer MLP projection head. L2-normalize z before loss.
  4. NT-Xent loss: stack [z1; z2] into (2B, D), compute sim = z @ zT / tau, mask diagonal, cross-entropy with labels [B..2B-1, 0..B-1]
  5. Temperature tau: use 0.1 (SimCLR default). Too small → exploding gradients on hard negatives; too large → uniform distribution, no gradient signal.
  6. Evaluate: freeze encoder, train nn.Linear(h_dim, n_classes) — this is the linear probe. Do NOT attach to z.
  7. Mask R-CNN checklist: RoIAlign (not RoIPool) for mask branch; mask head output shape (N_proposals, K, 28, 28); use sigmoid per class, NOT softmax.

37. Where does each piece belong: Lesson 112: Instance Segmentation &…

Sorting

Sort into buckets

These are the pieces of Lesson 112: Instance Segmentation & Contrastive Learning, out of order. Put each one back under the part of the lesson it belongs to.

Mask R-CNN — instance segmentation
Instance vs semantic vs panoptic segmentation; Mask R-CNN: Faster R-CNN + mask head; Mask head — shape trace in PyTorch
Contrastive learning — SimCLR & InfoNCE
Self-supervised: learning without labels; InfoNCE loss — the math; InfoNCE from scratch — code + trace
Linear probe & self-supervised evaluation
Linear probe: measuring representation quality; SimCLR pretraining + linear probe on synthetic data; The contrastive learning recipe
s1
Mask R-CNN — instance segmentation is where Lesson 112: Instance Segmentation & Contrastive Learning puts Instance vs semantic vs panoptic segmentation, Mask R-CNN: Faster R-CNN + mask head, Mask head — shape trace in PyTorch. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Contrastive learning — SimCLR & InfoNCE is where Lesson 112: Instance Segmentation & Contrastive Learning puts Self-supervised: learning without labels, InfoNCE loss — the math, InfoNCE from scratch — code + trace. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Linear probe & self-supervised evaluation is where Lesson 112: Instance Segmentation & Contrastive Learning puts Linear probe: measuring representation quality, SimCLR pretraining + linear probe on synthetic data, The contrastive learning recipe. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

38. Rule out three: Check — InfoNCE temperature

Elimination

Eliminate the wrong options

In NT-Xent loss, decreasing tau from 0.5 to 0.07 while keeping embeddings fixed will:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Sharpen the softmax, making it harder for the model to achieve low loss unless the positive pair has much higher cosine similarity than all negatives
  • B. Reduce the loss to zero automatically, since dividing by a smaller tau amplifies the positive score disproportionately
  • C. Have no effect on the loss, because the temperature cancels in the log-softmax
  • D. Soften the distribution, spreading probability mass more evenly across all 2N-1 entries

Survives elimination: A

Why: Smaller tau scales all logits up uniformly, concentrating the softmax on the highest-similarity entry. If the positive pair is only slightly more similar than the hardest negative, the model cannot score well — the sharpened distribution demands a large margin. Verified: tau=0.07 gives loss 0.0009 for tight positives (cosine ~0.999) but 2.468 for loose positives.

39. Check — InfoNCE temperature

Check

Work through the temperature effect before clicking.

Check your understanding

In NT-Xent loss, decreasing tau from 0.5 to 0.07 while keeping embeddings fixed will:

  • A. Sharpen the softmax, making it harder for the model to achieve low loss unless the positive pair has much higher cosine similarity than all negatives (correct)
  • B. Reduce the loss to zero automatically, since dividing by a smaller tau amplifies the positive score disproportionately
  • C. Have no effect on the loss, because the temperature cancels in the log-softmax
  • D. Soften the distribution, spreading probability mass more evenly across all 2N-1 entries

Answer: A

Why: Smaller tau scales all logits up uniformly, concentrating the softmax on the highest-similarity entry. If the positive pair is only slightly more similar than the hardest negative, the model cannot score well — the sharpened distribution demands a large margin. Verified: tau=0.07 gives loss 0.0009 for tight positives (cosine ~0.999) but 2.468 for loose positives.

Why B tempts people
Both numerator and denominator of the softmax are scaled by 1/tau; the positive does not gain a multiplicative advantage — the ratio depends only on the cosine similarity differences, not on tau alone.
Why C tempts people
The temperature does NOT cancel in log-softmax because it scales all terms uniformly before the softmax, changing the relative weights. A uniform scaling does not vanish inside a nonlinear log.
Why D tempts people
Smaller tau concentrates probability mass, not spreads it. Larger tau (approaching infinity) produces a uniform distribution over all entries.

40. Answer it before you see the options: Check — mask head output shape

Prediction

Predict first

A Mask R-CNN mask head receives RoI-aligned features of shape (N, 256, 14, 14). After two Conv2d layers and a ConvTranspose2d(stride=2), then a 1×1 Conv2d to 80 output channels (COCO classes), what is the output shape?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: (N, 80, 28, 28)

Why: ConvTranspose2d with kernel=2, stride=2 doubles spatial resolution: 14 → 28. The final 1×1 Conv2d changes channels from 64 to 80 (one mask per class) without changing spatial size. Result: (N, 80, 28, 28). At inference, sigmoid is applied per channel to get per-class binary masks.

41. Check — mask head output shape

Check

Trace the shapes before clicking.

Check your understanding

A Mask R-CNN mask head receives RoI-aligned features of shape (N, 256, 14, 14). After two Conv2d layers and a ConvTranspose2d(stride=2), then a 1×1 Conv2d to 80 output channels (COCO classes), what is the output shape?

  • A. (N, 80, 28, 28) (correct)
  • B. (N, 80, 14, 14)
  • C. (N, 256, 28, 28)
  • D. (N, 1, 28, 28)

Answer: A

Why: ConvTranspose2d with kernel=2, stride=2 doubles spatial resolution: 14 → 28. The final 1×1 Conv2d changes channels from 64 to 80 (one mask per class) without changing spatial size. Result: (N, 80, 28, 28). At inference, sigmoid is applied per channel to get per-class binary masks.

Why B tempts people
(N, 80, 14, 14) forgets the ConvTranspose2d upsample step — the mask head explicitly doubles spatial resolution to recover fine-grained pixel detail lost in the RoI feature grid.
Why C tempts people
(N, 256, 28, 28) confuses the input channel count (256 from FPN) with the output — the conv layers reduce channels from 256 to 64 before the final prediction head.
Why D tempts people
(N, 1, 28, 28) is the shape of a class-agnostic mask head (one mask regardless of class). Standard Mask R-CNN predicts K masks, one per class, and selects the predicted class's mask at inference.

42. Rule out three: Check — linear probe vs fine-tuning

Elimination

Eliminate the wrong options

You have a SimCLR-pretrained ResNet-50 encoder f and projection head g. You want to classify 500 labeled images. Which approach gives the best accuracy?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Train nn.Linear(h_dim, n_cls) on frozen h = f(x), discarding g
  • B. Train nn.Linear(z_dim, n_cls) on frozen z = g(f(x)), discarding f's gradient
  • C. Fine-tune all parameters of f and g end-to-end on the 500 labeled examples
  • D. Train nn.Linear(z_dim, n_cls) on z but also fine-tune g, keeping f frozen

Survives elimination: A

Why: The canonical SimCLR protocol attaches the linear probe to h = f(x), discarding the projection head g. The encoder h retains all information about the image; g is trained to be invariant to augmentation-irrelevant details, so z contains less information than h. Fine-tuning all of f+g on 500 labels (C) risks overfitting at this scale.

43. Check — linear probe vs fine-tuning

Check

Apply the representation quality argument.

Check your understanding

You have a SimCLR-pretrained ResNet-50 encoder f and projection head g. You want to classify 500 labeled images. Which approach gives the best accuracy?

  • A. Train nn.Linear(h_dim, n_cls) on frozen h = f(x), discarding g (correct)
  • B. Train nn.Linear(z_dim, n_cls) on frozen z = g(f(x)), discarding f's gradient
  • C. Fine-tune all parameters of f and g end-to-end on the 500 labeled examples
  • D. Train nn.Linear(z_dim, n_cls) on z but also fine-tune g, keeping f frozen

Answer: A

Why: The canonical SimCLR protocol attaches the linear probe to h = f(x), discarding the projection head g. The encoder h retains all information about the image; g is trained to be invariant to augmentation-irrelevant details, so z contains less information than h. Fine-tuning all of f+g on 500 labels (C) risks overfitting at this scale.

Why B tempts people
Using z = g(f(x)) gives strictly worse representations than h — the projection head is explicitly trained to discard information the contrastive loss deems irrelevant (e.g. color, texture), which is information downstream tasks often need.
Why C tempts people
Fine-tuning all parameters on 500 labeled examples typically overfits for a ResNet-50. The standard practice is either a frozen linear probe or fine-tuning only the last few layers.
Why D tempts people
Fine-tuning g while keeping f frozen still uses z as the representation, inheriting the information-loss problem of g. Discarding g entirely and using h is the correct protocol.

44. Your turn: SimCLR from scratch

Section

Project

45. Project: implement InfoNCE + SimCLR pipeline

Concept

Build a complete SimCLR training loop on synthetic 64-dim data. Three milestones: InfoNCE loss → full SimCLR pretrain → linear probe evaluation.

#milestonekey deliverable
1Implement nt_xent_loss(z1, z2, tau) from scratch — no libraryverify loss=0.0009 at tau=0.07 with tight positives
2Pretrain TinyEncoder+ProjectionHead with NT-Xent for 60 epochs (no labels)observe loss descent from ~5.56 to ~4.20
3Freeze encoder, train linear head on 160 labeled samples, report test accuracycompare SSL (62.5%) vs supervised (72.5%) baseline

Build rules: L2-normalize embeddings before any cosine similarity; mask the diagonal (self-similarity) from the (2B, 2B) matrix; attach the linear probe to h = encoder(x), NOT to z = projector(h).

46. Break it if you can: Project: implement InfoNCE + SimCLR pipeline

Counterexample

Discussion prompt

Build a complete SimCLR training loop on synthetic 64-dim data. Three milestones: InfoNCE loss → full SimCLR pretrain → linear probe evaluation.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

47. Milestone 1 — NT-Xent loss implementation

Worked example

Your turn: implement nt_xent_loss. Key question: what are the labels in the cross-entropy call? For index i in [0..B-1], the positive is at index i+B in the 2B row; for index i in [B..2B-1], the positive is i-B.

Hint: torch.cat([torch.arange(B, 2*B), torch.arange(B)]) builds the label vector directly. After masking the diagonal with -1e9, nn.CrossEntropyLoss handles the rest.

import torch, torch.nn as nn

def nt_xent_loss(z1, z2, tau=0.5):
    B = z1.shape[0]
    z1 = nn.functional.normalize(z1, dim=-1)
    z2 = nn.functional.normalize(z2, dim=-1)
    z  = torch.cat([z1, z2], dim=0)      # (2B, D)
    sim = z @ z.T / tau                   # (2B, 2B)
    sim.masked_fill_(torch.eye(2*B, dtype=torch.bool), -1e9)
    labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
    return nn.CrossEntropyLoss()(sim, labels)

torch.manual_seed(42)
B, D = 6, 8
z1 = torch.randn(B, D)
z2 = z1 + 0.05 * torch.randn_like(z1)    # tight positives
print(f'tau=0.07: {nt_xent_loss(z1, z2, 0.07):.4f}')  # 0.0009
print(f'tau=0.50: {nt_xent_loss(z1, z2, 0.50):.4f}')  # 0.8428
print(f'tau=1.00: {nt_xent_loss(z1, z2, 1.00):.4f}')  # 1.4959
taulosswhat it means
0.070.0009tight positives dominate; nearly zero loss
0.100.0083SimCLR default; slightly harder
0.500.8428softer; non-trivial even for tight pairs
1.001.4959almost uniform softmax; little gradient signal

48. Milestones 2 & 3 — pretrain + linear probe

Worked example

Your turn: add data augmentation (Gaussian noise views), run 60 epochs of SSL, then attach a frozen-encoder linear probe. Predict the final test accuracy before running — is 60 epochs enough to beat supervised?

Hint: augment each batch with x + 0.1*torch.randn_like(x) twice (different seeds). After pretraining, call encoder.eval() and extract h_tr = encoder(X_tr) with torch.no_grad() before training the linear head.

import torch, torch.nn as nn
# (TinyEncoder, ProjectionHead, nt_xent_loss defined in earlier milestones)
# --- Data ---
torch.manual_seed(42)
vec0 = torch.zeros(64); vec0[0] = 2.0
vec1 = torch.zeros(64); vec1[0] = -2.0
X = torch.cat([torch.randn(100,64)+vec0, torch.randn(100,64)+vec1])
y = torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm = torch.randperm(200, generator=torch.Generator().manual_seed(42))
X_tr, X_te = X[perm[:160]], X[perm[160:]]
y_tr, y_te = y[perm[:160]], y[perm[160:]]
# --- SSL pretrain ---
torch.manual_seed(42)
enc = TinyEncoder(64, 32, 16); prj = ProjectionHead(16, 16, 8)
opt = torch.optim.Adam(list(enc.parameters())+list(prj.parameters()), lr=1e-3)
for epoch in range(60):
    g1 = torch.Generator().manual_seed(epoch)
    g2 = torch.Generator().manual_seed(epoch+100)
    v1 = enc(X_tr + 0.1*torch.randn(X_tr.shape, generator=g1))
    v2 = enc(X_tr + 0.1*torch.randn(X_tr.shape, generator=g2))
    l = nt_xent_loss(prj(v1), prj(v2), tau=0.5)
    opt.zero_grad(); l.backward(); opt.step()
# --- Linear probe ---
with torch.no_grad():
    h_tr = enc(X_tr); h_te = enc(X_te)   # attach to h, NOT z
lin = nn.Linear(16, 2)
opt2 = torch.optim.Adam(lin.parameters(), lr=1e-2)
for _ in range(200):
    opt2.zero_grad()
    nn.CrossEntropyLoss()(lin(h_tr),y_tr).backward(); opt2.step()
with torch.no_grad():
    acc = (lin(h_te).argmax(1)==y_te).float().mean()
print(f'SSL linear probe: {acc*100:.1f}%')   # 62.5%
methodtest accuracylabels used in training?
SSL pretrain + linear probe62.5%no (pretraining); yes (linear head only)
Supervised MLP (same arch)72.5%yes (full training)
SimCLR on ImageNet (ResNet-50)76.5%no (pretraining); yes (linear probe)

49. What each one costs: Milestones 2 & 3 — pretrain + linear probe

Trade off

Comparison matrix

From Milestones 2 & 3 — pretrain + linear probe: every row here is a choice with a cost. Fill the labels used in training? column, then say which row you would actually pick and what you give up for it.

methodtest accuracylabels used in training?
SSL pretrain + linear probe62.5%no (pretraining); yes (linear head only)
Supervised MLP (same arch)72.5%yes (full training)
SimCLR on ImageNet (ResNet-50)76.5%no (pretraining); yes (linear probe)

50. The full program

Concept

import torch, torch.nn as nn

class TinyEncoder(nn.Module):
    def __init__(self, in_d=64, hid=32, out_d=16):
        super().__init__()
        self.net = nn.Sequential(nn.Linear(in_d,hid), nn.ReLU(), nn.Linear(hid,out_d))
    def forward(self, x): return self.net(x)

class ProjectionHead(nn.Module):
    def __init__(self, in_d=16, hid=16, out_d=8):
        super().__init__()
        self.net = nn.Sequential(nn.Linear(in_d,hid), nn.ReLU(), nn.Linear(hid,out_d))
    def forward(self, x): return self.net(x)

def nt_xent(z1, z2, tau=0.5):
    B = z1.shape[0]
    z1 = nn.functional.normalize(z1, dim=-1)
    z2 = nn.functional.normalize(z2, dim=-1)
    z  = torch.cat([z1, z2])
    sim = z @ z.T / tau
    sim.masked_fill_(torch.eye(2*B, dtype=torch.bool), -1e9)
    labels = torch.cat([torch.arange(B, 2*B), torch.arange(B)])
    return nn.CrossEntropyLoss()(sim, labels)

torch.manual_seed(42)
v0=torch.zeros(64); v0[0]=2.0; v1=torch.zeros(64); v1[0]=-2.0
X=torch.cat([torch.randn(100,64)+v0, torch.randn(100,64)+v1])
y=torch.cat([torch.zeros(100,dtype=torch.long), torch.ones(100,dtype=torch.long)])
perm=torch.randperm(200, generator=torch.Generator().manual_seed(42))
Xt,Xe,yt,ye=X[perm[:160]],X[perm[160:]],y[perm[:160]],y[perm[160:]]

torch.manual_seed(42)
enc=TinyEncoder(); prj=ProjectionHead()
opt=torch.optim.Adam(list(enc.parameters())+list(prj.parameters()), lr=1e-3)
for ep in range(60):
    g1=torch.Generator().manual_seed(ep)
    g2=torch.Generator().manual_seed(ep+100)
    l=nt_xent(prj(enc(Xt+0.1*torch.randn(Xt.shape,generator=g1))),
              prj(enc(Xt+0.1*torch.randn(Xt.shape,generator=g2))), tau=0.5)
    opt.zero_grad(); l.backward(); opt.step()
    if ep%10==0 or ep==59: print(f'ep {ep}: {l.item():.4f}')

with torch.no_grad(): ht=enc(Xt); he=enc(Xe)
lin=nn.Linear(16,2); opt2=torch.optim.Adam(lin.parameters(),lr=1e-2)
for _ in range(200):
    opt2.zero_grad(); nn.CrossEntropyLoss()(lin(ht),yt).backward(); opt2.step()
with torch.no_grad():
    print(f'Linear probe: {(lin(he).argmax(1)==ye).float().mean()*100:.1f}%')  # 62.5%
design choicevaluerationale
tau (temperature)0.5 (training), 0.07 (tight-pair demo)0.5 trains stably; 0.07 is SimCLR paper default
augmentationGaussian noise std=0.1minimal; real SimCLR uses crop+flip+color jitter
projection head16→16→8 with ReLUdiscard after pretraining; probe on h, not z
linear probe lr1e-2 (Adam)much higher than SSL lr; only 2 parameters per class
SSL vs supervised gap62.5% vs 72.5%closes at larger N and more epochs (scale is key)

51. Fill in: rationale for The full program

Comparison

Comparison matrix

From The full program: refill the rationale column from what you know. The rest of the table is as it appeared.

design choicevaluerationale
tau (temperature)0.5 (training), 0.07 (tight-pair demo)0.5 trains stably; 0.07 is SimCLR paper default
augmentationGaussian noise std=0.1minimal; real SimCLR uses crop+flip+color jitter
projection head16→16→8 with ReLUdiscard after pretraining; probe on h, not z
linear probe lr1e-2 (Adam)much higher than SSL lr; only 2 parameters per class
SSL vs supervised gap62.5% vs 72.5%closes at larger N and more epochs (scale is key)

52. Show it off

Concept

Out loud, slides closed: (1) describe the full SimCLR pipeline from a raw batch of images to a scalar NT-Xent loss, naming every tensor shape; (2) explain why you attach the linear probe to h = f(x) and not to z = g(f(x)); (3) explain what happens to NT-Xent loss as tau → 0 and as tau → ∞.

Stretch (from the lesson plan): implement the InfoNCE mutual information lower bound — show that minimizing InfoNCE is equivalent to maximizing a lower bound on I(h; x). Also: run the linear probe for multiple label fractions (10%, 25%, 50%, 100%) and plot SSL vs supervised accuracy as a function of label count.

53. Connect it up: Lesson 112: Instance Segmentation & Contrastive Learning

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Mask R-CNN — instance segmentation · Contrastive learning — SimCLR & InfoNCE · Linear probe & self-supervised evaluation · Your turn: SimCLR from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

54. What you can do now

Recap

conceptthe one thing to remember
Mask R-CNNFaster R-CNN + RoIAlign + conv→deconv mask head; K masks output
Panoptic seg.things (instances) + stuff (semantic); metric = PQ = SQ × RQ
InfoNCE / NT-Xent−log(exp(sim(zi,zj)/τ) / Σk exp(sim(zi,zk)/τ)); labels = positive index
tausmaller = sharper softmax; SimCLR uses 0.1; too small → gradient explosion
Linear probefreeze f, train nn.Linear on h (not z); measures representation quality

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 112 — Instance Segmentation & Contrastive Learning — Barron · USAAIO Round 2 Preparation, 2026
  2. He et al. 'Mask R-CNN' (ICCV 2017) — arXiv:1703.06870
  3. Chen et al. 'A Simple Framework for Contrastive Learning of Visual Representations' (SimCLR, ICML 2020) — arXiv:2002.05709
  4. Kirillov et al. 'Panoptic Segmentation' (CVPR 2019) — arXiv:1801.00868
  5. InfoNCE loss, NT-Xent, mask head shapes, linear probe accuracy verified with torch 2.7.1+cpu, synthetic data, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108