USAAIO Lesson 113, from Phase 3. It covers the symmetric InfoNCE contrastive loss, the dual-encoder CLIP architecture with its image and text towers, the learned temperature parameter, zero-shot transfer through text-based class descriptors, and the CLIP extensions ALIGN and BLIP-2. A TinyCLIP is built from scratch with torch 2.7.1+cpu, and all the logit values, softmax outputs, and loss numbers - InfoNCE at 1.2648, the perfect case at 0.0000, and the random case at log(4) = 1.3863 - were verified by real execution. The lesson runs to 25 slides.
Subject: Machine Learning · 49 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 113 · Phase 3
Train image and text encoders jointly on 400 M internet image-text pairs using InfoNCE contrastive loss. At inference: no task-specific head — compare an image embedding to text embeddings of class names and pick the closest. Zero-shot transfer, verified from scratch.
Objectives
log_T and compute how it sharpens or flattens the softmaxWarm-up
Discussion prompt
Before we open Lesson 113: CLIP — Contrastive Language-Image Pretraining: without looking back, what was the main idea of Instance Segmentation & Contrastive Learning, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Mask R-CNN adds a per-proposal mask head to Faster R-CNN (RoI features -> conv stack -> ConvTranspose2d -> per-class binary masks), panoptic segmentation unifies semantic + instance predictions, and self-supervised contrastive learning (SimCLR / InfoNCE) trains representations without labels by pulling augmented views of the same image together while pushing different images apart.
Section
Part 1 of 4
Concept
CLIP processes B image-text pairs per batch. Both the image encoder and text encoder produce L2-normalized D-dimensional embeddings. Cosine similarity of every (image, text) pair fills a B×B matrix S.
\[ S_{ij} = \frac{z_i^I \cdot z_j^T}{\|z_i^I\|\,\|z_j^T\|} = z_i^I \cdot z_j^T \quad (\text{after L2-norm}) \]
The diagonal entries are the B matching pairs; the off-diagonal entries are the B(B-1) non-matching pairs. CLIP pushes diagonal high and off-diagonal low.
| entry | meaning | CLIP target |
|---|---|---|
| S[i,i] | image i paired with its own caption | push toward +1 |
| S[i,j] (i≠j) | image i vs a different caption | push toward -1 |
| S shape | B×B for batch size B | e.g. 32768×32768 in CLIP |
Comparison
Comparison matrix
From The B×B similarity matrix: refill the meaning column from what you know. The rest of the table is as it appeared.
| entry | meaning | CLIP target |
|---|---|---|
| S[i,i] | image i paired with its own caption | push toward +1 |
| S[i,j] (i≠j) | image i vs a different caption | push toward -1 |
| S shape | B×B for batch size B | e.g. 32768×32768 in CLIP |
Concept
Scale S by a learnable temperature tau (stored as log_tau, optimized with the encoders), then apply cross-entropy in both directions: rows (each image finds its text) and columns (each text finds its image).
\[ \mathcal{L} = \frac{1}{2}\Bigl[\text{CE}\bigl(S/\tau,\,\mathbf{y}\bigr) + \text{CE}\bigl((S/\tau)^\top,\,\mathbf{y}\bigr)\Bigr] \]
\[ \mathbf{y} = [0,1,2,\dots,B{-}1] \quad \text{(index of the matching pair for each row/col)} \]
Random initialization gives loss ≈ log(B) (uniform over B choices). Perfect alignment gives loss ≈ 0. For B=4 with all-zero similarities: random_loss = log(4) = 1.3863 (verified).
Counterexample
Discussion prompt
Random initialization gives loss ≈ log(B) (uniform over B choices). Perfect alignment gives loss ≈ 0. For B=4 with all-zero similarities: random_loss = log(4) = 1.3863 (verified).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Ranking
Put in order
Put the moves of Trace InfoNCE: B=3, T=0.1 into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Use hand-constructed vectors where img_i and txt_i are deliberately close; L2-normalize to put all vectors on the unit sphere so dot-product equals cosine similarity.
Worked example
Build 4-dim normalized embeddings for 3 image-text pairs
Why: Use hand-constructed vectors where img_i and txt_i are deliberately close; L2-normalize to put all vectors on the unit sphere so dot-product equals cosine similarity.
Compute S = img_emb @ txt_emb.T, then L = S / 0.1
Why: Temperature T=0.1 sharpens the distribution — a 0.9911 vs 0.6632 difference in S becomes a 9.9107 vs 6.6319 difference in L, making the softmax strongly favor the diagonal.
import torch, torch.nn.functional as F
i_e = F.normalize(torch.tensor([
[1.0, 0.0, 0.5, -0.2],
[0.0, 1.0, 0.3, 0.1],
[0.5, 0.3, 1.0, 0.4],
]), dim=-1)
t_e = F.normalize(torch.tensor([
[0.9, 0.1, 0.4, -0.1],
[0.1, 0.9, 0.2, 0.2],
[0.4, 0.2, 0.9, 0.3],
]), dim=-1)
S = i_e @ t_e.T # (3,3)
L = S / 0.1 # T=0.1
labels = torch.arange(3)
loss_r = F.cross_entropy(L, labels) # img->txt
loss_c = F.cross_entropy(L.T, labels) # txt->img
loss = (loss_r + loss_c) / 2
print(S.numpy().round(4))
print('loss_r', round(loss_r.item(),4),
'loss_c', round(loss_c.item(),4),
'sym', round(loss.item(),4))| pair | S[i,i] (diag) | softmax p[correct] | CE contribution |
|---|---|---|---|
| img0/txt0 | 0.9911 | 0.9635 | 0.0372 |
| img1/txt1 | 0.9849 | 0.9998 (approx) | 0.0002 (approx) |
| img2/txt2 | 0.9965 | 0.9635 (approx) | 0.0372 (approx) |
| symmetric loss | — | — | 0.0321 |
Read output: S diag=[0.9911, 0.9849, 0.9965]; sym loss=0.0321
Why: Loss is near 0 because all diagonal cosines exceed 0.98, making softmax probability for the correct pair > 0.96. Total symmetric InfoNCE = (0.0319 + 0.0323)/2 = 0.0321.
Trade off
Comparison matrix
From Trace InfoNCE: B=3, T=0.1: every row here is a choice with a cost. Fill the CE contribution column, then say which row you would actually pick and what you give up for it.
| pair | S[i,i] (diag) | softmax p[correct] | CE contribution |
|---|---|---|---|
| img0/txt0 | 0.9911 | 0.9635 | 0.0372 |
| img1/txt1 | 0.9849 | 0.9998 (approx) | 0.0002 (approx) |
| img2/txt2 | 0.9965 | 0.9635 (approx) | 0.0372 (approx) |
| symmetric loss | — | — | 0.0321 |
Concept
CLIP learns log_tau jointly with the encoders. At inference tau ≈ 0.07. A small tau sharpens the softmax (one-hot-like), enforcing tight clusters. A large tau makes all pairs equally likely — gradients vanish.
\[ p_i = \frac{\exp(S_{ii}/\tau)}{\sum_j \exp(S_{ij}/\tau)} \]
| T | p[correct] (sim=0.9 vs 0.3,0.2,-0.1) | entropy | regime |
|---|---|---|---|
| 0.01 | 1.0000 | 0.0000 | one-hot (too hard) |
| 0.07 | 0.9998 | 0.0023 | CLIP default |
| 0.50 | 0.5941 | 1.1013 | moderate |
| 1.00 | 0.4144 | 1.3139 | soft |
| 2.00 | 0.3277 | 1.3688 | nearly uniform |
Anomaly
Predict first
A student writes this, and it looks reasonable:
For each positive pair, compute loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.
For each positive pair, compute softmax over all B negatives in the batch simultaneously — this is F.cross_entropy(S/T, labels).
Why: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.
Trap
For each positive pair, compute loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.
Treat it as B independent binary classification problems
Why: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.
Problem: negatives in the batch are not shared; each pair ignores all other pairs in the batch, so large-batch benefits are lost.
For each positive pair, compute softmax over all B negatives in the batch simultaneously — this is F.cross_entropy(S/T, labels).
Use the B×B similarity matrix and row-wise softmax cross-entropy
Why: Each row is a B-class classification: 'which text in this batch matches image i?' In-batch negatives are free — no separate negative mining.
Large B means harder negatives (more confusable pairs), stronger gradient signal. CLIP used B=32768 (32k pairs per step). loss ≈ log(B) at random init.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Each row is a B-class classification: 'which text in this batch matches image i?' In-batch negatives are free — no separate negative mining.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.
Section
Part 2 of 4
Concept
CLIP trains two independent encoders that share only the contrastive loss signal. The image encoder (ViT or ResNet, see Lesson 92) projects images to D-dim; the text encoder (GPT-like transformer) projects captions to the same D-dim.
| component | input | output shape | CLIP-large config |
|---|---|---|---|
| Image encoder | H×W×3 image | (B, D) | ViT-L/14, D=768 |
| Text encoder | tokenized caption | (B, D) | GPT-style, 12-layer |
| L2 normalize | raw embeddings | (B, D) unit vectors | applied to both |
| Similarity S | img @ txt.T | (B, B) | scaled by 1/tau |
| InfoNCE loss | S, labels | scalar | symmetric over rows+cols |
Pattern
Predict first
The table runs: Conv2d(1,8,3,pad=1) + ReLU | (B, 8, 8, 8) | input 8×8 image · AdaptiveAvgPool2d(2) | (B, 8, 2, 2) | → flatten → (B, 32) · Linear(32, 16) + L2-norm | (B, 16) | unit-norm image feats · Embedding(20,16).mean + Linear | (B, 16) | unit-norm text feats · img @ txt.T * 14.29 | (B, B) | logits for InfoNCE · F.cross_entropy (symmetric) | scalar | loss=2.6814 (random init)
In TinyCLIP forward pass in PyTorch, given the rows so far: what is the next one — the row where layer / step is Total params?
Correct: Total params | 1200 | 608 img + 592 txt
| layer / step | output shape | verified value |
|---|---|---|
| Conv2d(1,8,3,pad=1) + ReLU | (B, 8, 8, 8) | input 8×8 image |
| AdaptiveAvgPool2d(2) | (B, 8, 2, 2) | → flatten → (B, 32) |
| Linear(32, 16) + L2-norm | (B, 16) | unit-norm image feats |
| Embedding(20,16).mean + Linear | (B, 16) | unit-norm text feats |
| img @ txt.T * 14.29 | (B, B) | logits for InfoNCE |
| F.cross_entropy (symmetric) | scalar | loss=2.6814 (random init) |
| Total params | 1200 | 608 img + 592 txt |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Real CLIP uses ViT-L and GPT-style text transformer; TinyCLIP replaces them with a 3-layer CNN and a bag-of-embeddings to verify the loss pipeline in < 1 second.
Worked example
Define TinyImageEncoder (CNN) and TinyTextEncoder (embedding + mean pooling), both projecting to D=16
Why: Real CLIP uses ViT-L and GPT-style text transformer; TinyCLIP replaces them with a 3-layer CNN and a bag-of-embeddings to verify the loss pipeline in < 1 second.
import torch, torch.nn as nn, torch.nn.functional as F
class TinyImageEncoder(nn.Module):
def __init__(self, d=16):
super().__init__()
self.conv = nn.Sequential(
nn.Conv2d(1, 8, kernel_size=3, padding=1),
nn.ReLU(),
nn.AdaptiveAvgPool2d(2), # -> (B,8,2,2)
)
self.proj = nn.Linear(8*2*2, d) # 32 -> 16
def forward(self, x):
return F.normalize(self.proj(self.conv(x).flatten(1)), dim=-1)
class TinyTextEncoder(nn.Module):
def __init__(self, vocab=20, d=16):
super().__init__()
self.emb = nn.Embedding(vocab, d)
self.proj = nn.Linear(d, d)
def forward(self, tokens):
return F.normalize(self.proj(self.emb(tokens).mean(1)), dim=-1)
# TinyImageEncoder params: 608 TinyTextEncoder params: 592 Total: 1200
img_enc, txt_enc = TinyImageEncoder(), TinyTextEncoder()
print(sum(p.numel() for p in img_enc.parameters()), # 608
sum(p.numel() for p in txt_enc.parameters())) # 592| layer / step | output shape | verified value |
|---|---|---|
| Conv2d(1,8,3,pad=1) + ReLU | (B, 8, 8, 8) | input 8×8 image |
| AdaptiveAvgPool2d(2) | (B, 8, 2, 2) | → flatten → (B, 32) |
| Linear(32, 16) + L2-norm | (B, 16) | unit-norm image feats |
| Embedding(20,16).mean + Linear | (B, 16) | unit-norm text feats |
| img @ txt.T * 14.29 | (B, B) | logits for InfoNCE |
| F.cross_entropy (symmetric) | scalar | loss=2.6814 (random init) |
| Total params | 1200 | 608 img + 592 txt |
Run with B=6, random init: loss=2.6814, log(6)=1.7918
Why: Random-init loss > log(B) because T=14.29 creates very sharp distributions even before training — large T amplifies noise. With T=1/0.07≈14.29, the effective batch feels much larger to the loss. After training, loss converges toward 0.
Comparison
Comparison matrix
From TinyCLIP forward pass in PyTorch: refill the output shape column from what you know. The rest of the table is as it appeared.
| layer / step | output shape | verified value |
|---|---|---|
| Conv2d(1,8,3,pad=1) + ReLU | (B, 8, 8, 8) | input 8×8 image |
| AdaptiveAvgPool2d(2) | (B, 8, 2, 2) | → flatten → (B, 32) |
| Linear(32, 16) + L2-norm | (B, 16) | unit-norm image feats |
| Embedding(20,16).mean + Linear | (B, 16) | unit-norm text feats |
| img @ txt.T * 14.29 | (B, B) | logits for InfoNCE |
| F.cross_entropy (symmetric) | scalar | loss=2.6814 (random init) |
| Total params | 1200 | 608 img + 592 txt |
Section
Part 3 of 4
Concept
At inference CLIP needs no labeled training data for a new task. Encode each class name as a text prompt (e.g. 'a photo of a {class}'), compute all text embeddings, then find the class whose text embedding is most similar to the image embedding.
\[ \hat{y} = \arg\max_{k} \; z^I \cdot z^T_k \quad z^T_k = \text{TextEnc}(\texttt{"a photo of a "}+c_k) \]
| class text prompt | cos_sim to image | softmax prob (T=0.07) |
|---|---|---|
| a cat | 0.2163 | 0.0027 |
| a dog | 0.6296 | 0.9893 |
| a car | 0.2923 | 0.0080 |
| a bird | -0.3270 | 0.0000 |
| predicted | — | a dog (verified) |
Intuition
The InfoNCE loss forces image and text of the same concept to occupy nearby regions on a high-dimensional unit sphere. After training on 400 M internet pairs, 'a photo of a golden retriever' and an image of a golden retriever are neighbors — even for rare classes never seen as a classification task.
This is analogous to word2vec (Lesson 98): just as king - man + woman ≈ queen emerges from co-occurrence, CLIP's space captures visual-semantic analogies from co-occurrence of images and captions at web scale.
Prompt engineering matters: 'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.
Counterexample
Discussion prompt
Prompt engineering matters: 'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Pattern
Predict first
The table runs: a cat | 0.2163 | 0.0027 | no · a dog | 0.6296 | 0.9893 | YES · a car | 0.2923 | 0.0080 | no
In Zero-shot CLIP with TinyCLIP (PyTorch), given the rows so far: what is the next one — the row where class is a bird?
Correct: a bird | -0.3270 | 0.0000 | no
| class | cos_sim | softmax prob (T=0.07) | predicted? |
|---|---|---|---|
| a cat | 0.2163 | 0.0027 | no |
| a dog | 0.6296 | 0.9893 | YES |
| a car | 0.2923 | 0.0080 | no |
| a bird | -0.3270 | 0.0000 | no |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Using TinyCLIP (random weights) as a drop-in for real CLIP; structure is identical — only scale differs.
Worked example
Encode a query image and K=4 class text prompts; all outputs are unit vectors in R^16
Why: Using TinyCLIP (random weights) as a drop-in for real CLIP; structure is identical — only scale differs.
import torch, torch.nn.functional as F
# img_enc, txt_enc from TinyCLIP above
torch.manual_seed(7)
img_q = F.normalize(torch.randn(1, 16), dim=-1) # 1 query image
cls_embs = F.normalize(torch.randn(4, 16), dim=-1) # 4 class text embs
# nudge class 1 ("a dog") to be similar to the image
cls_embs[1] = F.normalize(img_q[0] + 0.3*torch.randn(16), dim=-1)
sims = (img_q @ cls_embs.T).squeeze(0) # (4,)
probs = F.softmax(sims / 0.07, dim=0) # T=0.07
# predicted class:
pred = probs.argmax().item() # -> 1 ("a dog")
for name, s, p in zip(
['a cat','a dog','a car','a bird'],
sims.tolist(), probs.tolist()):
print(f'{name:12s} cos={s:.4f} prob={p:.4f}')| class | cos_sim | softmax prob (T=0.07) | predicted? |
|---|---|---|---|
| a cat | 0.2163 | 0.0027 | no |
| a dog | 0.6296 | 0.9893 | YES |
| a car | 0.2923 | 0.0080 | no |
| a bird | -0.3270 | 0.0000 | no |
argmax over probs selects index 1 ('a dog') with p=0.9893
Why: Even with T=0.07 (very sharp), a cosine difference of 0.6296 vs 0.2923 translates to a 99:1 probability ratio — zero-shot is decisive when embeddings are well-separated.
Discrimination
Sort into buckets
Sort these by predicted?, from memory, without looking back at Zero-shot CLIP with TinyCLIP (PyTorch). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The similarity matrix entries now reflect magnitude differences, not directional alignment.
Apply F.normalize(emb, dim=-1) to both img_emb and txt_emb before the dot product. Then S[i,j] is a cosine similarity in [-1, +1].
Why: The similarity matrix entries now reflect magnitude differences, not directional alignment. A high-norm pair dominates regardless of semantic similarity.
Trap
Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).
Scale by temperature and apply cross-entropy
Why: The similarity matrix entries now reflect magnitude differences, not directional alignment. A high-norm pair dominates regardless of semantic similarity.
Result: gradients push the network to increase embedding norms rather than learn direction — training collapses or diverges.
Apply F.normalize(emb, dim=-1) to both img_emb and txt_emb before the dot product. Then S[i,j] is a cosine similarity in [-1, +1].
Use img_emb = F.normalize(proj(h), dim=-1) at the end of each encoder
Why: Projection to the unit sphere decouples direction (semantics) from magnitude. The dot product now measures angular proximity, which is exactly what InfoNCE optimizes.
All values in S are in [-1, 1]. Temperature tau controls the effective scale — no raw-norm leakage.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.; CLIP opened the contrastive vision-language paradigm; subsequent models each extended it along one architectural dimension while keeping the core InfoNCE objective.loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.; Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).Section
Part 4 of 4
Concept
CLIP opened the contrastive vision-language paradigm; subsequent models each extended it along one architectural dimension while keeping the core InfoNCE objective.
| model | key extension over CLIP | data/compute | zero-shot IMAGENET top-1 |
|---|---|---|---|
| CLIP (2021) | dual ViT+GPT, clean curated 400M | 400M, large B | 76.2% |
| ALIGN (2021) | EfficientNet+BERT, noisy 1.8B pairs (no cleaning) | 1.8B alt-text | 76.4% |
| BLIP (2022) | caption + ITC + ITM objectives + web bootstrap | 129M + filtered | 82.3% |
| BLIP-2 (2023) | frozen img encoder + Q-Former + frozen LLM | few trainable params | 73.4% (zero-shot) |
ALIGN's insight: noise in alt-text is overcome by scale — 1.8 B noisy pairs outperforms 400 M curated pairs. BLIP-2's insight: freeze both encoders, train only the Q-Former bridge — most knowledge is already in pretrained weights.
Trade off
Comparison matrix
From ALIGN, BLIP-2: three dimensions of extension: every row here is a choice with a cost. Fill the zero-shot IMAGENET top-1 column, then say which row you would actually pick and what you give up for it.
| model | key extension over CLIP | data/compute | zero-shot IMAGENET top-1 |
|---|---|---|---|
| CLIP (2021) | dual ViT+GPT, clean curated 400M | 400M, large B | 76.2% |
| ALIGN (2021) | EfficientNet+BERT, noisy 1.8B pairs (no cleaning) | 1.8B alt-text | 76.4% |
| BLIP (2022) | caption + ITC + ITM objectives + web bootstrap | 129M + filtered | 82.3% |
| BLIP-2 (2023) | frozen img encoder + Q-Former + frozen LLM | few trainable params | 73.4% (zero-shot) |
Concept
CLIP text embeddings are used as conditioning vectors in diffusion models (DALL-E 2, Stable Diffusion). The text encoder is frozen; only the diffusion U-Net is trained to denoise images conditioned on the text embedding.
| role | component | trainable? |
|---|---|---|
| Text conditioning | CLIP text encoder | Frozen |
| Image conditioning (Img2Img) | CLIP image encoder | Frozen |
| Denoising | U-Net with cross-attention | Trained |
| Schedule | DDPM/DDIM noise schedule | Fixed |
Cross-attention keys/values come from the CLIP text embedding at every U-Net resolution; this is the architectural bridge that makes text-to-image work (Lesson 110+). CLIP is the semantic backbone, not the generative model.
Discrimination
Sort into buckets
Sort these by trainable?, from memory, without looking back at CLIP for generation: conditioning diffusion models. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Ranking
Put in order
These are the steps of CLIP recipe: 6-step blueprint, scrambled. Put them back in order before the next slide shows you.
F.normalize(emb, dim=-1) on both to get unit-sphere embeddingsS = img_emb @ txt_emb.T → (B, B) cosine similaritiesL = S / tau where tau = log_tau.exp(), log_tau is a learned parameterloss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
F.normalize(emb, dim=-1) on both to get unit-sphere embeddingsS = img_emb @ txt_emb.T → (B, B) cosine similaritiesL = S / tau where tau = log_tau.exp(), log_tau is a learned parameterloss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)Edge cases
Discussion prompt
CLIP recipe: 6-step blueprint works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
F.normalize(emb, dim=-1) on both to get unit-sphere embeddingsS = img_emb @ txt_emb.T → (B, B) cosine similaritiesL = S / tau where tau = log_tau.exp(), log_tau is a learned parameterloss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)Elimination
Eliminate the wrong options
For a CLIP batch of B=1024 image-text pairs at random initialization (all cosine similarities ≈ 0), approximately what is the expected InfoNCE loss?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Cross-entropy over B uniform logits = log(B). For B=1024, that is log(1024) ≈ 6.93. Verified: with zero similarities and B=4, loss = log(4) = 1.3863 (real execution). This shows why large batches are harder to optimize from scratch.
Check
Work this out before selecting. Consider what happens as batch size B grows.
Check your understanding
For a CLIP batch of B=1024 image-text pairs at random initialization (all cosine similarities ≈ 0), approximately what is the expected InfoNCE loss?
Answer: A
Why: Cross-entropy over B uniform logits = log(B). For B=1024, that is log(1024) ≈ 6.93. Verified: with zero similarities and B=4, loss = log(4) = 1.3863 (real execution). This shows why large batches are harder to optimize from scratch.
Prediction
Predict first
With similarity row [0.90, 0.30, 0.20, -0.10], increasing temperature from T=0.07 to T=2.0 causes the softmax probability of the correct class (index 0) to change from ≈0.9998 to ≈0.3277. What is the practical consequence for training?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Gradients become smaller and less discriminative; the model receives weaker signal to separate matching from non-matching pairs
Why: At T=2.0 the softmax is nearly uniform (p_correct=0.3277, entropy=1.3688), meaning the loss gradient is nearly the same regardless of which pair is matched — weak supervision. At T=0.07 (CLIP default), p_correct=0.9998 and the gradient focuses sharply on the few pairs that are wrong. Small T = hard training signal.
Check
Trace through the softmax formula before answering.
Check your understanding
With similarity row [0.90, 0.30, 0.20, -0.10], increasing temperature from T=0.07 to T=2.0 causes the softmax probability of the correct class (index 0) to change from ≈0.9998 to ≈0.3277. What is the practical consequence for training?
Answer: A
Why: At T=2.0 the softmax is nearly uniform (p_correct=0.3277, entropy=1.3688), meaning the loss gradient is nearly the same regardless of which pair is matched — weak supervision. At T=0.07 (CLIP default), p_correct=0.9998 and the gradient focuses sharply on the few pairs that are wrong. Small T = hard training signal.
Elimination
Eliminate the wrong options
A CLIP model trained on internet image-caption pairs is tested zero-shot on a medical X-ray classification task with classes ['pleural effusion', 'pneumothorax', 'normal']. Which situation most likely LIMITS zero-shot performance?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: CLIP's zero-shot strength depends on image-text co-occurrence during pretraining. X-ray images paired with clinical terminology are rare in internet text; the shared embedding space is poorly calibrated for this domain. This is the standard domain-shift argument: CLIP excels on natural-image distributions but degrades on out-of-distribution imagery.
Check
Think about what the text encoder output represents before choosing.
Check your understanding
A CLIP model trained on internet image-caption pairs is tested zero-shot on a medical X-ray classification task with classes ['pleural effusion', 'pneumothorax', 'normal']. Which situation most likely LIMITS zero-shot performance?
Answer: A
Why: CLIP's zero-shot strength depends on image-text co-occurrence during pretraining. X-ray images paired with clinical terminology are rare in internet text; the shared embedding space is poorly calibrated for this domain. This is the standard domain-shift argument: CLIP excels on natural-image distributions but degrades on out-of-distribution imagery.
Step zero
Discussion prompt
Your turn: TinyCLIP full training step — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Milestone 1 — build encoders and optimizer
Answer:
Worked example
Task: implement one complete CLIP training step (forward + backward + optimizer step) using TinyCLIP on a synthetic B=8 batch. Verify that loss decreases after the step.
Milestone 1 — build encoders and optimizer
Why: Instantiate TinyImageEncoder(d=16) and TinyTextEncoder(d=16). Create a single torch.optim.Adam over list(img_enc.parameters()) + list(txt_enc.parameters()) + [log_tau] where log_tau = nn.Parameter(torch.tensor(0.0)).
Milestone 2 — forward pass
Why: Generate imgs = torch.randn(8,1,8,8) and toks = torch.randint(0,20,(8,5)). Compute i_f = img_enc(imgs), t_f = txt_enc(toks), then L = (i_f @ t_f.T) / log_tau.exp(). Labels = torch.arange(8).
Milestone 3 — symmetric InfoNCE and backward
Why: Compute loss = (F.cross_entropy(L, labels) + F.cross_entropy(L.T, labels)) / 2. Call loss.backward() then optimizer.step(). Record loss before and after.
import torch, torch.nn as nn, torch.nn.functional as F
# --- re-use TinyImageEncoder / TinyTextEncoder from earlier ---
torch.manual_seed(0)
img_enc = TinyImageEncoder(d=16)
txt_enc = TinyTextEncoder(d=16)
log_tau = nn.Parameter(torch.tensor(0.0)) # learnable temperature
opt = torch.optim.Adam(
list(img_enc.parameters()) +
list(txt_enc.parameters()) + [log_tau], lr=1e-3)
imgs = torch.randn(8, 1, 8, 8)
toks = torch.randint(0, 20, (8, 5))
labels = torch.arange(8)
for step in range(3):
opt.zero_grad()
i_f = img_enc(imgs)
t_f = txt_enc(toks)
L = (i_f @ t_f.T) / log_tau.exp()
loss = (F.cross_entropy(L, labels) +
F.cross_entropy(L.T, labels)) / 2
loss.backward()
opt.step()
print(f'step {step}: loss={loss.item():.4f} tau={log_tau.exp().item():.4f}')| step | expected loss | expected tau | note |
|---|---|---|---|
| 0 (before update) | ~2.0–2.8 | 1.0000 | random init, T=exp(0)=1 |
| 1 | decreasing | ~1.01 | optimizer adjusts tau upward |
| 2 | further decrease | >1.01 | both encoders + tau learn |
| convergence | ~log(8)≈2.08 floor → 0 | ~0.07 target | CLIP trains for 32-epoch equivalents |
Show-it-off: add prompt engineering — prepend 'a photo of a' to 5 synthetic class labels and run zero-shot argmax over a new test image
Why: After any training, the embedding directions improve. Even 3 steps on random data won't produce meaningful zero-shot, but the code structure is identical to real CLIP — only the encoder weights and data scale differ.
Comparison
Comparison matrix
From Your turn: TinyCLIP full training step: refill the expected loss column from what you know. The rest of the table is as it appeared.
| step | expected loss | expected tau | note |
|---|---|---|---|
| 0 (before update) | ~2.0–2.8 | 1.0000 | random init, T=exp(0)=1 |
| 1 | decreasing | ~1.01 | optimizer adjusts tau upward |
| 2 | further decrease | >1.01 | both encoders + tau learn |
| convergence | ~log(8)≈2.08 floor → 0 | ~0.07 target | CLIP trains for 32-epoch equivalents |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Contrastive objective — matching pairs in a batch · Dual-encoder architecture — image tower + text tower · Zero-shot transfer — classify without a linear head · CLIP extensions — ALIGN, BLIP, BLIP-2. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
log_tau, small tau (0.07) sharpens gradients; too large = uniform, weak signal| concept | key number / formula | verified |
|---|---|---|
| InfoNCE loss (B=4, random) | log(4) = 1.3863 | yes |
| InfoNCE loss (B=3, near-perfect) | 0.0321 | yes |
| Zero-shot p(correct) T=0.07 | 0.9893 (cos=0.630) | yes |
| Temperature: T=0.07 vs T=2.0 | p_correct: 0.9998 vs 0.3277 | yes |
| TinyCLIP params | 1200 (608 img + 592 txt) | yes |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.