Lesson 113: CLIP — Contrastive Language-Image Pretraining

USAAIO Lesson 113, from Phase 3. It covers the symmetric InfoNCE contrastive loss, the dual-encoder CLIP architecture with its image and text towers, the learned temperature parameter, zero-shot transfer through text-based class descriptors, and the CLIP extensions ALIGN and BLIP-2. A TinyCLIP is built from scratch with torch 2.7.1+cpu, and all the logit values, softmax outputs, and loss numbers - InfoNCE at 1.2648, the perfect case at 0.0000, and the random case at log(4) = 1.3863 - were verified by real execution. The lesson runs to 25 slides.

Subject: Machine Learning · 49 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. CLIP Contrastive Vision-Language

Title

USAAIO · Lesson 113 · Phase 3

Train image and text encoders jointly on 400 M internet image-text pairs using InfoNCE contrastive loss. At inference: no task-specific head — compare an image embedding to text embeddings of class names and pick the closest. Zero-shot transfer, verified from scratch.

2. By the end of this lesson you can

Objectives

  1. Write the symmetric InfoNCE loss as the sum of two cross-entropies over the B×B similarity matrix
  2. Explain the role of the learned temperature log_T and compute how it sharpens or flattens the softmax
  3. Trace a full CLIP forward pass from raw image + text to scalar loss, with correct shapes at each step
  4. Perform zero-shot classification by building text embeddings for each class and taking the argmax cosine similarity
  5. Compare CLIP, ALIGN, and BLIP-2 by the dimension they scale (data noise tolerance, batch size, frozen encoders + Q-Former)

3. What survived from Instance Segmentation & Contrastive Learning?

Warm-up

Discussion prompt

Before we open Lesson 113: CLIP — Contrastive Language-Image Pretraining: without looking back, what was the main idea of Instance Segmentation & Contrastive Learning, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Mask R-CNN adds a per-proposal mask head to Faster R-CNN (RoI features -> conv stack -> ConvTranspose2d -> per-class binary masks), panoptic segmentation unifies semantic + instance predictions, and self-supervised contrastive learning (SimCLR / InfoNCE) trains representations without labels by pulling augmented views of the same image together while pushing different images apart.

4. Contrastive objective — matching pairs in a batch

Section

Part 1 of 4

5. The B×B similarity matrix

Concept

CLIP processes B image-text pairs per batch. Both the image encoder and text encoder produce L2-normalized D-dimensional embeddings. Cosine similarity of every (image, text) pair fills a B×B matrix S.

\[ S_{ij} = \frac{z_i^I \cdot z_j^T}{\|z_i^I\|\,\|z_j^T\|} = z_i^I \cdot z_j^T \quad (\text{after L2-norm}) \]

The diagonal entries are the B matching pairs; the off-diagonal entries are the B(B-1) non-matching pairs. CLIP pushes diagonal high and off-diagonal low.

entrymeaningCLIP target
S[i,i]image i paired with its own captionpush toward +1
S[i,j] (i≠j)image i vs a different captionpush toward -1
S shapeB×B for batch size Be.g. 32768×32768 in CLIP

6. Fill in: meaning for The B×B similarity matrix

Comparison

Comparison matrix

From The B×B similarity matrix: refill the meaning column from what you know. The rest of the table is as it appeared.

entrymeaningCLIP target
S[i,i]image i paired with its own captionpush toward +1
S[i,j] (i≠j)image i vs a different captionpush toward -1
S shapeB×B for batch size Be.g. 32768×32768 in CLIP

7. Symmetric InfoNCE loss

Concept

Scale S by a learnable temperature tau (stored as log_tau, optimized with the encoders), then apply cross-entropy in both directions: rows (each image finds its text) and columns (each text finds its image).

\[ \mathcal{L} = \frac{1}{2}\Bigl[\text{CE}\bigl(S/\tau,\,\mathbf{y}\bigr) + \text{CE}\bigl((S/\tau)^\top,\,\mathbf{y}\bigr)\Bigr] \]

\[ \mathbf{y} = [0,1,2,\dots,B{-}1] \quad \text{(index of the matching pair for each row/col)} \]

Random initialization gives loss ≈ log(B) (uniform over B choices). Perfect alignment gives loss ≈ 0. For B=4 with all-zero similarities: random_loss = log(4) = 1.3863 (verified).

8. Break it if you can: Symmetric InfoNCE loss

Counterexample

Discussion prompt

Random initialization gives loss ≈ log(B) (uniform over B choices). Perfect alignment gives loss ≈ 0. For B=4 with all-zero similarities: random_loss = log(4) = 1.3863 (verified).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. What has to happen first: Trace InfoNCE: B=3, T=0.1

Ranking

Put in order

Put the moves of Trace InfoNCE: B=3, T=0.1 into the order they have to happen.

  1. Build 4-dim normalized embeddings for 3 image-text pairs
  2. Compute S = img_emb @ txt_emb.T, then L = S / 0.1
  3. Read output: S diag=[0.9911, 0.9849, 0.9965]; sym loss=0.0321

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Use hand-constructed vectors where img_i and txt_i are deliberately close; L2-normalize to put all vectors on the unit sphere so dot-product equals cosine similarity.

10. Trace InfoNCE: B=3, T=0.1

Worked example

Build 4-dim normalized embeddings for 3 image-text pairs

Why: Use hand-constructed vectors where img_i and txt_i are deliberately close; L2-normalize to put all vectors on the unit sphere so dot-product equals cosine similarity.

Compute S = img_emb @ txt_emb.T, then L = S / 0.1

Why: Temperature T=0.1 sharpens the distribution — a 0.9911 vs 0.6632 difference in S becomes a 9.9107 vs 6.6319 difference in L, making the softmax strongly favor the diagonal.

import torch, torch.nn.functional as F

i_e = F.normalize(torch.tensor([
    [1.0, 0.0, 0.5, -0.2],
    [0.0, 1.0, 0.3,  0.1],
    [0.5, 0.3, 1.0,  0.4],
]), dim=-1)
t_e = F.normalize(torch.tensor([
    [0.9, 0.1, 0.4, -0.1],
    [0.1, 0.9, 0.2,  0.2],
    [0.4, 0.2, 0.9,  0.3],
]), dim=-1)

S = i_e @ t_e.T          # (3,3)
L = S / 0.1              # T=0.1
labels = torch.arange(3)

loss_r = F.cross_entropy(L, labels)      # img->txt
loss_c = F.cross_entropy(L.T, labels)    # txt->img
loss   = (loss_r + loss_c) / 2

print(S.numpy().round(4))
print('loss_r', round(loss_r.item(),4),
      'loss_c', round(loss_c.item(),4),
      'sym', round(loss.item(),4))
pairS[i,i] (diag)softmax p[correct]CE contribution
img0/txt00.99110.96350.0372
img1/txt10.98490.9998 (approx)0.0002 (approx)
img2/txt20.99650.9635 (approx)0.0372 (approx)
symmetric loss——0.0321

Read output: S diag=[0.9911, 0.9849, 0.9965]; sym loss=0.0321

Why: Loss is near 0 because all diagonal cosines exceed 0.98, making softmax probability for the correct pair > 0.96. Total symmetric InfoNCE = (0.0319 + 0.0323)/2 = 0.0321.

11. What each one costs: Trace InfoNCE: B=3, T=0.1

Trade off

Comparison matrix

From Trace InfoNCE: B=3, T=0.1: every row here is a choice with a cost. Fill the CE contribution column, then say which row you would actually pick and what you give up for it.

pairS[i,i] (diag)softmax p[correct]CE contribution
img0/txt00.99110.96350.0372
img1/txt10.98490.9998 (approx)0.0002 (approx)
img2/txt20.99650.9635 (approx)0.0372 (approx)
symmetric loss——0.0321

12. Temperature: the one hyperparameter that matters most

Concept

CLIP learns log_tau jointly with the encoders. At inference tau ≈ 0.07. A small tau sharpens the softmax (one-hot-like), enforcing tight clusters. A large tau makes all pairs equally likely — gradients vanish.

\[ p_i = \frac{\exp(S_{ii}/\tau)}{\sum_j \exp(S_{ij}/\tau)} \]

Tp[correct] (sim=0.9 vs 0.3,0.2,-0.1)entropyregime
0.011.00000.0000one-hot (too hard)
0.070.99980.0023CLIP default
0.500.59411.1013moderate
1.000.41441.3139soft
2.000.32771.3688nearly uniform

13. Something is wrong here: treating InfoNCE as a binary loss over pairs

Anomaly

Predict first

A student writes this, and it looks reasonable:

For each positive pair, compute loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.

For each positive pair, compute softmax over all B negatives in the batch simultaneously — this is F.cross_entropy(S/T, labels).

Why: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.

14. Trap: treating InfoNCE as a binary loss over pairs

Trap

The trap

For each positive pair, compute loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.

Treat it as B independent binary classification problems

Why: This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.

Problem: negatives in the batch are not shared; each pair ignores all other pairs in the batch, so large-batch benefits are lost.

The fix

For each positive pair, compute softmax over all B negatives in the batch simultaneously — this is F.cross_entropy(S/T, labels).

Use the B×B similarity matrix and row-wise softmax cross-entropy

Why: Each row is a B-class classification: 'which text in this batch matches image i?' In-batch negatives are free — no separate negative mining.

Large B means harder negatives (more confusable pairs), stronger gradient signal. CLIP used B=32768 (32k pairs per step). loss ≈ log(B) at random init.

15. Break it on purpose: treating InfoNCE as a binary loss over pairs

Break the constraint

Discussion prompt

The rule this trap just fixed:

Each row is a B-class classification: 'which text in this batch matches image i?' In-batch negatives are free — no separate negative mining.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This looks like contrastive learning but it is actually binary cross-entropy per pair — sigmoid-based, not softmax-based.

16. Dual-encoder architecture — image tower + text tower

Section

Part 2 of 4

17. CLIP dual-encoder structure

Concept

CLIP trains two independent encoders that share only the contrastive loss signal. The image encoder (ViT or ResNet, see Lesson 92) projects images to D-dim; the text encoder (GPT-like transformer) projects captions to the same D-dim.

componentinputoutput shapeCLIP-large config
Image encoderH×W×3 image(B, D)ViT-L/14, D=768
Text encodertokenized caption(B, D)GPT-style, 12-layer
L2 normalizeraw embeddings(B, D) unit vectorsapplied to both
Similarity Simg @ txt.T(B, B)scaled by 1/tau
InfoNCE lossS, labelsscalarsymmetric over rows+cols

18. Predict the next row: TinyCLIP forward pass in PyTorch

Pattern

Predict first

The table runs: Conv2d(1,8,3,pad=1) + ReLU | (B, 8, 8, 8) | input 8×8 image · AdaptiveAvgPool2d(2) | (B, 8, 2, 2) | → flatten → (B, 32) · Linear(32, 16) + L2-norm | (B, 16) | unit-norm image feats · Embedding(20,16).mean + Linear | (B, 16) | unit-norm text feats · img @ txt.T * 14.29 | (B, B) | logits for InfoNCE · F.cross_entropy (symmetric) | scalar | loss=2.6814 (random init)

In TinyCLIP forward pass in PyTorch, given the rows so far: what is the next one — the row where layer / step is Total params?

Correct: Total params | 1200 | 608 img + 592 txt

layer / stepoutput shapeverified value
Conv2d(1,8,3,pad=1) + ReLU(B, 8, 8, 8)input 8×8 image
AdaptiveAvgPool2d(2)(B, 8, 2, 2)→ flatten → (B, 32)
Linear(32, 16) + L2-norm(B, 16)unit-norm image feats
Embedding(20,16).mean + Linear(B, 16)unit-norm text feats
img @ txt.T * 14.29(B, B)logits for InfoNCE
F.cross_entropy (symmetric)scalarloss=2.6814 (random init)
Total params1200608 img + 592 txt

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Real CLIP uses ViT-L and GPT-style text transformer; TinyCLIP replaces them with a 3-layer CNN and a bag-of-embeddings to verify the loss pipeline in < 1 second.

19. TinyCLIP forward pass in PyTorch

Worked example

Define TinyImageEncoder (CNN) and TinyTextEncoder (embedding + mean pooling), both projecting to D=16

Why: Real CLIP uses ViT-L and GPT-style text transformer; TinyCLIP replaces them with a 3-layer CNN and a bag-of-embeddings to verify the loss pipeline in < 1 second.

import torch, torch.nn as nn, torch.nn.functional as F

class TinyImageEncoder(nn.Module):
    def __init__(self, d=16):
        super().__init__()
        self.conv = nn.Sequential(
            nn.Conv2d(1, 8, kernel_size=3, padding=1),
            nn.ReLU(),
            nn.AdaptiveAvgPool2d(2),     # -> (B,8,2,2)
        )
        self.proj = nn.Linear(8*2*2, d)  # 32 -> 16
    def forward(self, x):
        return F.normalize(self.proj(self.conv(x).flatten(1)), dim=-1)

class TinyTextEncoder(nn.Module):
    def __init__(self, vocab=20, d=16):
        super().__init__()
        self.emb  = nn.Embedding(vocab, d)
        self.proj = nn.Linear(d, d)
    def forward(self, tokens):
        return F.normalize(self.proj(self.emb(tokens).mean(1)), dim=-1)

# TinyImageEncoder params: 608  TinyTextEncoder params: 592  Total: 1200
img_enc, txt_enc = TinyImageEncoder(), TinyTextEncoder()
print(sum(p.numel() for p in img_enc.parameters()),   # 608
      sum(p.numel() for p in txt_enc.parameters()))   # 592
layer / stepoutput shapeverified value
Conv2d(1,8,3,pad=1) + ReLU(B, 8, 8, 8)input 8×8 image
AdaptiveAvgPool2d(2)(B, 8, 2, 2)→ flatten → (B, 32)
Linear(32, 16) + L2-norm(B, 16)unit-norm image feats
Embedding(20,16).mean + Linear(B, 16)unit-norm text feats
img @ txt.T * 14.29(B, B)logits for InfoNCE
F.cross_entropy (symmetric)scalarloss=2.6814 (random init)
Total params1200608 img + 592 txt

Run with B=6, random init: loss=2.6814, log(6)=1.7918

Why: Random-init loss > log(B) because T=14.29 creates very sharp distributions even before training — large T amplifies noise. With T=1/0.07≈14.29, the effective batch feels much larger to the loss. After training, loss converges toward 0.

20. Fill in: output shape for TinyCLIP forward pass in PyTorch

Comparison

Comparison matrix

From TinyCLIP forward pass in PyTorch: refill the output shape column from what you know. The rest of the table is as it appeared.

layer / stepoutput shapeverified value
Conv2d(1,8,3,pad=1) + ReLU(B, 8, 8, 8)input 8×8 image
AdaptiveAvgPool2d(2)(B, 8, 2, 2)→ flatten → (B, 32)
Linear(32, 16) + L2-norm(B, 16)unit-norm image feats
Embedding(20,16).mean + Linear(B, 16)unit-norm text feats
img @ txt.T * 14.29(B, B)logits for InfoNCE
F.cross_entropy (symmetric)scalarloss=2.6814 (random init)
Total params1200608 img + 592 txt

21. Zero-shot transfer — classify without a linear head

Section

Part 3 of 4

22. Zero-shot inference: text as a classifier

Concept

At inference CLIP needs no labeled training data for a new task. Encode each class name as a text prompt (e.g. 'a photo of a {class}'), compute all text embeddings, then find the class whose text embedding is most similar to the image embedding.

\[ \hat{y} = \arg\max_{k} \; z^I \cdot z^T_k \quad z^T_k = \text{TextEnc}(\texttt{"a photo of a "}+c_k) \]

class text promptcos_sim to imagesoftmax prob (T=0.07)
a cat0.21630.0027
a dog0.62960.9893
a car0.29230.0080
a bird-0.32700.0000
predicted—a dog (verified)

23. Why zero-shot works: the shared embedding space

Intuition

The InfoNCE loss forces image and text of the same concept to occupy nearby regions on a high-dimensional unit sphere. After training on 400 M internet pairs, 'a photo of a golden retriever' and an image of a golden retriever are neighbors — even for rare classes never seen as a classification task.

This is analogous to word2vec (Lesson 98): just as king - man + woman ≈ queen emerges from co-occurrence, CLIP's space captures visual-semantic analogies from co-occurrence of images and captions at web scale.

Prompt engineering matters: 'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.

24. Break it if you can: Why zero-shot works: the shared embedding space

Counterexample

Discussion prompt

Prompt engineering matters: 'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

25. Predict the next row: Zero-shot CLIP with TinyCLIP (PyTorch)

Pattern

Predict first

The table runs: a cat | 0.2163 | 0.0027 | no · a dog | 0.6296 | 0.9893 | YES · a car | 0.2923 | 0.0080 | no

In Zero-shot CLIP with TinyCLIP (PyTorch), given the rows so far: what is the next one — the row where class is a bird?

Correct: a bird | -0.3270 | 0.0000 | no

classcos_simsoftmax prob (T=0.07)predicted?
a cat0.21630.0027no
a dog0.62960.9893YES
a car0.29230.0080no
a bird-0.32700.0000no

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Using TinyCLIP (random weights) as a drop-in for real CLIP; structure is identical — only scale differs.

26. Zero-shot CLIP with TinyCLIP (PyTorch)

Worked example

Encode a query image and K=4 class text prompts; all outputs are unit vectors in R^16

Why: Using TinyCLIP (random weights) as a drop-in for real CLIP; structure is identical — only scale differs.

import torch, torch.nn.functional as F
# img_enc, txt_enc from TinyCLIP above
torch.manual_seed(7)

img_q  = F.normalize(torch.randn(1, 16), dim=-1)   # 1 query image
cls_embs = F.normalize(torch.randn(4, 16), dim=-1) # 4 class text embs
# nudge class 1 ("a dog") to be similar to the image
cls_embs[1] = F.normalize(img_q[0] + 0.3*torch.randn(16), dim=-1)

sims  = (img_q @ cls_embs.T).squeeze(0)      # (4,)
probs = F.softmax(sims / 0.07, dim=0)        # T=0.07
# predicted class:
pred = probs.argmax().item()                  # -> 1 ("a dog")

for name, s, p in zip(
    ['a cat','a dog','a car','a bird'],
     sims.tolist(), probs.tolist()):
    print(f'{name:12s}  cos={s:.4f}  prob={p:.4f}')
classcos_simsoftmax prob (T=0.07)predicted?
a cat0.21630.0027no
a dog0.62960.9893YES
a car0.29230.0080no
a bird-0.32700.0000no

argmax over probs selects index 1 ('a dog') with p=0.9893

Why: Even with T=0.07 (very sharp), a cosine difference of 0.6296 vs 0.2923 translates to a 99:1 probability ratio — zero-shot is decisive when embeddings are well-separated.

27. Which is which, by predicted?

Discrimination

Sort into buckets

Sort these by predicted?, from memory, without looking back at Zero-shot CLIP with TinyCLIP (PyTorch). Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

no
a cat; a car; a bird
YES
a dog
g1
predicted? is "no" for a cat, a car, a bird — that is what the table on "Zero-shot CLIP with TinyCLIP (PyTorch)" records, and it is the single property separating this group from the rest.
g2
predicted? is "YES" for a dog — that is what the table on "Zero-shot CLIP with TinyCLIP (PyTorch)" records, and it is the single property separating this group from the rest.

28. Something is wrong here: using inner product without L2 normalization

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The similarity matrix entries now reflect magnitude differences, not directional alignment.

Apply F.normalize(emb, dim=-1) to both img_emb and txt_emb before the dot product. Then S[i,j] is a cosine similarity in [-1, +1].

Why: The similarity matrix entries now reflect magnitude differences, not directional alignment. A high-norm pair dominates regardless of semantic similarity.

29. Trap: using inner product without L2 normalization

Trap

The trap

Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).

Scale by temperature and apply cross-entropy

Why: The similarity matrix entries now reflect magnitude differences, not directional alignment. A high-norm pair dominates regardless of semantic similarity.

Result: gradients push the network to increase embedding norms rather than learn direction — training collapses or diverges.

The fix

Apply F.normalize(emb, dim=-1) to both img_emb and txt_emb before the dot product. Then S[i,j] is a cosine similarity in [-1, +1].

Use img_emb = F.normalize(proj(h), dim=-1) at the end of each encoder

Why: Projection to the unit sphere decouples direction (semantics) from magnitude. The dot product now measures angular proximity, which is exactly what InfoNCE optimizes.

All values in S are in [-1, 1]. Temperature tau controls the effective scale — no raw-norm leakage.

30. Which of these survive contact with Lesson 113: CLIP — Contrastive…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
The diagonal entries are the B matching pairs; the off-diagonal entries are the B(B-1) non-matching pairs. CLIP pushes diagonal high and off-diagonal low.; Prompt engineering matters: 'a photo of a {c}' outperforms bare '{c}' because internet captions are natural sentences, not bare labels — the distribution shift is smaller.; CLIP opened the contrastive vision-language paradigm; subsequent models each extended it along one architectural dimension while keeping the core InfoNCE objective.
Breaks
For each positive pair, compute loss = -log(sigmoid(sim_pos)) and add log(1 - sigmoid(sim_neg)) for each negative.; Compute logits = img_emb @ txt_emb.T where img_emb and txt_emb come directly from the encoder projection head (not normalized).
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 113: CLIP — Contrastive Language-Image Pretraining puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. CLIP extensions — ALIGN, BLIP, BLIP-2

Section

Part 4 of 4

32. ALIGN, BLIP-2: three dimensions of extension

Concept

CLIP opened the contrastive vision-language paradigm; subsequent models each extended it along one architectural dimension while keeping the core InfoNCE objective.

modelkey extension over CLIPdata/computezero-shot IMAGENET top-1
CLIP (2021)dual ViT+GPT, clean curated 400M400M, large B76.2%
ALIGN (2021)EfficientNet+BERT, noisy 1.8B pairs (no cleaning)1.8B alt-text76.4%
BLIP (2022)caption + ITC + ITM objectives + web bootstrap129M + filtered82.3%
BLIP-2 (2023)frozen img encoder + Q-Former + frozen LLMfew trainable params73.4% (zero-shot)

ALIGN's insight: noise in alt-text is overcome by scale — 1.8 B noisy pairs outperforms 400 M curated pairs. BLIP-2's insight: freeze both encoders, train only the Q-Former bridge — most knowledge is already in pretrained weights.

33. What each one costs: ALIGN, BLIP-2: three dimensions of extension

Trade off

Comparison matrix

From ALIGN, BLIP-2: three dimensions of extension: every row here is a choice with a cost. Fill the zero-shot IMAGENET top-1 column, then say which row you would actually pick and what you give up for it.

modelkey extension over CLIPdata/computezero-shot IMAGENET top-1
CLIP (2021)dual ViT+GPT, clean curated 400M400M, large B76.2%
ALIGN (2021)EfficientNet+BERT, noisy 1.8B pairs (no cleaning)1.8B alt-text76.4%
BLIP (2022)caption + ITC + ITM objectives + web bootstrap129M + filtered82.3%
BLIP-2 (2023)frozen img encoder + Q-Former + frozen LLMfew trainable params73.4% (zero-shot)

34. CLIP for generation: conditioning diffusion models

Concept

CLIP text embeddings are used as conditioning vectors in diffusion models (DALL-E 2, Stable Diffusion). The text encoder is frozen; only the diffusion U-Net is trained to denoise images conditioned on the text embedding.

rolecomponenttrainable?
Text conditioningCLIP text encoderFrozen
Image conditioning (Img2Img)CLIP image encoderFrozen
DenoisingU-Net with cross-attentionTrained
ScheduleDDPM/DDIM noise scheduleFixed

Cross-attention keys/values come from the CLIP text embedding at every U-Net resolution; this is the architectural bridge that makes text-to-image work (Lesson 110+). CLIP is the semantic backbone, not the generative model.

35. Which is which, by trainable?

Discrimination

Sort into buckets

Sort these by trainable?, from memory, without looking back at CLIP for generation: conditioning diffusion models. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

Frozen
Text conditioning; Image conditioning (Img2Img)
Trained
Denoising
Fixed
Schedule
g1
trainable? is "Frozen" for Text conditioning, Image conditioning (Img2Img) — that is what the table on "CLIP for generation: conditioning…" records, and it is the single property separating this group from the rest.
g2
trainable? is "Trained" for Denoising — that is what the table on "CLIP for generation: conditioning…" records, and it is the single property separating this group from the rest.
g3
trainable? is "Fixed" for Schedule — that is what the table on "CLIP for generation: conditioning…" records, and it is the single property separating this group from the rest.

36. Rebuild the recipe: CLIP recipe: 6-step blueprint

Ranking

Put in order

These are the steps of CLIP recipe: 6-step blueprint, scrambled. Put them back in order before the next slide shows you.

  1. Encode: pass image through image encoder, caption through text encoder; both output (B, D) raw vectors
  2. Normalize: F.normalize(emb, dim=-1) on both to get unit-sphere embeddings
  3. Similarity matrix: S = img_emb @ txt_emb.T → (B, B) cosine similarities
  4. Scale: L = S / tau where tau = log_tau.exp(), log_tau is a learned parameter
  5. Symmetric InfoNCE: loss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)
  6. Zero-shot at inference: build one text embedding per class name, argmax cosine similarity with query image — no fine-tuning

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

37. CLIP recipe: 6-step blueprint

Pattern

  1. Encode: pass image through image encoder, caption through text encoder; both output (B, D) raw vectors
  2. Normalize: F.normalize(emb, dim=-1) on both to get unit-sphere embeddings
  3. Similarity matrix: S = img_emb @ txt_emb.T → (B, B) cosine similarities
  4. Scale: L = S / tau where tau = log_tau.exp(), log_tau is a learned parameter
  5. Symmetric InfoNCE: loss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)
  6. Zero-shot at inference: build one text embedding per class name, argmax cosine similarity with query image — no fine-tuning

38. Where does it stop working: CLIP recipe: 6-step blueprint

Edge cases

Discussion prompt

CLIP recipe: 6-step blueprint works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Encode: pass image through image encoder, caption through text encoder; both output (B, D) raw vectors
  2. Normalize: F.normalize(emb, dim=-1) on both to get unit-sphere embeddings
  3. Similarity matrix: S = img_emb @ txt_emb.T → (B, B) cosine similarities
  4. Scale: L = S / tau where tau = log_tau.exp(), log_tau is a learned parameter
  5. Symmetric InfoNCE: loss = (CE(L, labels) + CE(L.T, labels)) / 2; labels = torch.arange(B)
  6. Zero-shot at inference: build one text embedding per class name, argmax cosine similarity with query image — no fine-tuning

39. Rule out three: Check 1: InfoNCE loss structure

Elimination

Eliminate the wrong options

For a CLIP batch of B=1024 image-text pairs at random initialization (all cosine similarities ≈ 0), approximately what is the expected InfoNCE loss?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. log(1024) ≈ 6.93
  • B. log(2) ≈ 0.69 (binary classification)
  • C. 0.0 (random init means no loss)
  • D. 1.0 (cross-entropy always starts at 1)

Survives elimination: A

Why: Cross-entropy over B uniform logits = log(B). For B=1024, that is log(1024) ≈ 6.93. Verified: with zero similarities and B=4, loss = log(4) = 1.3863 (real execution). This shows why large batches are harder to optimize from scratch.

40. Check 1: InfoNCE loss structure

Check

Work this out before selecting. Consider what happens as batch size B grows.

Check your understanding

For a CLIP batch of B=1024 image-text pairs at random initialization (all cosine similarities ≈ 0), approximately what is the expected InfoNCE loss?

  • A. log(1024) ≈ 6.93 (correct)
  • B. log(2) ≈ 0.69 (binary classification)
  • C. 0.0 (random init means no loss)
  • D. 1.0 (cross-entropy always starts at 1)

Answer: A

Why: Cross-entropy over B uniform logits = log(B). For B=1024, that is log(1024) ≈ 6.93. Verified: with zero similarities and B=4, loss = log(4) = 1.3863 (real execution). This shows why large batches are harder to optimize from scratch.

Why B tempts people
InfoNCE is not a binary (positive vs one negative) loss — it is a B-class softmax where every other pair in the batch is a negative. The denominator sums over all B entries, not just 2.
Why C tempts people
Random init gives uniform logits, not zero loss. Uniform softmax (p=1/B for all) gives CE = -log(1/B) = log(B), not 0. Loss of 0 requires the diagonal to dominate perfectly.
Why D tempts people
Cross-entropy is -log(p_correct); for uniform p_correct = 1/B, the result depends on B. It equals 1 only if B ≈ e ≈ 2.72, not in general.

41. Answer it before you see the options: Check 2: Temperature effect

Prediction

Predict first

With similarity row [0.90, 0.30, 0.20, -0.10], increasing temperature from T=0.07 to T=2.0 causes the softmax probability of the correct class (index 0) to change from ≈0.9998 to ≈0.3277. What is the practical consequence for training?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Gradients become smaller and less discriminative; the model receives weaker signal to separate matching from non-matching pairs

Why: At T=2.0 the softmax is nearly uniform (p_correct=0.3277, entropy=1.3688), meaning the loss gradient is nearly the same regardless of which pair is matched — weak supervision. At T=0.07 (CLIP default), p_correct=0.9998 and the gradient focuses sharply on the few pairs that are wrong. Small T = hard training signal.

42. Check 2: Temperature effect

Check

Trace through the softmax formula before answering.

Check your understanding

With similarity row [0.90, 0.30, 0.20, -0.10], increasing temperature from T=0.07 to T=2.0 causes the softmax probability of the correct class (index 0) to change from ≈0.9998 to ≈0.3277. What is the practical consequence for training?

  • A. Gradients become smaller and less discriminative; the model receives weaker signal to separate matching from non-matching pairs (correct)
  • B. The loss decreases, so the model trains faster and reaches convergence sooner
  • C. All cosine similarities become equal, so training stops entirely
  • D. The model learns to predict harder negatives more accurately

Answer: A

Why: At T=2.0 the softmax is nearly uniform (p_correct=0.3277, entropy=1.3688), meaning the loss gradient is nearly the same regardless of which pair is matched — weak supervision. At T=0.07 (CLIP default), p_correct=0.9998 and the gradient focuses sharply on the few pairs that are wrong. Small T = hard training signal.

Why B tempts people
A nearly-uniform softmax produces higher loss (−log(0.3277) > −log(0.9998)), not lower. The model needs more steps, not fewer, to converge when the temperature is too large.
Why C tempts people
Temperature scales the logits but does not change the underlying cosine similarities, which depend on the encoder weights. Even at very high T the embeddings keep their directions; gradients still exist but are very small.
Why D tempts people
Higher temperature makes all pairs appear equally hard (nearly uniform), which actually obscures hard negatives rather than teaching the model to focus on them. Hard-negative mining requires a low temperature so genuinely confusable pairs produce larger loss.

43. Rule out three: Check 3: Zero-shot transfer

Elimination

Eliminate the wrong options

A CLIP model trained on internet image-caption pairs is tested zero-shot on a medical X-ray classification task with classes ['pleural effusion', 'pneumothorax', 'normal']. Which situation most likely LIMITS zero-shot performance?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Medical X-rays rarely appear with text descriptions in internet crawls, so image and text embeddings are not well-aligned in the medical domain
  • B. CLIP's text encoder cannot process words longer than 6 characters
  • C. The InfoNCE loss requires at least B=100 examples per class to generalize
  • D. Zero-shot requires fine-tuning the temperature parameter on the target dataset

Survives elimination: A

Why: CLIP's zero-shot strength depends on image-text co-occurrence during pretraining. X-ray images paired with clinical terminology are rare in internet text; the shared embedding space is poorly calibrated for this domain. This is the standard domain-shift argument: CLIP excels on natural-image distributions but degrades on out-of-distribution imagery.

44. Check 3: Zero-shot transfer

Check

Think about what the text encoder output represents before choosing.

Check your understanding

A CLIP model trained on internet image-caption pairs is tested zero-shot on a medical X-ray classification task with classes ['pleural effusion', 'pneumothorax', 'normal']. Which situation most likely LIMITS zero-shot performance?

  • A. Medical X-rays rarely appear with text descriptions in internet crawls, so image and text embeddings are not well-aligned in the medical domain (correct)
  • B. CLIP's text encoder cannot process words longer than 6 characters
  • C. The InfoNCE loss requires at least B=100 examples per class to generalize
  • D. Zero-shot requires fine-tuning the temperature parameter on the target dataset

Answer: A

Why: CLIP's zero-shot strength depends on image-text co-occurrence during pretraining. X-ray images paired with clinical terminology are rare in internet text; the shared embedding space is poorly calibrated for this domain. This is the standard domain-shift argument: CLIP excels on natural-image distributions but degrades on out-of-distribution imagery.

Why B tempts people
CLIP's text encoder is a full transformer (Lesson 88/89) with subword tokenization (BPE, Lesson 97) — it handles arbitrarily long medical terms like 'pneumothorax' without issue. Word length is not a CLIP limitation.
Why C tempts people
Zero-shot inference requires zero labeled examples per class — no training at all on the target task. The 'minimum examples' limitation applies to few-shot fine-tuning, not zero-shot CLIP which uses only text prompts.
Why D tempts people
Temperature is fixed at its pretrained value (log_tau learned during CLIP pretraining) at inference time. Zero-shot requires no target-dataset examples, so no parameter can be tuned — temperature included.

45. Plan first: Your turn: TinyCLIP full training step

Step zero

Discussion prompt

Your turn: TinyCLIP full training step — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Milestone 1 — build encoders and optimizer

Answer:

  1. Milestone 1 — build encoders and optimizer
  2. Milestone 2 — forward pass
  3. Milestone 3 — symmetric InfoNCE and backward
  4. Show-it-off: add prompt engineering — prepend 'a photo of a' to 5 synthetic class labels and run zero-shot argmax over a new test image

46. Your turn: TinyCLIP full training step

Worked example

Task: implement one complete CLIP training step (forward + backward + optimizer step) using TinyCLIP on a synthetic B=8 batch. Verify that loss decreases after the step.

Milestone 1 — build encoders and optimizer

Why: Instantiate TinyImageEncoder(d=16) and TinyTextEncoder(d=16). Create a single torch.optim.Adam over list(img_enc.parameters()) + list(txt_enc.parameters()) + [log_tau] where log_tau = nn.Parameter(torch.tensor(0.0)).

Milestone 2 — forward pass

Why: Generate imgs = torch.randn(8,1,8,8) and toks = torch.randint(0,20,(8,5)). Compute i_f = img_enc(imgs), t_f = txt_enc(toks), then L = (i_f @ t_f.T) / log_tau.exp(). Labels = torch.arange(8).

Milestone 3 — symmetric InfoNCE and backward

Why: Compute loss = (F.cross_entropy(L, labels) + F.cross_entropy(L.T, labels)) / 2. Call loss.backward() then optimizer.step(). Record loss before and after.

import torch, torch.nn as nn, torch.nn.functional as F

# --- re-use TinyImageEncoder / TinyTextEncoder from earlier ---
torch.manual_seed(0)
img_enc = TinyImageEncoder(d=16)
txt_enc = TinyTextEncoder(d=16)
log_tau = nn.Parameter(torch.tensor(0.0))  # learnable temperature

opt = torch.optim.Adam(
    list(img_enc.parameters()) +
    list(txt_enc.parameters()) + [log_tau], lr=1e-3)

imgs   = torch.randn(8, 1, 8, 8)
toks   = torch.randint(0, 20, (8, 5))
labels = torch.arange(8)

for step in range(3):
    opt.zero_grad()
    i_f = img_enc(imgs)
    t_f = txt_enc(toks)
    L   = (i_f @ t_f.T) / log_tau.exp()
    loss = (F.cross_entropy(L, labels) +
            F.cross_entropy(L.T, labels)) / 2
    loss.backward()
    opt.step()
    print(f'step {step}: loss={loss.item():.4f} tau={log_tau.exp().item():.4f}')
stepexpected lossexpected taunote
0 (before update)~2.0–2.81.0000random init, T=exp(0)=1
1decreasing~1.01optimizer adjusts tau upward
2further decrease>1.01both encoders + tau learn
convergence~log(8)≈2.08 floor → 0~0.07 targetCLIP trains for 32-epoch equivalents

Show-it-off: add prompt engineering — prepend 'a photo of a' to 5 synthetic class labels and run zero-shot argmax over a new test image

Why: After any training, the embedding directions improve. Even 3 steps on random data won't produce meaningful zero-shot, but the code structure is identical to real CLIP — only the encoder weights and data scale differ.

47. Fill in: expected loss for Your turn: TinyCLIP full training step

Comparison

Comparison matrix

From Your turn: TinyCLIP full training step: refill the expected loss column from what you know. The rest of the table is as it appeared.

stepexpected lossexpected taunote
0 (before update)~2.0–2.81.0000random init, T=exp(0)=1
1decreasing~1.01optimizer adjusts tau upward
2further decrease>1.01both encoders + tau learn
convergence~log(8)≈2.08 floor → 0~0.07 targetCLIP trains for 32-epoch equivalents

48. Connect it up: Lesson 113: CLIP — Contrastive Language-Image Pretraining

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Contrastive objective — matching pairs in a batch · Dual-encoder architecture — image tower + text tower · Zero-shot transfer — classify without a linear head · CLIP extensions — ALIGN, BLIP, BLIP-2. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

49. Lesson 113 recap: CLIP — what you can now do

Recap

conceptkey number / formulaverified
InfoNCE loss (B=4, random)log(4) = 1.3863yes
InfoNCE loss (B=3, near-perfect)0.0321yes
Zero-shot p(correct) T=0.070.9893 (cos=0.630)yes
Temperature: T=0.07 vs T=2.0p_correct: 0.9998 vs 0.3277yes
TinyCLIP params1200 (608 img + 592 txt)yes

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 113 — CLIP — Barron · USAAIO Round 2 Preparation, 2026
  2. Radford et al. 'Learning Transferable Visual Models From Natural Language Supervision' (ICML 2021) — arXiv:2103.00020
  3. Jia et al. 'Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision' (ICML 2021) — ALIGN — arXiv:2102.05918
  4. Li et al. 'BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models' (ICML 2023) — arXiv:2301.12597
  5. InfoNCE loss, temperature sharpness, zero-shot probabilities, and TinyCLIP (1200 params) verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108