Lesson 31: t-SNE & UMAP

USAAIO Lesson 31, from Week 11, fully worked. It starts from PCA's linear wall and the manifold hypothesis, then builds t-SNE from the ground up on a four-point running example: squared distances, the conditional similarities p_{j|i} with a per-point sigma, symmetrization into the joint P, perplexity as 2 to the power of the entropy, the Student-t low-dimensional kernel Q derived and contrasted with the Gaussian that caused crowding, and the KL(P||Q) objective with its gradient. Every P, Q, and KL number is recomputed by hand and reproduced from scratch in NumPy. UMAP is then contrasted as a graph method, PCA and t-SNE are run on the digits dataset - 0.61 against 0.98 by KNN - and the no-transform and geometry traps are proven in code. Every snippet runs standalone, and every number came from real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 106 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. t-SNE & UMAP

Title

USAAIO · Lesson 31 · Week 11 (Dimensionality Reduction)

When a linear projection isn't enough. We build t-SNE from the ground up on a 4-point example — similarities P in high-D, Q in 2-D, and the KL(P‖Q) they minimize — derive the Student-t kernel that fixes crowding, contrast UMAP, and prove in code why these maps are a microscope, not a feature extractor.

2. By the end of this lesson you can

Objectives

  1. State PCA's linear limitation and the manifold hypothesis, and say what a neighborhood-preserving method buys you
  2. Build the high-D similarity matrix P by hand — distances → conditional p_{j|i} → symmetrized joint — and reproduce it in NumPy
  3. Define perplexity as 2^{H(P_i)} and explain how the per-point σ_i is tuned to hit it
  4. Derive the Student-t low-D kernel Q, and show numerically why it cures the crowding that killed Gaussian SNE
  5. Write the t-SNE objective min D_KL(P‖Q), compute it entry by entry, and read off its gradient
  6. Contrast UMAP (graph-based, more global, faster) and prove in code why t-SNE/UMAP coords can't be reused downstream

3. What survived from Support Vector Machines?

Warm-up

Discussion prompt

Before we open Lesson 31: t-SNE & UMAP: without looking back, what was the main idea of Support Vector Machines, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the primal max-margin SVM, the dual QP in α with 0≤α≤C, why only support vectors define the boundary, the decision function f(x)=Σαᵢyᵢk(xᵢ,x)+b, and the soft-margin C trade-off. Fit an SVM, identify the support vectors, and reconstruct the decision function from the duals.

4. Beyond linear

Section

Part 1 of 8 — why PCA hits a wall

5. PCA can only rotate and flatten

Concept

PCA (Lesson 19) finds orthogonal directions of maximum variance and projects onto them. Every operation it does is linear — a rotation followed by dropping axes. It cannot bend space.

\[ Z = X W, \qquad W \in \mathbb{R}^{D \times 2} \;\; (\text{a fixed linear map}) \]

If the structure you care about is curved, a flat projection smears it. That is the wall this lesson climbs over.

6. Break it if you can: PCA can only rotate and flatten

Counterexample

Discussion prompt

PCA (Lesson 19) finds orthogonal directions of maximum variance and projects onto them. Every operation it does is linear — a rotation followed by dropping axes. It cannot bend space.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

If the structure you care about is curved, a flat projection smears it. That is the wall this lesson climbs over.

7. Picture it first: The swiss roll: distance lies

Picture it

Figure (svg): A spiral swiss-roll cross-section; two points on adjacent coils are near in straight-line distance but far along the coil.

PCA sees the short straight arrow; the true separation runs the long way along the coil.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture a sheet of paper rolled into a spiral. Two points on different layers of the roll are close in straight-line 3-D distance but far apart along the sheet — the distance that actually matters.

8. The swiss roll: distance lies

Intuition

Picture a sheet of paper rolled into a spiral. Two points on different layers of the roll are close in straight-line 3-D distance but far apart along the sheet — the distance that actually matters.

Figure (svg): A spiral swiss-roll cross-section; two points on adjacent coils are near in straight-line distance but far along the coil.

PCA sees the short straight arrow; the true separation runs the long way along the coil.

PCA keeps the misleading straight-line arrow. We want a method that keeps who your neighbors are on the sheet.

9. By analogy: The swiss roll: distance lies

Analogy

Discussion prompt

Explain The swiss roll: distance lies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

PCA keeps the misleading straight-line arrow. We want a method that keeps who your neighbors are on the sheet.

10. The manifold hypothesis

Concept

manifold hypothesis — Real high-dimensional data (images, text, audio) does not fill its ambient space — it concentrates on a much lower-dimensional, usually curved, surface (a manifold) sitting inside it.

A 28×28 image lives in ℝ^{784}, but the set of digit images is a thin, curved sheet inside that cube. Neighborhood-preserving methods are built to follow that sheet where a linear projection can't.

11. The one idea behind t-SNE

Intuition

Forget coordinates. Ask a softer question: for each point, which other points are its close neighbors, and how close? Turn that into a probability — 'if I'm at point i, how likely am I to pick j as a neighbor?'

Do this in the original space and again in a 2-D map. Then slide the 2-D points around until the two neighbor-probability tables match. That is the whole algorithm — the rest is making 'match' precise.

12. Why probabilities, not raw distances

Intuition

Why convert distances into probabilities at all? Because raw distances are on an arbitrary scale that differs between the 64-D space and the 2-D map — you can't compare ‖xᵢ−xⱼ‖ in one space to ‖yᵢ−yⱼ‖ in the other.

Normalizing each into a probability distribution puts both spaces on the same footing — every row sums to 1 — so 'make them match' becomes a well-defined objective (a divergence between distributions) instead of an apples-to-oranges distance comparison.

13. High-D similarities: P

Section

Part 2 of 8 — one running example

14. Our running example: four points

Concept

So every number stays checkable by hand, our 'high-D' space is a line with four points in two tight pairs: {0, 1} sit together, and {2, 3} sit together far away.

pointcoordinate xcoordinate y
000
110
250
360

The right answer for any neighbor method: keep 0–1 together and 2–3 together, with the two pairs apart. We'll watch the math discover exactly that.

15. Fill in: coordinate x for Our running example: four points

Comparison

Comparison matrix

From Our running example: four points: refill the coordinate x column from what you know. The rest of the table is as it appeared.

pointcoordinate xcoordinate y
000
110
250
360

16. Guess the shape of the answer: Step 1 — squared distances in high-D

Estimation

Predict first

t-SNE starts from pairwise squared Euclidean distances ‖xᵢ − xⱼ‖². With points on a line these are just squared coordinate gaps. Compute the whole matrix in NumPy:

Commit before you compute: what does Step 1 — squared distances in high-D come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Row 0: gaps to (1,5,6) are (1,5,6) → squares (1,25,36)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. d(0,1)²=1²=1, d(0,2)²=5²=25, d(0,3)²=6²=36.

17. Step 1 — squared distances in high-D

Worked example

t-SNE starts from pairwise squared Euclidean distances ‖xᵢ − xⱼ‖². With points on a line these are just squared coordinate gaps. Compute the whole matrix in NumPy:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
# broadcast: D2[i,j] = ||xi - xj||^2
D2 = np.sum((X[:,None,:] - X[None,:,:])**2, axis=2)
print(D2)

Row 0: gaps to (1,5,6) are (1,5,6) → squares (1,25,36)

Why: d(0,1)²=1²=1, d(0,2)²=5²=25, d(0,3)²=6²=36. The nearest neighbor of 0 is 1 by a wide margin — exactly the structure we want P to capture.

from \ to0123
0012536
1101625
2251601
3362510

18. What each one costs: Step 1 — squared distances in high-D

Trade off

Comparison matrix

From Step 1 — squared distances in high-D: every row here is a choice with a cost. Fill the 2 column, then say which row you would actually pick and what you give up for it.

from \ to0123
0012536
1101625
2251601
3362510

19. From distance to similarity probability

Concept

Turn each distance into a similarity with a Gaussian bump centered on point i: near points score high, far points score near zero. Then normalize row i so it's a probability distribution over the other points.

\[ p_{j\mid i} = \frac{\exp\!\big(-\lVert x_i - x_j\rVert^2 / 2\sigma_i^2\big)}{\sum_{k \neq i} \exp\!\big(-\lVert x_i - x_k\rVert^2 / 2\sigma_i^2\big)}, \qquad p_{i\mid i} = 0 \]

p_{j|i} reads: standing at i, the probability I'd pick j as my neighbor. The bandwidth σᵢ sets how far 'neighbor' reaches — we'll fix σᵢ² = 2 for now so the arithmetic is clean, and tune it properly in Part 3.

20. What has to be given first: Step 2 — conditional similarities p_{j|i}

Missing information

Discussion prompt

With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

= 0.778801, 0.001930, 0.000123. The denominator Z₀ = their sum = 0.780855, so p(1|0)=0.778801/0.780855=0.99737 — point 1 grabs essentially all of point 0's neighbor mass.

21. Step 2 — conditional similarities p_{j|i}

Worked example

With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2 / 4.0)          # 2*sigma^2 = 4
np.fill_diagonal(W, 0.0)       # p_{i|i} = 0
Pc = W / W.sum(axis=1, keepdims=True)   # normalize each ROW
print(Pc[0])

Row 0 numerators: exp(−1/4), exp(−25/4), exp(−36/4)

Why: = 0.778801, 0.001930, 0.000123. The denominator Z₀ = their sum = 0.780855, so p(1|0)=0.778801/0.780855=0.99737 — point 1 grabs essentially all of point 0's neighbor mass.

p_{j|0}j=0j=1j=2j=3
value00.997370.002470.00016
readingself=0the neighbor≈0≈0

22. Work backwards from the answer: Step 2 — conditional similarities p_{j|i}

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Row 0 numerators: exp(−1/4), exp(−25/4), exp(−36/4)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:

23. Something is wrong here: p_{j|i} is not symmetric

Anomaly

Predict first

A student writes this, and it looks reasonable:

Similarity is mutual, so surely p_{j|i} = p_{i|j} — one number per pair.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: They differ, because each row uses its OWN denominator Zᵢ.

Conditionals are row-normalized, so p_{j|i} and p_{i|j} generally differ. We fix this by symmetrizing in the next step.

Why: They differ, because each row uses its OWN denominator Zᵢ. Point 0's row sums to Z₀=0.78085 while point 2's row sums to Z₂=0.79905 — different totals. The unnormalized weight exp(−25/4) is shared, so p(2|0)=exp(−25/4)/0.78085=0.002472 but p(0|2)=exp(−25/4)/0.79905=0.002416. Same numerator, different Z ⇒ p(2|0)≠p(0|2).

24. Trap: p_{j|i} is not symmetric

Trap

The trap

Similarity is mutual, so surely p_{j|i} = p_{i|j} — one number per pair.

Assume p(2|0) = p(0|2)

Why: They differ, because each row uses its OWN denominator Zᵢ. Point 0's row sums to Z₀=0.78085 while point 2's row sums to Z₂=0.79905 — different totals. The unnormalized weight exp(−25/4) is shared, so p(2|0)=exp(−25/4)/0.78085=0.002472 but p(0|2)=exp(−25/4)/0.79905=0.002416. Same numerator, different Z ⇒ p(2|0)≠p(0|2).

The fix

Conditionals are row-normalized, so p_{j|i} and p_{i|j} generally differ. We fix this by symmetrizing in the next step.

Define the JOINT P_ij = (p_{j|i} + p_{i|j}) / (2n)

Why: Averaging the two directions and dividing by n makes P symmetric AND sum to 1 over all pairs — the object t-SNE actually matches. Never feed raw conditionals to the objective.

25. Predict the next row: Step 3 — symmetrize into the joint P

Pattern

Predict first

The table runs: 0 | 0 | 0.2465 | 0.0006 | 0.0000 · 1 | 0.2465 | 0 | 0.0057 | 0.0006 · 2 | 0.0006 | 0.0057 | 0 | 0.2465

In Step 3 — symmetrize into the joint P, given the rows so far: what is the next one — the row where P_ij is 3?

Correct: 3 | 0.0000 | 0.0006 | 0.2465 | 0

P_ij0123
000.24650.00060.0000
10.246500.00570.0006
20.00060.005700.2465
30.00000.00060.24650

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two within-pair entries P[0,1] and P[2,3] each carry ≈0.2465 of the total mass; cross-pair entries like P[1,2] are ≈0.0057.

26. Step 3 — symmetrize into the joint P

Worked example

Average the two directions and divide by 2n (here n = 4) so the whole matrix is symmetric and sums to 1:

\[ P_{ij} = \frac{p_{j\mid i} + p_{i\mid j}}{2n}, \qquad \sum_{i \neq j} P_{ij} = 1 \]

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n = len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
P = (Pc + Pc.T) / (2*n)      # symmetric joint
print(np.round(P, 6)); print('sum =', P.sum())

P[0,1] = (0.99737 + 0.974662)/8 = 0.246504

Why: The two within-pair entries P[0,1] and P[2,3] each carry ≈0.2465 of the total mass; cross-pair entries like P[1,2] are ≈0.0057. P encodes 'two tight pairs' numerically.

P_ij0123
000.24650.00060.0000
10.246500.00570.0006
20.00060.005700.2465
30.00000.00060.24650

27. Watch it run: Step 3 — symmetrize into the joint P

Pattern

Step through it

Step through Step 3 — symmetrize into the joint P one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: P_ij is 0
  2. Step 2: P_ij is 1
  3. Step 3: P_ij is 2
  4. Step 4: P_ij is 3

28. Perplexity

Section

Part 3 of 8 — the per-point bandwidth

29. Why every point needs its own σ

Concept

A dense cluster needs a small σ (neighbors are close); a sparse region needs a large σ (neighbors are far). One global bandwidth would over-smooth the dense parts and starve the sparse parts.

So t-SNE picks a separate σᵢ for each point — chosen so that every point ends up with roughly the same effective number of neighbors. That target number is the perplexity.

30. Perplexity = 2 to the entropy

Concept

The Shannon entropy H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:

\[ \mathrm{Perp}(P_i) = 2^{H(P_i)}, \qquad H(P_i) = -\!\sum_{j} p_{j\mid i}\,\log_2 p_{j\mid i} \]

If pᵢ spreads evenly over k neighbors, Perp = k. So perplexity is a smooth, differentiable stand-in for 'how many neighbors each point effectively has' — typically set between 5 and 50.

31. Teach it back: Perplexity = 2 to the entropy

Explain it

Discussion prompt

Explain Perplexity = 2 to the entropy to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The Shannon entropy H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:

32. Restore the missing line: Compute the entropy of a row

Fill the middle

Fill in the blanks

From Compute the entropy of a row — one line has had its right-hand side removed. Put it back.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
p0 =
Pc[0]; p0 = p0[p0 > 0]**
H = -np.sum(p0 * np.log2(p0))
print(round(H, 6), round(2**H, 6))

Why: p0 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Almost all mass sits on one neighbor, so entropy is near zero and perplexity near 1 — 'point 0 effectively has one neighbor'.

33. Compute the entropy of a row

Worked example

Perplexity is 2^{H}, so first compute the entropy H(P₀) of point 0's conditional row [0.99737, 0.00247, 0.00016] (excluding self). A near-one-hot distribution has tiny entropy. Self-contained:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
p0 = Pc[0]; p0 = p0[p0 > 0]
H = -np.sum(p0 * np.log2(p0))
print(round(H, 6), round(2**H, 6))

H(P₀) = 0.027195 bits → perplexity 2^H = 1.019

Why: Almost all mass sits on one neighbor, so entropy is near zero and perplexity near 1 — 'point 0 effectively has one neighbor'. This matches the σ²=2 row of the sweep table exactly.

quantityvalue
dominant p(1|0)0.99737
H(P₀) = −Σ p log₂ p0.027195 bits
perplexity = 2^H1.019029

34. Where does each piece belong: Lesson 31: t-SNE & UMAP

Sorting

Sort into buckets

These are the pieces of Lesson 31: t-SNE & UMAP, out of order. Put each one back under the part of the lesson it belongs to.

Beyond linear
PCA can only rotate and flatten; The swiss roll: distance lies; The manifold hypothesis
High-D similarities: P
Our running example: four points; Step 1 — squared distances in high-D; From distance to similarity probability
Perplexity
Why every point needs its own σ; Perplexity = 2 to the entropy; Compute the entropy of a row
s1
Beyond linear is where Lesson 31: t-SNE & UMAP puts PCA can only rotate and flatten, The swiss roll: distance lies, The manifold hypothesis. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
High-D similarities: P is where Lesson 31: t-SNE & UMAP puts Our running example: four points, Step 1 — squared distances in high-D, From distance to similarity probability. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Perplexity is where Lesson 31: t-SNE & UMAP puts Why every point needs its own σ, Perplexity = 2 to the entropy, Compute the entropy of a row. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

35. Finish it with less help: Perplexity rises with σ — traced

Faded example

Fill in the blanks

Perplexity rises with σ — traced, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])2, axis=2)
row =
D2[0]** # distances from point 0
def perplexity(sig2):
w = np.exp(-row/(2*sig2)); w[row==0] = 0.0
p = w/w.sum(); p = p[p>0]
return 2*(-np.sum(pnp.log2(p)))
for s2 in [0.5, 2.0, 8.0, 50.0]:
print(s2, round(perplexity(s2), 4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Perplexity climbs monotonically with σ².

36. Perplexity rises with σ — traced

Worked example

Fix point 0's distances and sweep σ². Small σ → all mass on the one nearest point (perplexity ≈ 1); large σ → mass spreads to all three others (perplexity → 3). Self-contained:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
row = D2[0]                       # distances from point 0
def perplexity(sig2):
    w = np.exp(-row/(2*sig2)); w[row==0] = 0.0
    p = w/w.sum(); p = p[p>0]
    return 2**(-np.sum(p*np.log2(p)))
for s2 in [0.5, 2.0, 8.0, 50.0]:
    print(s2, round(perplexity(s2), 4))

σ²: 0.5 → 1.000, 2 → 1.019, 8 → 2.062, 50 → 2.967

Why: Perplexity climbs monotonically with σ². At σ²=0.5 point 0 has one effective neighbor; at σ²=50 it has ~3 (all other points). This monotonicity is what lets a binary search hit any target.

σ²perplexity(row 0)effective neighbors
0.51.0000just point 1
2.01.0190≈ point 1
8.02.0619≈ 2 points
50.02.9671≈ all 3 others

37. What happens as it grows: Perplexity rises with σ — traced

Scale up

Step through it

Step through Perplexity rises with σ — traced and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: σ² is 0.5
  2. Step 2: σ² is 2.0
  3. Step 3: σ² is 8.0
  4. Step 4: σ² is 50.0

38. Finish it with less help: Binary-search σ to hit the target

Faded example

Fill in the blanks

Binary-search σ to hit the target, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
row = np.sum((X[0]-X)2, axis=1)**
def perp(sig2):
w = np.exp(-row/(2*sig2)); w[row==0]=0.0
p = w/w.sum(); p = p[p>0]
return 2*(-np.sum(pnp.log2(p)))
lo, hi = 1e-6, 1e6
for _ in range(60):
mid = (lo+hi)/2
if perp(mid) < 2.0: lo = mid
else: hi = mid
print(round(mid,4), round(perp(mid),4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The search brackets the answer and halves the interval 60 times.

39. Binary-search σ to hit the target

Worked example

Because perplexity increases monotonically in σ, t-SNE binary-searches σᵢ until Perp(Pᵢ) equals the requested value. Target perplexity 2.0 for point 0:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
row = np.sum((X[0]-X)**2, axis=1)
def perp(sig2):
    w = np.exp(-row/(2*sig2)); w[row==0]=0.0
    p = w/w.sum(); p = p[p>0]
    return 2**(-np.sum(p*np.log2(p)))
lo, hi = 1e-6, 1e6
for _ in range(60):
    mid = (lo+hi)/2
    if perp(mid) < 2.0: lo = mid
    else: hi = mid
print(round(mid,4), round(perp(mid),4))

Converges to σ² = 7.6062, achieving perplexity 2.0000

Why: The search brackets the answer and halves the interval 60 times. Real t-SNE runs exactly this search for EVERY point, giving each its own σᵢ before building P.

target perplexityfound σ²achieved perplexity
1.0 (min)→ 0→ 1
2.07.60622.0000
3.0 (max, n−1)→ ∞→ 3

40. Watch it run: Binary-search σ to hit the target

Pattern

Step through it

Step through Binary-search σ to hit the target one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: target perplexity is 1.0 (min)
  2. Step 2: target perplexity is 2.0
  3. Step 3: target perplexity is 3.0 (max, n−1)

41. Low-D Q & the t-kernel

Section

Part 4 of 8 — the map side

42. The map, and its similarities Q

Concept

Now the 2-D map. Each high-D point xᵢ gets a low-D position yᵢ (what we solve for). Build a similarity Q on the y's the same spirit as P — but with a different kernel and a single global normalization.

\[ q_{ij} = \frac{\big(1 + \lVert y_i - y_j\rVert^2\big)^{-1}}{\sum_{k \neq l}\big(1 + \lVert y_k - y_l\rVert^2\big)^{-1}} \]

That (1 + d²)⁻¹ is a Student-t distribution with one degree of freedom (a Cauchy). Why not a Gaussian like in high-D? That choice is the single most important idea in t-SNE — next.

43. The crowding problem

Intuition

In high dimensions there is room: a point can have many neighbors all roughly equidistant. Squash that into 2-D and there isn't enough space — moderately-distant points get shoved on top of each other. The map turns into a single blob.

The original SNE used a Gaussian in 2-D and suffered exactly this crowding. t-SNE's fix: give the low-D kernel a heavy tail so far-apart points can stay far apart without paying a huge similarity penalty.

44. SNE → t-SNE: two fixes

Concept

The original SNE (Hinton & Roweis, 2002) already matched Gaussian similarities by minimizing KL, but it was hard to optimize and crowded. t-SNE (2008) makes exactly two changes.

  1. Symmetrize the objective: match one joint P (not asymmetric conditionals per point) — a simpler gradient
  2. Swap the low-D kernel for a heavy-tailed Student-t — the crowding cure we're about to trace

Everything else — the per-point σ, the perplexity target, gradient descent on KL — is inherited from SNE. The 't' in t-SNE is literally the t-distribution.

45. Restore the missing line: Gaussian vs Student-t tails — traced

Fill the middle

Fill in the blanks

From Gaussian vs Student-t tails — traced — one line has had its right-hand side removed. Put it back.

import numpy as np
for d in [1., 2., 4., 8.]:
gaussian = np.exp(-d2 / 2.0) # SNE's 2-D kernel
student_t =
1.0 / (1.0 + d2) # t-SNE's 2-D kernel
print(int(d), round(gaussian,6), round(student_t,6))

Why: student_t is what everything below it consumes, so the wrong expression here fails later and somewhere else. The t-kernel assigns 175× more similarity to a distance-4 pair.

46. Gaussian vs Student-t tails — traced

Worked example

Compare the two kernels at growing distance. The Gaussian collapses to ~0; the Student-t decays gently. Self-contained:

import numpy as np
for d in [1., 2., 4., 8.]:
    gaussian  = np.exp(-d**2 / 2.0)   # SNE's 2-D kernel
    student_t = 1.0 / (1.0 + d**2)    # t-SNE's 2-D kernel
    print(int(d), round(gaussian,6), round(student_t,6))

At d=4: Gaussian = 0.000335, Student-t = 0.058824

Why: The t-kernel assigns 175× more similarity to a distance-4 pair. Under the Gaussian, a pair that far apart contributes almost nothing — so the optimizer yanks it inward to matter, causing crowding. The heavy tail removes that pull.

distance dGaussian e^{−d²/2}Student-t 1/(1+d²)ratio t/g
10.6065310.5000000.82×
20.1353350.2000001.48×
40.0003350.058824175×
80.0000000.015385huge

47. Watch it run: Gaussian vs Student-t tails — traced

Pattern

Step through it

Step through Gaussian vs Student-t tails — traced one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: distance d is 1
  2. Step 2: distance d is 2
  3. Step 3: distance d is 4
  4. Step 4: distance d is 8

48. Guess the shape of the answer: Same similarity, more room

Estimation

Predict first

Flip it around: to produce a target similarity q = 0.1, how far apart may a pair sit under each kernel? Solve e^{−d²/2}=0.1 and 1/(1+d²)=0.1:

Commit before you compute: what does Same similarity, more room come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Gaussian d = 2.146, Student-t d = 3.000 (1.40× farther)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For the SAME similarity, the t-kernel lets a pair sit 40% farther apart.

49. Same similarity, more room

Worked example

Flip it around: to produce a target similarity q = 0.1, how far apart may a pair sit under each kernel? Solve e^{−d²/2}=0.1 and 1/(1+d²)=0.1:

import numpy as np
q = 0.1
d_gauss = np.sqrt(-2*np.log(q))   # Gaussian: e^{-d^2/2}=q
d_t     = np.sqrt(1/q - 1)         # Student-t: 1/(1+d^2)=q
print(round(d_gauss,4), round(d_t,4), round(d_t/d_gauss,2))

Gaussian d = 2.146, Student-t d = 3.000 (1.40× farther)

Why: For the SAME similarity, the t-kernel lets a pair sit 40% farther apart. Multiplied across every moderate pair, that extra room is what spreads the clusters out instead of crushing them together.

kerneldistance to hit q=0.1effect on map
Gaussian (SNE)2.146clusters crowd inward
Student-t (t-SNE)3.000clusters breathe
gain1.40×uncrowded

50. Compute Q for a candidate map

Worked example

Place the four points on a line as a candidate map Y = [0, 0.5, 3, 3.5] — same pairing as high-D. Build Q with the t-kernel and a single global normalizer. Self-contained:

import numpy as np
Y = np.array([[0.],[0.5],[3.],[3.5]])
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0 + Dy2); np.fill_diagonal(Kt, 0.0)
Q = Kt / Kt.sum()          # ONE global denominator
print(np.round(Q, 6)); print('sum =', Q.sum())

q(0,1) = 0.8 / 4.026805 = 0.198669

Why: The kernel value for the close pair (d²=0.25) is 1/1.25=0.8; divided by the total 4.0268 gives 0.1987. Q is symmetric and sums to 1, just like P — now they are comparable.

Q_ij0123
000.19870.02480.0187
10.198700.03430.0248
20.02480.034300.1987
30.01870.02480.19870

51. What happens as it grows: Compute Q for a candidate map

Scale up

Step through it

Step through Compute Q for a candidate map and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: Q_ij is 0
  2. Step 2: Q_ij is 1
  3. Step 3: Q_ij is 2
  4. Step 4: Q_ij is 3

52. The objective: KL(P‖Q)

Section

Part 5 of 8 — matching the two tables

53. Kullback–Leibler divergence

Concept

We need a number that says how far apart the two neighbor tables P and Q are. That's the KL divergence (Lesson 14): the cost of using Q to describe data drawn from P.

\[ C = D_{KL}(P \,\|\, Q) = \sum_{i \neq j} P_{ij} \, \log \frac{P_{ij}}{Q_{ij}} \]

It is ≥ 0, and = 0 iff P = Q everywhere. t-SNE moves the y's to drive C down — that's the entire training loop.

54. KL is asymmetric — and that's a feature

Intuition

Look at a single term P_{ij} log(P_{ij}/Q_{ij}). It's weighted by P_{ij}. So a pair that is close in high-D (large P) is expensive to get wrong — the map is punished hard for pulling true neighbors apart.

But a pair that is far in high-D (P ≈ 0) contributes almost nothing, even if the map accidentally places them close. Net effect: t-SNE fiercely preserves local neighborhoods and is relaxed about global distances. Remember that asymmetry — it explains every t-SNE 'trap' later.

55. Compute KL(P‖Q) entry by entry

Worked example

Rebuild both P and Q from scratch, then sum P·log(P/Q) over all ordered pairs. Fully self-contained:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1, keepdims=True); P = (Pc+Pc.T)/(2*n)
Y = np.array([[0.],[0.5],[3.],[3.5]])
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
m = ~np.eye(n, dtype=bool)
print('KL =', round(np.sum(P[m]*np.log(P[m]/Q[m])), 6))

The four within-pair terms dominate at +0.0532 each

Why: P(0,1)=0.2465 > Q(0,1)=0.1987, so log(P/Q)>0 and the term is +0.0532; the four such ordered within-pair terms sum to +0.213. The cross-pair terms are slightly negative, netting KL = 0.182689.

pair (i,j)P_ijQ_ijP·log(P/Q)
(0,1) within0.24650.1987+0.053181
(2,3) within0.24650.1987+0.053181
(1,2) cross0.00570.0343−0.010246
(0,2) cross0.00060.0248−0.002264
Σ (all ordered pairs)——0.182689

56. Restore the missing line: A worse map has higher KL

Fill the middle

Fill in the blanks

From A worse map has higher KL — one line has had its right-hand side removed. Put it back.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1,keepdims=True); P = (Pc+Pc.T)/(2*n)
def kl_of(Y):
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
m = ~np.eye(n,dtype=bool)
return np.sum(P[m]*np.log(P[m]/Q[m]))
good = np.array([[0.],[0.5],[3.],[3.5]])
bad = np.array([[0.],[0.5],[0.2],[0.7]])
print(round(kl_of(good),4), round(kl_of(bad),4))

Why: good is what everything below it consumes, so the wrong expression here fails later and somewhere else. Overlapping the clusters puts high Q on pairs that have near-zero P — the KL almost sextuples.

57. A worse map has higher KL

Worked example

Prove the objective actually prefers the right layout. Take a bad map that overlaps the two clusters (Y = [0, 0.5, 0.2, 0.7]) and recompute KL against the same P:

import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1,keepdims=True); P = (Pc+Pc.T)/(2*n)
def kl_of(Y):
    Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
    Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
    m = ~np.eye(n,dtype=bool)
    return np.sum(P[m]*np.log(P[m]/Q[m]))
good = np.array([[0.],[0.5],[3.],[3.5]])
bad  = np.array([[0.],[0.5],[0.2],[0.7]])
print(round(kl_of(good),4), round(kl_of(bad),4))

good KL = 0.1827, bad KL = 1.0870

Why: Overlapping the clusters puts high Q on pairs that have near-zero P — the KL almost sextuples. Gradient descent on KL therefore pushes the layout AWAY from the bad map toward the separated one. The objective works.

candidate map YclustersKL(P‖Q)
[0, 0.5, 3, 3.5]separated (correct)0.1827
[0, 0.5, 0.2, 0.7]overlapping (wrong)1.0870
→ optimizer prefersseparatedlower KL

58. Fill in: KL(P‖Q) for A worse map has higher KL

Comparison

Comparison matrix

From A worse map has higher KL: refill the KL(P‖Q) column from what you know. The rest of the table is as it appeared.

candidate map YclustersKL(P‖Q)
[0, 0.5, 3, 3.5]separated (correct)0.1827
[0, 0.5, 0.2, 0.7]overlapping (wrong)1.0870
→ optimizer prefersseparatedlower KL

59. The gradient that moves the points

Concept

t-SNE minimizes C by gradient descent on the map positions yᵢ. The gradient has a clean, force-like form:

\[ \frac{\partial C}{\partial y_i} = 4 \sum_{j} (P_{ij} - Q_{ij})\,\big(1 + \lVert y_i - y_j\rVert^2\big)^{-1}\,(y_i - y_j) \]

Read it as springs: when P_{ij} > Q_{ij} (truer neighbors than the map shows) the term pulls yᵢ toward yⱼ; when P_{ij} < Q_{ij} it pushes them apart. The map settles where every spring balances.

60. Two knobs that make it converge

Concept

Raw gradient descent on KL gets stuck in tangled local minima. sklearn's TSNE uses two tricks by default so the map untangles cleanly.

trickdefaultwhat it does
early_exaggeration12.0scales P up early → clusters clump tight, then relax
learning_rateautostep size for the position updates
max_iter1000gradient-descent steps before stopping

Early exaggeration multiplies P by 12 for the first phase, over-emphasizing true neighbors so clusters form; then it's turned off to fine-tune spacing. You rarely touch these — but knowing they exist explains why a t-SNE run has phases.

61. Rebuild the recipe: The t-SNE algorithm, end to end

Ranking

Put in order

These are the steps of The t-SNE algorithm, end to end, scrambled. Put them back in order before the next slide shows you.

  1. Distances: compute ‖xᵢ − xⱼ‖² in high-D
  2. Bandwidths: binary-search each σᵢ so Perp(Pᵢ) = target perplexity
  3. High-D P: conditional p_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)
  4. Low-D Q: Student-t kernel (1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizer
  5. Optimize: gradient-descend min D_KL(P‖Q) on the positions yᵢ
  6. Read locally: trust who-is-near-whom; distrust cluster sizes and gaps

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

62. The t-SNE algorithm, end to end

Pattern

  1. Distances: compute ‖xᵢ − xⱼ‖² in high-D
  2. Bandwidths: binary-search each σᵢ so Perp(Pᵢ) = target perplexity
  3. High-D P: conditional p_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)
  4. Low-D Q: Student-t kernel (1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizer
  5. Optimize: gradient-descend min D_KL(P‖Q) on the positions yᵢ
  6. Read locally: trust who-is-near-whom; distrust cluster sizes and gaps

63. Where does it stop working: The t-SNE algorithm, end to end

Edge cases

Discussion prompt

The t-SNE algorithm, end to end works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Distances: compute ‖xᵢ − xⱼ‖² in high-D
  2. Bandwidths: binary-search each σᵢ so Perp(Pᵢ) = target perplexity
  3. High-D P: conditional p_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)
  4. Low-D Q: Student-t kernel (1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizer
  5. Optimize: gradient-descend min D_KL(P‖Q) on the positions yᵢ
  6. Read locally: trust who-is-near-whom; distrust cluster sizes and gaps

64. UMAP

Section

Part 6 of 8 — the graph view

65. UMAP: a fuzzy neighbor graph

Concept

UMAP reaches the same goal from graph theory. Build a weighted k-nearest-neighbor graph in high-D (edge weights = fuzzy membership strengths), then lay that graph out in 2-D so its edge structure is preserved.

Its low-D objective is a cross-entropy over edges (both present and absent), optimized with negative sampling — closer in spirit to how word embeddings train than to t-SNE's KL, but the neighbor-preserving intent is the same.

66. The neighbor graph on our four points

Worked example

UMAP's first move is a k-nearest-neighbor graph. Build it on our running example with k=1 (each point's single nearest non-self neighbor) and read off the edges. Self-contained:

import numpy as np
from sklearn.neighbors import NearestNeighbors
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
nn = NearestNeighbors(n_neighbors=2).fit(X)   # 2 = self + 1
dist, idx = nn.kneighbors(X)
for i in range(4):
    print(i, '->', idx[i][1], 'dist', round(dist[i][1], 3))

Edges: 0↔1 and 2↔3, each at distance 1

Why: The graph recovers the two tight pairs exactly — same neighborhood structure P encoded, now as explicit edges. UMAP then lays out THIS graph in 2-D, minimizing an edge cross-entropy rather than a similarity KL.

pointnearest neighboredge distance
011.000
101.000
231.000
321.000

67. What stays fixed: The neighbor graph on our four points

Invariant

Step through it

Step through The neighbor graph on our four points one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: point is 0
  2. Step 2: point is 1
  3. Step 3: point is 2
  4. Step 4: point is 3

68. UMAP vs t-SNE, honestly

Concept

axist-SNEUMAP
knobperplexity (5–50)n_neighbors + min_dist
objectiveKL of similaritiesgraph cross-entropy
global structureweaksomewhat better
speed / scaleslower (O(N log N) w/ Barnes-Hut)faster, scales to millions
new pointsno transformapproximate transform, still shaky

UMAP preserves a bit more global layout and runs faster, so it's the default for large datasets. But it shares t-SNE's core caveats — the map is still for eyes, not features.

69. Teach it back: UMAP vs t-SNE, honestly

Explain it

Discussion prompt

Explain UMAP vs t-SNE, honestly to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

UMAP preserves a bit more global layout and runs faster, so it's the default for large datasets. But it shares t-SNE's core caveats — the map is still for eyes, not features.

70. The knobs change the story

Concept

t-SNE's perplexity and UMAP's n_neighbors both set the local-vs-global balance. Small values emphasize tight local clusters; large values reveal coarser global shape.

Extreme settings produce misleading maps — a too-small perplexity can shatter one real cluster into fake shards, a too-large one can merge distinct clusters. Always sweep a few values before trusting a picture.

71. Which method when

Concept

methodreach for it when you need
PCAfast linear reduction, reusable features for ML, denoising
t-SNEhigh-quality 2-D pictures of local cluster structure
UMAPfaster viz, a bit more global structure, larger datasets

One rule of thumb: PCA is the workhorse for features; t-SNE and UMAP are for the human eye. The next slides prove — in code — why you must not confuse the two.

72. By analogy: Which method when

Analogy

Discussion prompt

Explain Which method when by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

One rule of thumb: PCA is the workhorse for features; t-SNE and UMAP are for the human eye. The next slides prove — in code — why you must not confuse the two.

73. Reading & reusing the map

Section

Part 7 of 8 — the traps in code

74. Predict the next row: PCA vs t-SNE on digits

Pattern

Predict first

The table runs: raw pixels | 64 | 0.963 · PCA (linear) | 2 | 0.609

In PCA vs t-SNE on digits, given the rows so far: what is the next one — the row where representation is t-SNE (nonlinear)?

Correct: t-SNE (nonlinear) | 2 | 0.980

representationdimsKNN accuracy
raw pixels640.963
PCA (linear)20.609
t-SNE (nonlinear)20.980

Why: The relationship between the columns, not the individual numbers, is what generates the next row. t-SNE's neighborhood preservation packs each digit into a tight, separable island — its 2-D map is even MORE KNN-separable than the raw 64-D pixels (0.98 vs 0.96), while PCA's flat projection overlaps several digits down to 0.61.

75. PCA vs t-SNE on digits

Worked example

The payoff, on real data. Project the 64-D digits set to 2-D two ways, then measure how separable the ten classes are with a 3-fold KNN classifier on each 2-D embedding. Self-contained:

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca  = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
k = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('raw64', k(X), 'pca2', k(pca), 'tsne2', k(tsne))

raw 64-D = 0.963, PCA 2-D = 0.609, t-SNE 2-D = 0.980

Why: t-SNE's neighborhood preservation packs each digit into a tight, separable island — its 2-D map is even MORE KNN-separable than the raw 64-D pixels (0.98 vs 0.96), while PCA's flat projection overlaps several digits down to 0.61.

representationdimsKNN accuracy
raw pixels640.963
PCA (linear)20.609
t-SNE (nonlinear)20.980

76. What has to be given first: Proof: t-SNE has no transform

Missing information

Discussion prompt

That 0.98 is seductive — why not use the t-SNE coordinates as features? Because there is no mapping for new points. PCA learns a reusable projection; t-SNE learns only where these points go. Check the API itself:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

PCA stores components you can apply to unseen rows. TSNE only implements fit_transform on the training set — there is literally no method to embed a new point, because its position is defined only relative to the points it was optimized with.

77. Proof: t-SNE has no transform

Worked example

That 0.98 is seductive — why not use the t-SNE coordinates as features? Because there is no mapping for new points. PCA learns a reusable projection; t-SNE learns only where these points go. Check the API itself:

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
X, y = load_digits(return_X_y=True)
print('PCA  has transform:', hasattr(PCA(2).fit(X), 'transform'))
print('TSNE has transform:', hasattr(TSNE(2), 'transform'))

PCA.transform exists (True); TSNE.transform does not (False)

Why: PCA stores components you can apply to unseen rows. TSNE only implements fit_transform on the training set — there is literally no method to embed a new point, because its position is defined only relative to the points it was optimized with.

estimatorhas .transform?usable on new data?
PCATrueyes — reusable features
TSNEFalseno — training set only
UMAPapproximateshaky, not for production

78. Guess the shape of the answer: Proof: the layout is non-deterministic

Estimation

Predict first

Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:

Commit before you compute: what does Proof: the layout is non-deterministic come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: max coordinate difference between seeds ≈ 130.6

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The same data, same settings, different seed → a completely different coordinate system (the local neighborhoods are similar, but absolute positions swing by ~130 units).

79. Proof: the layout is non-deterministic

Worked example

Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:

import numpy as np
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
X, y = load_digits(return_X_y=True)
a = TSNE(2, init='random', perplexity=30, random_state=0).fit_transform(X)
b = TSNE(2, init='random', perplexity=30, random_state=1).fit_transform(X)
print('max coordinate diff:', round(np.abs(a-b).max(), 1))

max coordinate difference between seeds ≈ 130.6

Why: The same data, same settings, different seed → a completely different coordinate system (the local neighborhoods are similar, but absolute positions swing by ~130 units). Features that change this much between runs cannot go into a model.

runseedoutcome
A0one arbitrary orientation
B1different orientation/positions
max |A − B|—≈ 130.6 (huge)

80. Work backwards from the answer: Proof: the layout is non-deterministic

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

max coordinate difference between seeds ≈ 130.6

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:

81. Something is wrong here: reuse t-SNE coords downstream

Anomaly

Predict first

A student writes this, and it looks reasonable:

t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless.

Use t-SNE for visualization and exploration only. Model on PCA, raw, or learned features that have a stable transform.

Why: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.

82. Trap: reuse t-SNE coords downstream

Trap

The trap

t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.

model.fit(tsne_coords, y) → ship it

Why: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.

The fix

Use t-SNE for visualization and exploration only. Model on PCA, raw, or learned features that have a stable transform.

Visualize with t-SNE; model on PCA / raw / embeddings

Why: PCA's .transform embeds new points deterministically, so its coordinates generalize. t-SNE is a microscope for looking at your data, not a feature extractor for a pipeline.

83. Break it on purpose: reuse t-SNE coords downstream

Break the constraint

Discussion prompt

The rule this trap just fixed:

PCA's .transform embeds new points deterministically, so its coordinates generalize. t-SNE is a microscope for looking at your data, not a feature extractor for a pipeline.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.

84. Something is wrong here: reading the geometry literally

Anomaly

Predict first

A student writes this, and it looks reasonable:

Cluster B is twice the diameter of cluster A and sits far from it, so B's items are more varied and the two groups are very different.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary.

Read only local neighborhood structure. Treat cluster sizes and global distances as unreliable.

Why: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary. A 'bigger' or 'farther' cluster in the map may be neither in the data.

85. Trap: reading the geometry literally

Trap

The trap

Cluster B is twice the diameter of cluster A and sits far from it, so B's items are more varied and the two groups are very different.

Interpret cluster sizes and between-cluster gaps as real

Why: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary. A 'bigger' or 'farther' cluster in the map may be neither in the data.

The fix

Read only local neighborhood structure. Treat cluster sizes and global distances as unreliable.

Say 'these points are neighbors', not 'this cluster is bigger/farther'

Why: The method optimizes local similarity by construction. UMAP preserves a little more global layout, but the same caution holds — never quantify from the picture.

86. Which of these survive contact with Lesson 31: t-SNE & UMAP?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
If the structure you care about is curved, a flat projection smears it. That is the wall this lesson climbs over.; PCA keeps the misleading straight-line arrow. We want a method that keeps who your neighbors are on the sheet.; The Shannon entropy H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:
Breaks
Similarity is mutual, so surely p_{j|i} = p_{i|j} — one number per pair.; t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 31: t-SNE & UMAP puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

87. Rule out three: Check yourself — the objective

Elimination

Eliminate the wrong options

t-SNE finds its 2-D layout by minimizing:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. KL divergence between the high-D and 2-D pairwise-similarity distributions, D_KL(P‖Q)
  • B. the reconstruction error of a linear projection
  • C. the variance captured by each output axis
  • D. the sum of squared differences of all pairwise distances

Survives elimination: A

Why: t-SNE converts distances to similarity probabilities P (high-D, Gaussian) and Q (2-D, Student-t), then gradient-descends the positions to minimize D_KL(P‖Q) = Σ P log(P/Q) — preserving local neighborhoods.

88. Check yourself — the objective

Check

Back to first principles: what number is t-SNE driving down?

Check your understanding

t-SNE finds its 2-D layout by minimizing:

  • A. KL divergence between the high-D and 2-D pairwise-similarity distributions, D_KL(P‖Q) (correct)
  • B. the reconstruction error of a linear projection
  • C. the variance captured by each output axis
  • D. the sum of squared differences of all pairwise distances

Answer: A

Why: t-SNE converts distances to similarity probabilities P (high-D, Gaussian) and Q (2-D, Student-t), then gradient-descends the positions to minimize D_KL(P‖Q) = Σ P log(P/Q) — preserving local neighborhoods.

Why B tempts people
Minimizing reconstruction error of a linear projection is PCA. t-SNE is nonlinear and probability-based; it never reconstructs the inputs.
Why C tempts people
Maximizing captured variance per axis is PCA's objective. t-SNE has no notion of variance-per-axis — its axes are not even meaningfully oriented.
Why D tempts people
t-SNE deliberately does NOT preserve all pairwise distances (that's classical MDS). Its P-weighted KL preserves local similarity and distorts global distance on purpose.

89. Answer it before you see the options: Check yourself — why Student-t

Prediction

Predict first

Why does t-SNE use a heavy-tailed Student-t kernel in the 2-D map instead of a Gaussian?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Its heavy tail lets moderately-distant points stay far apart cheaply, relieving the crowding that plagued Gaussian SNE

Why: In 2-D there isn't room for all the high-D neighbors. The Gaussian's tail vanishes (0.0003 at d=4), so distant pairs contribute nothing and get pulled inward — crowding. The Student-t's 1/(1+d²) tail (0.059 at d=4) keeps distant pairs meaningful, so clusters spread out.

90. Check yourself — why Student-t

Check

Recall the tail comparison you traced.

Check your understanding

Why does t-SNE use a heavy-tailed Student-t kernel in the 2-D map instead of a Gaussian?

  • A. Its heavy tail lets moderately-distant points stay far apart cheaply, relieving the crowding that plagued Gaussian SNE (correct)
  • B. The Student-t is faster to evaluate than a Gaussian
  • C. It makes Q sum to 1, which a Gaussian cannot
  • D. It guarantees the clusters end up the same size

Answer: A

Why: In 2-D there isn't room for all the high-D neighbors. The Gaussian's tail vanishes (0.0003 at d=4), so distant pairs contribute nothing and get pulled inward — crowding. The Student-t's 1/(1+d²) tail (0.059 at d=4) keeps distant pairs meaningful, so clusters spread out.

Why B tempts people
Evaluation speed isn't the reason; both are cheap. The motivation is purely the tail shape and its effect on crowding.
Why C tempts people
Any positive kernel can be normalized to sum to 1 — a Gaussian Q sums to 1 too. Normalization is not why the t-kernel was chosen.
Why D tempts people
Nothing forces equal cluster sizes; in fact t-SNE distorts cluster sizes. The t-kernel is about the tail, not about equalizing diameters.

91. Answer it before you see the options: Check yourself — downstream use

Prediction

Predict first

Why are t-SNE coordinates a poor choice as features for a classifier that will see new data?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: There is no stable transform for new points, the layout is non-deterministic, and global geometry is distorted

Why: TSNE has no .transform (proved: hasattr → False), two seeds differ by ~130 units, and its inter-cluster distances are unreliable — so the coordinates don't generalize to unseen points.

92. Check yourself — downstream use

Check

You proved two of these in code.

Check your understanding

Why are t-SNE coordinates a poor choice as features for a classifier that will see new data?

  • A. There is no stable transform for new points, the layout is non-deterministic, and global geometry is distorted (correct)
  • B. They have too many dimensions to model efficiently
  • C. They are always perfectly linearly separable, which causes overfitting
  • D. scikit-learn raises an error if you fit a model on them

Answer: A

Why: TSNE has no .transform (proved: hasattr → False), two seeds differ by ~130 units, and its inter-cluster distances are unreliable — so the coordinates don't generalize to unseen points.

Why B tempts people
t-SNE outputs FEW dimensions (usually 2). The problem is generalization to new points, not an excess of dimensions.
Why C tempts people
The strong in-sample separability (0.98) is exactly what tempts people; the flaw is that it won't hold for unseen data, not that separability itself overfits.
Why D tempts people
sklearn will happily fit a model on any array — it won't error. The obstruction is conceptual (no out-of-sample map), not a library guard.

93. Rule out three: Check yourself — pick the tool

Elimination

Eliminate the wrong options

You need 2-D features to feed a model that will later score new, unseen data. Best choice?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. PCA — it learns a projection (.transform) you can apply to new points
  • B. t-SNE — it separates the training clusters best
  • C. UMAP — it's the fastest of the three
  • D. Any of them; for 2-D features they're interchangeable

Survives elimination: A

Why: PCA fits a linear map and stores it, so .transform embeds unseen rows deterministically — the requirement for reusable features. t-SNE/UMAP are visualization tools without a reliable out-of-sample mapping.

94. Check yourself — pick the tool

Check

Match the method to the job.

Check your understanding

You need 2-D features to feed a model that will later score new, unseen data. Best choice?

  • A. PCA — it learns a projection (.transform) you can apply to new points (correct)
  • B. t-SNE — it separates the training clusters best
  • C. UMAP — it's the fastest of the three
  • D. Any of them; for 2-D features they're interchangeable

Answer: A

Why: PCA fits a linear map and stores it, so .transform embeds unseen rows deterministically — the requirement for reusable features. t-SNE/UMAP are visualization tools without a reliable out-of-sample mapping.

Why B tempts people
t-SNE's separation is in-sample only and has no transform, so it fails the moment new data arrives — despite the tempting training-set score.
Why C tempts people
Speed isn't the criterion here; UMAP's transform is only approximate and unstable, so it's still the wrong pick for production features.
Why D tempts people
They are NOT interchangeable: only PCA gives a stable, reusable projection for new data. That is the whole point of this lesson.

95. Your turn: compare embeddings

Section

Part 8 of 8 — the project

96. Project: PCA vs t-SNE separability

Concept

Embed the digits set to 2-D with both PCA and t-SNE, score each with a KNN classifier, and draw the right conclusion about what each tool is for. You derived every piece — now assemble it.

#requirementtool
1PCA to 2-D + KNN scorePCA(n_components=2)
2t-SNE to 2-D + KNN scoreTSNE(perplexity=30)
3Compare, and state the correct use of eachcross_val_score

Build rules: type every line yourself, always set random_state for t-SNE (it's stochastic), and remember the KNN number is a separability diagnostic, not a deployable model.

97. Break it if you can: Project: PCA vs t-SNE separability

Counterexample

Discussion prompt

Embed the digits set to 2-D with both PCA and t-SNE, score each with a KNN classifier, and draw the right conclusion about what each tool is for. You derived every piece — now assemble it.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, always set random_state for t-SNE (it's stochastic), and remember the KNN number is a separability diagnostic, not a deployable model.

98. Milestone 1 — the PCA embedding

Worked example

Your turn: project digits to 2-D with PCA and score a 3-fold KNN on it. Predict out loud: above or below 0.7?

Hint: PCA(n_components=2).fit_transform(X), then cross_val_score(KNeighborsClassifier(), pca, y, cv=3).mean().

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca = PCA(n_components=2).fit_transform(X)
print(round(cross_val_score(KNeighborsClassifier(), pca, y, cv=3).mean(), 3))
embeddingdimsKNN accuracy
PCA 2-D20.609

99. Milestone 2 — the t-SNE embedding

Worked example

Your turn: project to 2-D with t-SNE and score the same KNN. Predict the jump over PCA before you run it.

Hint: TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X). The random_state makes it reproducible.

from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
print(round(cross_val_score(KNeighborsClassifier(), tsne, y, cv=3).mean(), 3))
embeddingdimsKNN accuracy
t-SNE 2-D20.980

100. Milestone 3 — compare and conclude

Worked example

Your turn: put the two numbers (and the raw 64-D baseline) side by side. Which embedding preserves class structure in 2-D — and which would you actually deploy?

Hint: score X itself for the raw baseline. Expect t-SNE ≳ raw ≫ PCA on separability — yet PCA is the one you'd ship.

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca  = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
k = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('raw', k(X), 'pca', k(pca), 'tsne', k(tsne))
embeddingKNN accuracycorrect use
raw 64-D0.963baseline
PCA 2-D0.609reusable features (has transform)
t-SNE 2-D0.980visualization only (no transform)

101. What each one costs: Milestone 3 — compare and conclude

Trade off

Comparison matrix

From Milestone 3 — compare and conclude: every row here is a choice with a cost. Fill the correct use column, then say which row you would actually pick and what you give up for it.

embeddingKNN accuracycorrect use
raw 64-D0.963baseline
PCA 2-D0.609reusable features (has transform)
t-SNE 2-D0.980visualization only (no transform)

102. The full program

Concept

from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score

X, y = load_digits(return_X_y=True)
pca  = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
knn = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('PCA 2D:', knn(pca), ' t-SNE 2D:', knn(tsne))
print('PCA has transform:', hasattr(PCA(2).fit(X), 'transform'),
      '| TSNE has transform:', hasattr(TSNE(2), 'transform'))
printed linevalue
PCA 2D:0.609
t-SNE 2D:0.98
PCA has transform / TSNE has transformTrue / False

If t-SNE's 2-D map is far more separable than PCA's — yet you'd still model on PCA because only it has a transform — you know exactly what each tool is for.

103. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
PCA 2D:0.609
t-SNE 2D:0.98
PCA has transform / TSNE has transformTrue / False

104. Show it off

Concept

Slides closed, out loud: explain (1) which two distributions t-SNE matches and with which divergence, (2) why the Student-t tail prevents crowding, and (3) two concrete reasons you can't feed t-SNE coordinates to a downstream model.

Stretch (homework): scatter-plot the t-SNE map colored by digit, sweep perplexity to 5 and 200 and watch the map distort, then try UMAP (umap-learn). These same tools will reappear when you inspect latent spaces (Week 43) and learned embeddings (Week 47).

105. Connect it up: Lesson 31: t-SNE & UMAP

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Beyond linear · High-D similarities: P · Perplexity · Low-D Q & the t-kernel · The objective: KL(P‖Q) · UMAP. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

106. What you can do now

Recap

ideathe one thing to remember
high-D PGaussian similarities, per-point σ, symmetrized to sum 1
perplexity2^{entropy} = effective # neighbors; tune σ to it
low-D QStudent-t tail beats crowding — clusters breathe
objectiveminimize KL(P‖Q); P-weighted ⇒ local, not global
UMAPgraph cross-entropy: more global, faster
caveatviz only — no transform, non-deterministic geometry

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 31 (Week 11 — t-SNE/UMAP) — Barron · USAAIO Round 2 Preparation, 2026
  2. van der Maaten & Hinton, Visualizing Data using t-SNE — JMLR 9 (2008) 2579–2605
  3. McInnes, Healy & Melville, UMAP: Uniform Manifold Approximation and Projection — arXiv:1802.03426, 2018
  4. scikit-learn manifold.TSNE / decomposition.PCA
  5. Every P/Q/KL entry, perplexity, and KNN accuracy produced by real execution — numpy 2.2 + scikit-learn 1.9, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108