USAAIO Lesson 31, from Week 11, fully worked. It starts from PCA's linear wall and the manifold hypothesis, then builds t-SNE from the ground up on a four-point running example: squared distances, the conditional similarities p_{j|i} with a per-point sigma, symmetrization into the joint P, perplexity as 2 to the power of the entropy, the Student-t low-dimensional kernel Q derived and contrasted with the Gaussian that caused crowding, and the KL(P||Q) objective with its gradient. Every P, Q, and KL number is recomputed by hand and reproduced from scratch in NumPy. UMAP is then contrasted as a graph method, PCA and t-SNE are run on the digits dataset - 0.61 against 0.98 by KNN - and the no-transform and geometry traps are proven in code. Every snippet runs standalone, and every number came from real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 106 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 31 · Week 11 (Dimensionality Reduction)
When a linear projection isn't enough. We build t-SNE from the ground up on a 4-point example — similarities P in high-D, Q in 2-D, and the KL(P‖Q) they minimize — derive the Student-t kernel that fixes crowding, contrast UMAP, and prove in code why these maps are a microscope, not a feature extractor.
Objectives
P by hand — distances → conditional p_{j|i} → symmetrized joint — and reproduce it in NumPy2^{H(P_i)} and explain how the per-point σ_i is tuned to hit itQ, and show numerically why it cures the crowding that killed Gaussian SNEmin D_KL(P‖Q), compute it entry by entry, and read off its gradientWarm-up
Discussion prompt
Before we open Lesson 31: t-SNE & UMAP: without looking back, what was the main idea of Support Vector Machines, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the primal max-margin SVM, the dual QP in α with 0≤α≤C, why only support vectors define the boundary, the decision function f(x)=Σαᵢyᵢk(xᵢ,x)+b, and the soft-margin C trade-off. Fit an SVM, identify the support vectors, and reconstruct the decision function from the duals.
Section
Part 1 of 8 — why PCA hits a wall
Concept
PCA (Lesson 19) finds orthogonal directions of maximum variance and projects onto them. Every operation it does is linear — a rotation followed by dropping axes. It cannot bend space.
\[ Z = X W, \qquad W \in \mathbb{R}^{D \times 2} \;\; (\text{a fixed linear map}) \]
If the structure you care about is curved, a flat projection smears it. That is the wall this lesson climbs over.
Counterexample
Discussion prompt
PCA (Lesson 19) finds orthogonal directions of maximum variance and projects onto them. Every operation it does is linear — a rotation followed by dropping axes. It cannot bend space.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
If the structure you care about is curved, a flat projection smears it. That is the wall this lesson climbs over.
Picture it
Figure (svg): A spiral swiss-roll cross-section; two points on adjacent coils are near in straight-line distance but far along the coil.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture a sheet of paper rolled into a spiral. Two points on different layers of the roll are close in straight-line 3-D distance but far apart along the sheet — the distance that actually matters.
Intuition
Picture a sheet of paper rolled into a spiral. Two points on different layers of the roll are close in straight-line 3-D distance but far apart along the sheet — the distance that actually matters.
Figure (svg): A spiral swiss-roll cross-section; two points on adjacent coils are near in straight-line distance but far along the coil.
PCA keeps the misleading straight-line arrow. We want a method that keeps who your neighbors are on the sheet.
Analogy
Discussion prompt
Explain The swiss roll: distance lies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
PCA keeps the misleading straight-line arrow. We want a method that keeps who your neighbors are on the sheet.
Concept
manifold hypothesis — Real high-dimensional data (images, text, audio) does not fill its ambient space — it concentrates on a much lower-dimensional, usually curved, surface (a manifold) sitting inside it.
A 28×28 image lives in ℝ^{784}, but the set of digit images is a thin, curved sheet inside that cube. Neighborhood-preserving methods are built to follow that sheet where a linear projection can't.
Intuition
Forget coordinates. Ask a softer question: for each point, which other points are its close neighbors, and how close? Turn that into a probability — 'if I'm at point i, how likely am I to pick j as a neighbor?'
Do this in the original space and again in a 2-D map. Then slide the 2-D points around until the two neighbor-probability tables match. That is the whole algorithm — the rest is making 'match' precise.
Intuition
Why convert distances into probabilities at all? Because raw distances are on an arbitrary scale that differs between the 64-D space and the 2-D map — you can't compare ‖xᵢ−xⱼ‖ in one space to ‖yᵢ−yⱼ‖ in the other.
Normalizing each into a probability distribution puts both spaces on the same footing — every row sums to 1 — so 'make them match' becomes a well-defined objective (a divergence between distributions) instead of an apples-to-oranges distance comparison.
Section
Part 2 of 8 — one running example
Concept
So every number stays checkable by hand, our 'high-D' space is a line with four points in two tight pairs: {0, 1} sit together, and {2, 3} sit together far away.
| point | coordinate x | coordinate y |
|---|---|---|
| 0 | 0 | 0 |
| 1 | 1 | 0 |
| 2 | 5 | 0 |
| 3 | 6 | 0 |
The right answer for any neighbor method: keep 0–1 together and 2–3 together, with the two pairs apart. We'll watch the math discover exactly that.
Comparison
Comparison matrix
From Our running example: four points: refill the coordinate x column from what you know. The rest of the table is as it appeared.
| point | coordinate x | coordinate y |
|---|---|---|
| 0 | 0 | 0 |
| 1 | 1 | 0 |
| 2 | 5 | 0 |
| 3 | 6 | 0 |
Estimation
Predict first
t-SNE starts from pairwise squared Euclidean distances ‖xᵢ − xⱼ‖². With points on a line these are just squared coordinate gaps. Compute the whole matrix in NumPy:
Commit before you compute: what does Step 1 — squared distances in high-D come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Row 0: gaps to (1,5,6) are (1,5,6) → squares (1,25,36)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. d(0,1)²=1²=1, d(0,2)²=5²=25, d(0,3)²=6²=36.
Worked example
t-SNE starts from pairwise squared Euclidean distances ‖xᵢ − xⱼ‖². With points on a line these are just squared coordinate gaps. Compute the whole matrix in NumPy:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
# broadcast: D2[i,j] = ||xi - xj||^2
D2 = np.sum((X[:,None,:] - X[None,:,:])**2, axis=2)
print(D2)Row 0: gaps to (1,5,6) are (1,5,6) → squares (1,25,36)
Why: d(0,1)²=1²=1, d(0,2)²=5²=25, d(0,3)²=6²=36. The nearest neighbor of 0 is 1 by a wide margin — exactly the structure we want P to capture.
| from \ to | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 0 | 0 | 1 | 25 | 36 |
| 1 | 1 | 0 | 16 | 25 |
| 2 | 25 | 16 | 0 | 1 |
| 3 | 36 | 25 | 1 | 0 |
Trade off
Comparison matrix
From Step 1 — squared distances in high-D: every row here is a choice with a cost. Fill the 2 column, then say which row you would actually pick and what you give up for it.
| from \ to | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 0 | 0 | 1 | 25 | 36 |
| 1 | 1 | 0 | 16 | 25 |
| 2 | 25 | 16 | 0 | 1 |
| 3 | 36 | 25 | 1 | 0 |
Concept
Turn each distance into a similarity with a Gaussian bump centered on point i: near points score high, far points score near zero. Then normalize row i so it's a probability distribution over the other points.
\[ p_{j\mid i} = \frac{\exp\!\big(-\lVert x_i - x_j\rVert^2 / 2\sigma_i^2\big)}{\sum_{k \neq i} \exp\!\big(-\lVert x_i - x_k\rVert^2 / 2\sigma_i^2\big)}, \qquad p_{i\mid i} = 0 \]
p_{j|i} reads: standing at i, the probability I'd pick j as my neighbor. The bandwidth σᵢ sets how far 'neighbor' reaches — we'll fix σᵢ² = 2 for now so the arithmetic is clean, and tune it properly in Part 3.
Missing information
Discussion prompt
With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
= 0.778801, 0.001930, 0.000123. The denominator Z₀ = their sum = 0.780855, so p(1|0)=0.778801/0.780855=0.99737 — point 1 grabs essentially all of point 0's neighbor mass.
Worked example
With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2 / 4.0) # 2*sigma^2 = 4
np.fill_diagonal(W, 0.0) # p_{i|i} = 0
Pc = W / W.sum(axis=1, keepdims=True) # normalize each ROW
print(Pc[0])Row 0 numerators: exp(−1/4), exp(−25/4), exp(−36/4)
Why: = 0.778801, 0.001930, 0.000123. The denominator Z₀ = their sum = 0.780855, so p(1|0)=0.778801/0.780855=0.99737 — point 1 grabs essentially all of point 0's neighbor mass.
| p_{j|0} | j=0 | j=1 | j=2 | j=3 |
|---|---|---|---|---|
| value | 0 | 0.99737 | 0.00247 | 0.00016 |
| reading | self=0 | the neighbor | ≈0 | ≈0 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Row 0 numerators: exp(−1/4), exp(−25/4), exp(−36/4)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
With σᵢ² = 2, the exponent denominator is 2σᵢ² = 4. Exponentiate −D2/4, zero the diagonal, then row-normalize. Self-contained:
Anomaly
Predict first
A student writes this, and it looks reasonable:
Similarity is mutual, so surely p_{j|i} = p_{i|j} — one number per pair.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: They differ, because each row uses its OWN denominator Zᵢ.
Conditionals are row-normalized, so p_{j|i} and p_{i|j} generally differ. We fix this by symmetrizing in the next step.
Why: They differ, because each row uses its OWN denominator Zᵢ. Point 0's row sums to Z₀=0.78085 while point 2's row sums to Z₂=0.79905 — different totals. The unnormalized weight exp(−25/4) is shared, so p(2|0)=exp(−25/4)/0.78085=0.002472 but p(0|2)=exp(−25/4)/0.79905=0.002416. Same numerator, different Z ⇒ p(2|0)≠p(0|2).
Trap
Similarity is mutual, so surely p_{j|i} = p_{i|j} — one number per pair.
Assume p(2|0) = p(0|2)
Why: They differ, because each row uses its OWN denominator Zᵢ. Point 0's row sums to Z₀=0.78085 while point 2's row sums to Z₂=0.79905 — different totals. The unnormalized weight exp(−25/4) is shared, so p(2|0)=exp(−25/4)/0.78085=0.002472 but p(0|2)=exp(−25/4)/0.79905=0.002416. Same numerator, different Z ⇒ p(2|0)≠p(0|2).
Conditionals are row-normalized, so p_{j|i} and p_{i|j} generally differ. We fix this by symmetrizing in the next step.
Define the JOINT P_ij = (p_{j|i} + p_{i|j}) / (2n)
Why: Averaging the two directions and dividing by n makes P symmetric AND sum to 1 over all pairs — the object t-SNE actually matches. Never feed raw conditionals to the objective.
Pattern
Predict first
The table runs: 0 | 0 | 0.2465 | 0.0006 | 0.0000 · 1 | 0.2465 | 0 | 0.0057 | 0.0006 · 2 | 0.0006 | 0.0057 | 0 | 0.2465
In Step 3 — symmetrize into the joint P, given the rows so far: what is the next one — the row where P_ij is 3?
Correct: 3 | 0.0000 | 0.0006 | 0.2465 | 0
| P_ij | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 0 | 0 | 0.2465 | 0.0006 | 0.0000 |
| 1 | 0.2465 | 0 | 0.0057 | 0.0006 |
| 2 | 0.0006 | 0.0057 | 0 | 0.2465 |
| 3 | 0.0000 | 0.0006 | 0.2465 | 0 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two within-pair entries P[0,1] and P[2,3] each carry ≈0.2465 of the total mass; cross-pair entries like P[1,2] are ≈0.0057.
Worked example
Average the two directions and divide by 2n (here n = 4) so the whole matrix is symmetric and sums to 1:
\[ P_{ij} = \frac{p_{j\mid i} + p_{i\mid j}}{2n}, \qquad \sum_{i \neq j} P_{ij} = 1 \]
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n = len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
P = (Pc + Pc.T) / (2*n) # symmetric joint
print(np.round(P, 6)); print('sum =', P.sum())P[0,1] = (0.99737 + 0.974662)/8 = 0.246504
Why: The two within-pair entries P[0,1] and P[2,3] each carry ≈0.2465 of the total mass; cross-pair entries like P[1,2] are ≈0.0057. P encodes 'two tight pairs' numerically.
| P_ij | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 0 | 0 | 0.2465 | 0.0006 | 0.0000 |
| 1 | 0.2465 | 0 | 0.0057 | 0.0006 |
| 2 | 0.0006 | 0.0057 | 0 | 0.2465 |
| 3 | 0.0000 | 0.0006 | 0.2465 | 0 |
Pattern
Step through it
Step through Step 3 — symmetrize into the joint P one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 3 of 8 — the per-point bandwidth
Concept
A dense cluster needs a small σ (neighbors are close); a sparse region needs a large σ (neighbors are far). One global bandwidth would over-smooth the dense parts and starve the sparse parts.
So t-SNE picks a separate σᵢ for each point — chosen so that every point ends up with roughly the same effective number of neighbors. That target number is the perplexity.
Concept
The Shannon entropy H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:
\[ \mathrm{Perp}(P_i) = 2^{H(P_i)}, \qquad H(P_i) = -\!\sum_{j} p_{j\mid i}\,\log_2 p_{j\mid i} \]
If pᵢ spreads evenly over k neighbors, Perp = k. So perplexity is a smooth, differentiable stand-in for 'how many neighbors each point effectively has' — typically set between 5 and 50.
Explain it
Discussion prompt
Explain Perplexity = 2 to the entropy to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The Shannon entropy H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:
Fill the middle
Fill in the blanks
From Compute the entropy of a row — one line has had its right-hand side removed. Put it back.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
p0 = Pc[0]; p0 = p0[p0 > 0]**
H = -np.sum(p0 * np.log2(p0))
print(round(H, 6), round(2**H, 6))
Why: p0 is what everything below it consumes, so the wrong expression here fails later and somewhere else. Almost all mass sits on one neighbor, so entropy is near zero and perplexity near 1 — 'point 0 effectively has one neighbor'.
Worked example
Perplexity is 2^{H}, so first compute the entropy H(P₀) of point 0's conditional row [0.99737, 0.00247, 0.00016] (excluding self). A near-one-hot distribution has tiny entropy. Self-contained:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W, 0.0)
Pc = W / W.sum(axis=1, keepdims=True)
p0 = Pc[0]; p0 = p0[p0 > 0]
H = -np.sum(p0 * np.log2(p0))
print(round(H, 6), round(2**H, 6))H(P₀) = 0.027195 bits → perplexity 2^H = 1.019
Why: Almost all mass sits on one neighbor, so entropy is near zero and perplexity near 1 — 'point 0 effectively has one neighbor'. This matches the σ²=2 row of the sweep table exactly.
| quantity | value |
|---|---|
| dominant p(1|0) | 0.99737 |
| H(P₀) = −Σ p log₂ p | 0.027195 bits |
| perplexity = 2^H | 1.019029 |
Sorting
Sort into buckets
These are the pieces of Lesson 31: t-SNE & UMAP, out of order. Put each one back under the part of the lesson it belongs to.
Faded example
Fill in the blanks
Perplexity rises with σ — traced, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])2, axis=2)
row = D2[0]** # distances from point 0
def perplexity(sig2):
w = np.exp(-row/(2*sig2)); w[row==0] = 0.0
p = w/w.sum(); p = p[p>0]
return 2*(-np.sum(pnp.log2(p)))
for s2 in [0.5, 2.0, 8.0, 50.0]:
print(s2, round(perplexity(s2), 4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Perplexity climbs monotonically with σ².
Worked example
Fix point 0's distances and sweep σ². Small σ → all mass on the one nearest point (perplexity ≈ 1); large σ → mass spreads to all three others (perplexity → 3). Self-contained:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
row = D2[0] # distances from point 0
def perplexity(sig2):
w = np.exp(-row/(2*sig2)); w[row==0] = 0.0
p = w/w.sum(); p = p[p>0]
return 2**(-np.sum(p*np.log2(p)))
for s2 in [0.5, 2.0, 8.0, 50.0]:
print(s2, round(perplexity(s2), 4))σ²: 0.5 → 1.000, 2 → 1.019, 8 → 2.062, 50 → 2.967
Why: Perplexity climbs monotonically with σ². At σ²=0.5 point 0 has one effective neighbor; at σ²=50 it has ~3 (all other points). This monotonicity is what lets a binary search hit any target.
| σ² | perplexity(row 0) | effective neighbors |
|---|---|---|
| 0.5 | 1.0000 | just point 1 |
| 2.0 | 1.0190 | ≈ point 1 |
| 8.0 | 2.0619 | ≈ 2 points |
| 50.0 | 2.9671 | ≈ all 3 others |
Scale up
Step through it
Step through Perplexity rises with σ — traced and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Faded example
Fill in the blanks
Binary-search σ to hit the target, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
row = np.sum((X[0]-X)2, axis=1)**
def perp(sig2):
w = np.exp(-row/(2*sig2)); w[row==0]=0.0
p = w/w.sum(); p = p[p>0]
return 2*(-np.sum(pnp.log2(p)))
lo, hi = 1e-6, 1e6
for _ in range(60):
mid = (lo+hi)/2
if perp(mid) < 2.0: lo = mid
else: hi = mid
print(round(mid,4), round(perp(mid),4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The search brackets the answer and halves the interval 60 times.
Worked example
Because perplexity increases monotonically in σ, t-SNE binary-searches σᵢ until Perp(Pᵢ) equals the requested value. Target perplexity 2.0 for point 0:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
row = np.sum((X[0]-X)**2, axis=1)
def perp(sig2):
w = np.exp(-row/(2*sig2)); w[row==0]=0.0
p = w/w.sum(); p = p[p>0]
return 2**(-np.sum(p*np.log2(p)))
lo, hi = 1e-6, 1e6
for _ in range(60):
mid = (lo+hi)/2
if perp(mid) < 2.0: lo = mid
else: hi = mid
print(round(mid,4), round(perp(mid),4))Converges to σ² = 7.6062, achieving perplexity 2.0000
Why: The search brackets the answer and halves the interval 60 times. Real t-SNE runs exactly this search for EVERY point, giving each its own σᵢ before building P.
| target perplexity | found σ² | achieved perplexity |
|---|---|---|
| 1.0 (min) | → 0 | → 1 |
| 2.0 | 7.6062 | 2.0000 |
| 3.0 (max, n−1) | → ∞ | → 3 |
Pattern
Step through it
Step through Binary-search σ to hit the target one row at a time. What is driving the change, and what would the row after the last one be?
Section
Part 4 of 8 — the map side
Concept
Now the 2-D map. Each high-D point xᵢ gets a low-D position yᵢ (what we solve for). Build a similarity Q on the y's the same spirit as P — but with a different kernel and a single global normalization.
\[ q_{ij} = \frac{\big(1 + \lVert y_i - y_j\rVert^2\big)^{-1}}{\sum_{k \neq l}\big(1 + \lVert y_k - y_l\rVert^2\big)^{-1}} \]
That (1 + d²)⁻¹ is a Student-t distribution with one degree of freedom (a Cauchy). Why not a Gaussian like in high-D? That choice is the single most important idea in t-SNE — next.
Intuition
In high dimensions there is room: a point can have many neighbors all roughly equidistant. Squash that into 2-D and there isn't enough space — moderately-distant points get shoved on top of each other. The map turns into a single blob.
The original SNE used a Gaussian in 2-D and suffered exactly this crowding. t-SNE's fix: give the low-D kernel a heavy tail so far-apart points can stay far apart without paying a huge similarity penalty.
Concept
The original SNE (Hinton & Roweis, 2002) already matched Gaussian similarities by minimizing KL, but it was hard to optimize and crowded. t-SNE (2008) makes exactly two changes.
P (not asymmetric conditionals per point) — a simpler gradientEverything else — the per-point σ, the perplexity target, gradient descent on KL — is inherited from SNE. The 't' in t-SNE is literally the t-distribution.
Fill the middle
Fill in the blanks
From Gaussian vs Student-t tails — traced — one line has had its right-hand side removed. Put it back.
import numpy as np
for d in [1., 2., 4., 8.]:
gaussian = np.exp(-d2 / 2.0) # SNE's 2-D kernel
student_t = 1.0 / (1.0 + d2) # t-SNE's 2-D kernel
print(int(d), round(gaussian,6), round(student_t,6))
Why: student_t is what everything below it consumes, so the wrong expression here fails later and somewhere else. The t-kernel assigns 175× more similarity to a distance-4 pair.
Worked example
Compare the two kernels at growing distance. The Gaussian collapses to ~0; the Student-t decays gently. Self-contained:
import numpy as np
for d in [1., 2., 4., 8.]:
gaussian = np.exp(-d**2 / 2.0) # SNE's 2-D kernel
student_t = 1.0 / (1.0 + d**2) # t-SNE's 2-D kernel
print(int(d), round(gaussian,6), round(student_t,6))At d=4: Gaussian = 0.000335, Student-t = 0.058824
Why: The t-kernel assigns 175× more similarity to a distance-4 pair. Under the Gaussian, a pair that far apart contributes almost nothing — so the optimizer yanks it inward to matter, causing crowding. The heavy tail removes that pull.
| distance d | Gaussian e^{−d²/2} | Student-t 1/(1+d²) | ratio t/g |
|---|---|---|---|
| 1 | 0.606531 | 0.500000 | 0.82× |
| 2 | 0.135335 | 0.200000 | 1.48× |
| 4 | 0.000335 | 0.058824 | 175× |
| 8 | 0.000000 | 0.015385 | huge |
Pattern
Step through it
Step through Gaussian vs Student-t tails — traced one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Flip it around: to produce a target similarity q = 0.1, how far apart may a pair sit under each kernel? Solve e^{−d²/2}=0.1 and 1/(1+d²)=0.1:
Commit before you compute: what does Same similarity, more room come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Gaussian d = 2.146, Student-t d = 3.000 (1.40× farther)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For the SAME similarity, the t-kernel lets a pair sit 40% farther apart.
Worked example
Flip it around: to produce a target similarity q = 0.1, how far apart may a pair sit under each kernel? Solve e^{−d²/2}=0.1 and 1/(1+d²)=0.1:
import numpy as np
q = 0.1
d_gauss = np.sqrt(-2*np.log(q)) # Gaussian: e^{-d^2/2}=q
d_t = np.sqrt(1/q - 1) # Student-t: 1/(1+d^2)=q
print(round(d_gauss,4), round(d_t,4), round(d_t/d_gauss,2))Gaussian d = 2.146, Student-t d = 3.000 (1.40× farther)
Why: For the SAME similarity, the t-kernel lets a pair sit 40% farther apart. Multiplied across every moderate pair, that extra room is what spreads the clusters out instead of crushing them together.
| kernel | distance to hit q=0.1 | effect on map |
|---|---|---|
| Gaussian (SNE) | 2.146 | clusters crowd inward |
| Student-t (t-SNE) | 3.000 | clusters breathe |
| gain | 1.40× | uncrowded |
Worked example
Place the four points on a line as a candidate map Y = [0, 0.5, 3, 3.5] — same pairing as high-D. Build Q with the t-kernel and a single global normalizer. Self-contained:
import numpy as np
Y = np.array([[0.],[0.5],[3.],[3.5]])
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0 + Dy2); np.fill_diagonal(Kt, 0.0)
Q = Kt / Kt.sum() # ONE global denominator
print(np.round(Q, 6)); print('sum =', Q.sum())q(0,1) = 0.8 / 4.026805 = 0.198669
Why: The kernel value for the close pair (d²=0.25) is 1/1.25=0.8; divided by the total 4.0268 gives 0.1987. Q is symmetric and sums to 1, just like P — now they are comparable.
| Q_ij | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| 0 | 0 | 0.1987 | 0.0248 | 0.0187 |
| 1 | 0.1987 | 0 | 0.0343 | 0.0248 |
| 2 | 0.0248 | 0.0343 | 0 | 0.1987 |
| 3 | 0.0187 | 0.0248 | 0.1987 | 0 |
Scale up
Step through it
Step through Compute Q for a candidate map and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Section
Part 5 of 8 — matching the two tables
Concept
We need a number that says how far apart the two neighbor tables P and Q are. That's the KL divergence (Lesson 14): the cost of using Q to describe data drawn from P.
\[ C = D_{KL}(P \,\|\, Q) = \sum_{i \neq j} P_{ij} \, \log \frac{P_{ij}}{Q_{ij}} \]
It is ≥ 0, and = 0 iff P = Q everywhere. t-SNE moves the y's to drive C down — that's the entire training loop.
Intuition
Look at a single term P_{ij} log(P_{ij}/Q_{ij}). It's weighted by P_{ij}. So a pair that is close in high-D (large P) is expensive to get wrong — the map is punished hard for pulling true neighbors apart.
But a pair that is far in high-D (P ≈ 0) contributes almost nothing, even if the map accidentally places them close. Net effect: t-SNE fiercely preserves local neighborhoods and is relaxed about global distances. Remember that asymmetry — it explains every t-SNE 'trap' later.
Worked example
Rebuild both P and Q from scratch, then sum P·log(P/Q) over all ordered pairs. Fully self-contained:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1, keepdims=True); P = (Pc+Pc.T)/(2*n)
Y = np.array([[0.],[0.5],[3.],[3.5]])
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
m = ~np.eye(n, dtype=bool)
print('KL =', round(np.sum(P[m]*np.log(P[m]/Q[m])), 6))The four within-pair terms dominate at +0.0532 each
Why: P(0,1)=0.2465 > Q(0,1)=0.1987, so log(P/Q)>0 and the term is +0.0532; the four such ordered within-pair terms sum to +0.213. The cross-pair terms are slightly negative, netting KL = 0.182689.
| pair (i,j) | P_ij | Q_ij | P·log(P/Q) |
|---|---|---|---|
| (0,1) within | 0.2465 | 0.1987 | +0.053181 |
| (2,3) within | 0.2465 | 0.1987 | +0.053181 |
| (1,2) cross | 0.0057 | 0.0343 | −0.010246 |
| (0,2) cross | 0.0006 | 0.0248 | −0.002264 |
| Σ (all ordered pairs) | — | — | 0.182689 |
Fill the middle
Fill in the blanks
From A worse map has higher KL — one line has had its right-hand side removed. Put it back.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1,keepdims=True); P = (Pc+Pc.T)/(2*n)
def kl_of(Y):
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
m = ~np.eye(n,dtype=bool)
return np.sum(P[m]*np.log(P[m]/Q[m]))
good = np.array([[0.],[0.5],[3.],[3.5]])
bad = np.array([[0.],[0.5],[0.2],[0.7]])
print(round(kl_of(good),4), round(kl_of(bad),4))
Why: good is what everything below it consumes, so the wrong expression here fails later and somewhere else. Overlapping the clusters puts high Q on pairs that have near-zero P — the KL almost sextuples.
Worked example
Prove the objective actually prefers the right layout. Take a bad map that overlaps the two clusters (Y = [0, 0.5, 0.2, 0.7]) and recompute KL against the same P:
import numpy as np
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]]); n=len(X)
D2 = np.sum((X[:,None,:]-X[None,:,:])**2, axis=2)
W = np.exp(-D2/4.0); np.fill_diagonal(W,0.0)
Pc = W/W.sum(axis=1,keepdims=True); P = (Pc+Pc.T)/(2*n)
def kl_of(Y):
Dy2 = np.sum((Y[:,None,:]-Y[None,:,:])**2, axis=2)
Kt = 1.0/(1.0+Dy2); np.fill_diagonal(Kt,0.0); Q = Kt/Kt.sum()
m = ~np.eye(n,dtype=bool)
return np.sum(P[m]*np.log(P[m]/Q[m]))
good = np.array([[0.],[0.5],[3.],[3.5]])
bad = np.array([[0.],[0.5],[0.2],[0.7]])
print(round(kl_of(good),4), round(kl_of(bad),4))good KL = 0.1827, bad KL = 1.0870
Why: Overlapping the clusters puts high Q on pairs that have near-zero P — the KL almost sextuples. Gradient descent on KL therefore pushes the layout AWAY from the bad map toward the separated one. The objective works.
| candidate map Y | clusters | KL(P‖Q) |
|---|---|---|
| [0, 0.5, 3, 3.5] | separated (correct) | 0.1827 |
| [0, 0.5, 0.2, 0.7] | overlapping (wrong) | 1.0870 |
| → optimizer prefers | separated | lower KL |
Comparison
Comparison matrix
From A worse map has higher KL: refill the KL(P‖Q) column from what you know. The rest of the table is as it appeared.
| candidate map Y | clusters | KL(P‖Q) |
|---|---|---|
| [0, 0.5, 3, 3.5] | separated (correct) | 0.1827 |
| [0, 0.5, 0.2, 0.7] | overlapping (wrong) | 1.0870 |
| → optimizer prefers | separated | lower KL |
Concept
t-SNE minimizes C by gradient descent on the map positions yᵢ. The gradient has a clean, force-like form:
\[ \frac{\partial C}{\partial y_i} = 4 \sum_{j} (P_{ij} - Q_{ij})\,\big(1 + \lVert y_i - y_j\rVert^2\big)^{-1}\,(y_i - y_j) \]
Read it as springs: when P_{ij} > Q_{ij} (truer neighbors than the map shows) the term pulls yᵢ toward yⱼ; when P_{ij} < Q_{ij} it pushes them apart. The map settles where every spring balances.
Concept
Raw gradient descent on KL gets stuck in tangled local minima. sklearn's TSNE uses two tricks by default so the map untangles cleanly.
| trick | default | what it does |
|---|---|---|
| early_exaggeration | 12.0 | scales P up early → clusters clump tight, then relax |
| learning_rate | auto | step size for the position updates |
| max_iter | 1000 | gradient-descent steps before stopping |
Early exaggeration multiplies P by 12 for the first phase, over-emphasizing true neighbors so clusters form; then it's turned off to fine-tune spacing. You rarely touch these — but knowing they exist explains why a t-SNE run has phases.
Ranking
Put in order
These are the steps of The t-SNE algorithm, end to end, scrambled. Put them back in order before the next slide shows you.
‖xᵢ − xⱼ‖² in high-Dσᵢ so Perp(Pᵢ) = target perplexityp_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)(1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizermin D_KL(P‖Q) on the positions yᵢWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
‖xᵢ − xⱼ‖² in high-Dσᵢ so Perp(Pᵢ) = target perplexityp_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)(1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizermin D_KL(P‖Q) on the positions yᵢEdge cases
Discussion prompt
The t-SNE algorithm, end to end works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
‖xᵢ − xⱼ‖² in high-Dσᵢ so Perp(Pᵢ) = target perplexityp_{j|i} → symmetrize → joint Pᵢⱼ (sums to 1)(1+‖yᵢ−yⱼ‖²)⁻¹, one global normalizermin D_KL(P‖Q) on the positions yᵢSection
Part 6 of 8 — the graph view
Concept
UMAP reaches the same goal from graph theory. Build a weighted k-nearest-neighbor graph in high-D (edge weights = fuzzy membership strengths), then lay that graph out in 2-D so its edge structure is preserved.
Its low-D objective is a cross-entropy over edges (both present and absent), optimized with negative sampling — closer in spirit to how word embeddings train than to t-SNE's KL, but the neighbor-preserving intent is the same.
Worked example
UMAP's first move is a k-nearest-neighbor graph. Build it on our running example with k=1 (each point's single nearest non-self neighbor) and read off the edges. Self-contained:
import numpy as np
from sklearn.neighbors import NearestNeighbors
X = np.array([[0.,0.],[1.,0.],[5.,0.],[6.,0.]])
nn = NearestNeighbors(n_neighbors=2).fit(X) # 2 = self + 1
dist, idx = nn.kneighbors(X)
for i in range(4):
print(i, '->', idx[i][1], 'dist', round(dist[i][1], 3))Edges: 0↔1 and 2↔3, each at distance 1
Why: The graph recovers the two tight pairs exactly — same neighborhood structure P encoded, now as explicit edges. UMAP then lays out THIS graph in 2-D, minimizing an edge cross-entropy rather than a similarity KL.
| point | nearest neighbor | edge distance |
|---|---|---|
| 0 | 1 | 1.000 |
| 1 | 0 | 1.000 |
| 2 | 3 | 1.000 |
| 3 | 2 | 1.000 |
Invariant
Step through it
Step through The neighbor graph on our four points one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
| axis | t-SNE | UMAP |
|---|---|---|
| knob | perplexity (5–50) | n_neighbors + min_dist |
| objective | KL of similarities | graph cross-entropy |
| global structure | weak | somewhat better |
| speed / scale | slower (O(N log N) w/ Barnes-Hut) | faster, scales to millions |
| new points | no transform | approximate transform, still shaky |
UMAP preserves a bit more global layout and runs faster, so it's the default for large datasets. But it shares t-SNE's core caveats — the map is still for eyes, not features.
Explain it
Discussion prompt
Explain UMAP vs t-SNE, honestly to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
UMAP preserves a bit more global layout and runs faster, so it's the default for large datasets. But it shares t-SNE's core caveats — the map is still for eyes, not features.
Concept
t-SNE's perplexity and UMAP's n_neighbors both set the local-vs-global balance. Small values emphasize tight local clusters; large values reveal coarser global shape.
Extreme settings produce misleading maps — a too-small perplexity can shatter one real cluster into fake shards, a too-large one can merge distinct clusters. Always sweep a few values before trusting a picture.
Concept
| method | reach for it when you need |
|---|---|
| PCA | fast linear reduction, reusable features for ML, denoising |
| t-SNE | high-quality 2-D pictures of local cluster structure |
| UMAP | faster viz, a bit more global structure, larger datasets |
One rule of thumb: PCA is the workhorse for features; t-SNE and UMAP are for the human eye. The next slides prove — in code — why you must not confuse the two.
Analogy
Discussion prompt
Explain Which method when by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One rule of thumb: PCA is the workhorse for features; t-SNE and UMAP are for the human eye. The next slides prove — in code — why you must not confuse the two.
Section
Part 7 of 8 — the traps in code
Pattern
Predict first
The table runs: raw pixels | 64 | 0.963 · PCA (linear) | 2 | 0.609
In PCA vs t-SNE on digits, given the rows so far: what is the next one — the row where representation is t-SNE (nonlinear)?
Correct: t-SNE (nonlinear) | 2 | 0.980
| representation | dims | KNN accuracy |
|---|---|---|
| raw pixels | 64 | 0.963 |
| PCA (linear) | 2 | 0.609 |
| t-SNE (nonlinear) | 2 | 0.980 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. t-SNE's neighborhood preservation packs each digit into a tight, separable island — its 2-D map is even MORE KNN-separable than the raw 64-D pixels (0.98 vs 0.96), while PCA's flat projection overlaps several digits down to 0.61.
Worked example
The payoff, on real data. Project the 64-D digits set to 2-D two ways, then measure how separable the ten classes are with a 3-fold KNN classifier on each 2-D embedding. Self-contained:
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
k = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('raw64', k(X), 'pca2', k(pca), 'tsne2', k(tsne))raw 64-D = 0.963, PCA 2-D = 0.609, t-SNE 2-D = 0.980
Why: t-SNE's neighborhood preservation packs each digit into a tight, separable island — its 2-D map is even MORE KNN-separable than the raw 64-D pixels (0.98 vs 0.96), while PCA's flat projection overlaps several digits down to 0.61.
| representation | dims | KNN accuracy |
|---|---|---|
| raw pixels | 64 | 0.963 |
| PCA (linear) | 2 | 0.609 |
| t-SNE (nonlinear) | 2 | 0.980 |
Missing information
Discussion prompt
That 0.98 is seductive — why not use the t-SNE coordinates as features? Because there is no mapping for new points. PCA learns a reusable projection; t-SNE learns only where these points go. Check the API itself:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
PCA stores components you can apply to unseen rows. TSNE only implements fit_transform on the training set — there is literally no method to embed a new point, because its position is defined only relative to the points it was optimized with.
Worked example
That 0.98 is seductive — why not use the t-SNE coordinates as features? Because there is no mapping for new points. PCA learns a reusable projection; t-SNE learns only where these points go. Check the API itself:
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
X, y = load_digits(return_X_y=True)
print('PCA has transform:', hasattr(PCA(2).fit(X), 'transform'))
print('TSNE has transform:', hasattr(TSNE(2), 'transform'))PCA.transform exists (True); TSNE.transform does not (False)
Why: PCA stores components you can apply to unseen rows. TSNE only implements fit_transform on the training set — there is literally no method to embed a new point, because its position is defined only relative to the points it was optimized with.
| estimator | has .transform? | usable on new data? |
|---|---|---|
| PCA | True | yes — reusable features |
| TSNE | False | no — training set only |
| UMAP | approximate | shaky, not for production |
Estimation
Predict first
Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:
Commit before you compute: what does Proof: the layout is non-deterministic come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: max coordinate difference between seeds ≈ 130.6
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The same data, same settings, different seed → a completely different coordinate system (the local neighborhoods are similar, but absolute positions swing by ~130 units).
Worked example
Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:
import numpy as np
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
X, y = load_digits(return_X_y=True)
a = TSNE(2, init='random', perplexity=30, random_state=0).fit_transform(X)
b = TSNE(2, init='random', perplexity=30, random_state=1).fit_transform(X)
print('max coordinate diff:', round(np.abs(a-b).max(), 1))max coordinate difference between seeds ≈ 130.6
Why: The same data, same settings, different seed → a completely different coordinate system (the local neighborhoods are similar, but absolute positions swing by ~130 units). Features that change this much between runs cannot go into a model.
| run | seed | outcome |
|---|---|---|
| A | 0 | one arbitrary orientation |
| B | 1 | different orientation/positions |
| max |A − B| | — | ≈ 130.6 (huge) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
max coordinate difference between seeds ≈ 130.6
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Even the shape isn't stable: two runs with different seeds land points in wildly different places. Run t-SNE twice with init='random' and compare coordinates:
Anomaly
Predict first
A student writes this, and it looks reasonable:
t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless.
Use t-SNE for visualization and exploration only. Model on PCA, raw, or learned features that have a stable transform.
Why: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.
Trap
t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.
model.fit(tsne_coords, y) → ship it
Why: It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.
Use t-SNE for visualization and exploration only. Model on PCA, raw, or learned features that have a stable transform.
Visualize with t-SNE; model on PCA / raw / embeddings
Why: PCA's .transform embeds new points deterministically, so its coordinates generalize. t-SNE is a microscope for looking at your data, not a feature extractor for a pipeline.
Break the constraint
Discussion prompt
The rule this trap just fixed:
PCA's .transform embeds new points deterministically, so its coordinates generalize. t-SNE is a microscope for looking at your data, not a feature extractor for a pipeline.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
It has no .transform, so you literally cannot embed the next incoming digit; the layout changes every run (Δ≈130); and inter-cluster distances are meaningless. The 0.98 is an IN-SAMPLE separability diagnostic — it does not transfer to new data.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Cluster B is twice the diameter of cluster A and sits far from it, so B's items are more varied and the two groups are very different.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary.
Read only local neighborhood structure. Treat cluster sizes and global distances as unreliable.
Why: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary. A 'bigger' or 'farther' cluster in the map may be neither in the data.
Trap
Cluster B is twice the diameter of cluster A and sits far from it, so B's items are more varied and the two groups are very different.
Interpret cluster sizes and between-cluster gaps as real
Why: The KL objective is P-weighted, so it only pins down LOCAL neighbors; it equalizes dense regions and the gaps between clusters are largely arbitrary. A 'bigger' or 'farther' cluster in the map may be neither in the data.
Read only local neighborhood structure. Treat cluster sizes and global distances as unreliable.
Say 'these points are neighbors', not 'this cluster is bigger/farther'
Why: The method optimizes local similarity by construction. UMAP preserves a little more global layout, but the same caution holds — never quantify from the picture.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
H(Pᵢ) of the conditional row measures how spread-out point i's neighbor distribution is. Perplexity exponentiates it:p_{j|i} = p_{i|j} — one number per pair.; t-SNE separates the digits beautifully (0.98), so train the production classifier on the 2-D t-SNE features.Elimination
Eliminate the wrong options
t-SNE finds its 2-D layout by minimizing:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: t-SNE converts distances to similarity probabilities P (high-D, Gaussian) and Q (2-D, Student-t), then gradient-descends the positions to minimize D_KL(P‖Q) = Σ P log(P/Q) — preserving local neighborhoods.
Check
Back to first principles: what number is t-SNE driving down?
Check your understanding
t-SNE finds its 2-D layout by minimizing:
Answer: A
Why: t-SNE converts distances to similarity probabilities P (high-D, Gaussian) and Q (2-D, Student-t), then gradient-descends the positions to minimize D_KL(P‖Q) = Σ P log(P/Q) — preserving local neighborhoods.
Prediction
Predict first
Why does t-SNE use a heavy-tailed Student-t kernel in the 2-D map instead of a Gaussian?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Its heavy tail lets moderately-distant points stay far apart cheaply, relieving the crowding that plagued Gaussian SNE
Why: In 2-D there isn't room for all the high-D neighbors. The Gaussian's tail vanishes (0.0003 at d=4), so distant pairs contribute nothing and get pulled inward — crowding. The Student-t's 1/(1+d²) tail (0.059 at d=4) keeps distant pairs meaningful, so clusters spread out.
Check
Recall the tail comparison you traced.
Check your understanding
Why does t-SNE use a heavy-tailed Student-t kernel in the 2-D map instead of a Gaussian?
Answer: A
Why: In 2-D there isn't room for all the high-D neighbors. The Gaussian's tail vanishes (0.0003 at d=4), so distant pairs contribute nothing and get pulled inward — crowding. The Student-t's 1/(1+d²) tail (0.059 at d=4) keeps distant pairs meaningful, so clusters spread out.
Prediction
Predict first
Why are t-SNE coordinates a poor choice as features for a classifier that will see new data?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: There is no stable transform for new points, the layout is non-deterministic, and global geometry is distorted
Why: TSNE has no .transform (proved: hasattr → False), two seeds differ by ~130 units, and its inter-cluster distances are unreliable — so the coordinates don't generalize to unseen points.
Check
You proved two of these in code.
Check your understanding
Why are t-SNE coordinates a poor choice as features for a classifier that will see new data?
Answer: A
Why: TSNE has no .transform (proved: hasattr → False), two seeds differ by ~130 units, and its inter-cluster distances are unreliable — so the coordinates don't generalize to unseen points.
Elimination
Eliminate the wrong options
You need 2-D features to feed a model that will later score new, unseen data. Best choice?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: PCA fits a linear map and stores it, so .transform embeds unseen rows deterministically — the requirement for reusable features. t-SNE/UMAP are visualization tools without a reliable out-of-sample mapping.
Check
Match the method to the job.
Check your understanding
You need 2-D features to feed a model that will later score new, unseen data. Best choice?
Answer: A
Why: PCA fits a linear map and stores it, so .transform embeds unseen rows deterministically — the requirement for reusable features. t-SNE/UMAP are visualization tools without a reliable out-of-sample mapping.
Section
Part 8 of 8 — the project
Concept
Embed the digits set to 2-D with both PCA and t-SNE, score each with a KNN classifier, and draw the right conclusion about what each tool is for. You derived every piece — now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | PCA to 2-D + KNN score | PCA(n_components=2) |
| 2 | t-SNE to 2-D + KNN score | TSNE(perplexity=30) |
| 3 | Compare, and state the correct use of each | cross_val_score |
Build rules: type every line yourself, always set random_state for t-SNE (it's stochastic), and remember the KNN number is a separability diagnostic, not a deployable model.
Counterexample
Discussion prompt
Embed the digits set to 2-D with both PCA and t-SNE, score each with a KNN classifier, and draw the right conclusion about what each tool is for. You derived every piece — now assemble it.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, always set random_state for t-SNE (it's stochastic), and remember the KNN number is a separability diagnostic, not a deployable model.
Worked example
Your turn: project digits to 2-D with PCA and score a 3-fold KNN on it. Predict out loud: above or below 0.7?
Hint: PCA(n_components=2).fit_transform(X), then cross_val_score(KNeighborsClassifier(), pca, y, cv=3).mean().
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca = PCA(n_components=2).fit_transform(X)
print(round(cross_val_score(KNeighborsClassifier(), pca, y, cv=3).mean(), 3))| embedding | dims | KNN accuracy |
|---|---|---|
| PCA 2-D | 2 | 0.609 |
Worked example
Your turn: project to 2-D with t-SNE and score the same KNN. Predict the jump over PCA before you run it.
Hint: TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X). The random_state makes it reproducible.
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
print(round(cross_val_score(KNeighborsClassifier(), tsne, y, cv=3).mean(), 3))| embedding | dims | KNN accuracy |
|---|---|---|
| t-SNE 2-D | 2 | 0.980 |
Worked example
Your turn: put the two numbers (and the raw 64-D baseline) side by side. Which embedding preserves class structure in 2-D — and which would you actually deploy?
Hint: score X itself for the raw baseline. Expect t-SNE ≳ raw ≫ PCA on separability — yet PCA is the one you'd ship.
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
k = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('raw', k(X), 'pca', k(pca), 'tsne', k(tsne))| embedding | KNN accuracy | correct use |
|---|---|---|
| raw 64-D | 0.963 | baseline |
| PCA 2-D | 0.609 | reusable features (has transform) |
| t-SNE 2-D | 0.980 | visualization only (no transform) |
Trade off
Comparison matrix
From Milestone 3 — compare and conclude: every row here is a choice with a cost. Fill the correct use column, then say which row you would actually pick and what you give up for it.
| embedding | KNN accuracy | correct use |
|---|---|---|
| raw 64-D | 0.963 | baseline |
| PCA 2-D | 0.609 | reusable features (has transform) |
| t-SNE 2-D | 0.980 | visualization only (no transform) |
Concept
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import cross_val_score
X, y = load_digits(return_X_y=True)
pca = PCA(n_components=2).fit_transform(X)
tsne = TSNE(n_components=2, init='pca', perplexity=30, random_state=0).fit_transform(X)
knn = lambda Z: round(cross_val_score(KNeighborsClassifier(), Z, y, cv=3).mean(), 3)
print('PCA 2D:', knn(pca), ' t-SNE 2D:', knn(tsne))
print('PCA has transform:', hasattr(PCA(2).fit(X), 'transform'),
'| TSNE has transform:', hasattr(TSNE(2), 'transform'))| printed line | value |
|---|---|
| PCA 2D: | 0.609 |
| t-SNE 2D: | 0.98 |
| PCA has transform / TSNE has transform | True / False |
If t-SNE's 2-D map is far more separable than PCA's — yet you'd still model on PCA because only it has a transform — you know exactly what each tool is for.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| PCA 2D: | 0.609 |
| t-SNE 2D: | 0.98 |
| PCA has transform / TSNE has transform | True / False |
Concept
Slides closed, out loud: explain (1) which two distributions t-SNE matches and with which divergence, (2) why the Student-t tail prevents crowding, and (3) two concrete reasons you can't feed t-SNE coordinates to a downstream model.
Stretch (homework): scatter-plot the t-SNE map colored by digit, sweep perplexity to 5 and 200 and watch the map distort, then try UMAP (umap-learn). These same tools will reappear when you inspect latent spaces (Week 43) and learned embeddings (Week 47).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Beyond linear · High-D similarities: P · Perplexity · Low-D Q & the t-kernel · The objective: KL(P‖Q) · UMAP. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
p_{j|i} → symmetrized joint Pᵢⱼ — and reproduce it in NumPy (P[0,1]=0.2465)= 2^{H(Pᵢ)} and binary-search each σᵢ to hit itQ and show its heavy tail cures crowding (0.059 vs 0.0003 at d=4)min D_KL(P‖Q) (good map 0.183 < bad map 1.087) and read off its spring-like gradient| idea | the one thing to remember |
|---|---|
| high-D P | Gaussian similarities, per-point σ, symmetrized to sum 1 |
| perplexity | 2^{entropy} = effective # neighbors; tune σ to it |
| low-D Q | Student-t tail beats crowding — clusters breathe |
| objective | minimize KL(P‖Q); P-weighted ⇒ local, not global |
| UMAP | graph cross-entropy: more global, faster |
| caveat | viz only — no transform, non-deterministic geometry |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.