Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP)

USAAIO Lesson 58, from Phase 3. It builds Kernel PCA from scratch - the kernel matrix, centering, eigendecomposition, and projection - and covers the choice between an RBF and a polynomial kernel. It then covers the mechanics of t-SNE, including perplexity, the learning rate, the KL divergence, and why projecting a new point fails, and the theory of UMAP, with its fuzzy simplicial sets, cross-entropy optimization, and inductive mapping. It ends by comparing the methods on the Swiss roll. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 31 slides.

Subject: Machine Learning · 58 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Nonlinear Dimensionality Reduction: Kernel PCA, t-SNE, UMAP

Title

USAAIO · Lesson 58 · Phase 3

PCA lives in the original feature space — but real manifolds are curved. Today: three methods that work in the implicit or neighbourhood space instead.

2. By the end of this lesson you can

Objectives

  1. Implement Kernel PCA from scratch: build K, center it, eigendecompose, project
  2. Choose the right kernel (RBF for local/curved structure, polynomial for global polynomial manifolds)
  3. Explain t-SNE's mechanics: perplexity, learning_rate, KL divergence, and why it is transductive
  4. Describe UMAP's graph-based objective and why it supports inductive (new-point) projection
  5. Decide which method to use given data geometry and downstream requirements

3. What survived from Transfer Learning & Fine-Tuning?

Warm-up

Discussion prompt

Before we open Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP): without looking back, what was the main idea of Transfer Learning & Fine-Tuning, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

feature extraction vs fine-tuning, freezing pretrained layers, the 4-quadrant decision rule (data size x domain similarity), progressive layer unfreezing, differential learning rates (10x lower for pretrained layers), and data augmentation strategies (mixup, cutmix). Build a frozen-backbone classifier and progressively unfreeze it.

4. Kernel PCA: PCA in implicit space

Section

Part 1 of 3

5. The kernel trick for dimensionality reduction

Concept

Linear PCA maximises variance in the original feature space. If data lies on a curved manifold, the top linear directions miss the structure.

\[ \text{Kernel PCA: replace }X^\top X \text{ with the kernel matrix }K,\; K_{ij} = \kappa(x_i, x_j) \]

We never explicitly compute the lifted coordinates — the kernel function implicitly measures dot products in a (possibly infinite-dimensional) feature space.

6. Break it if you can: The kernel trick for dimensionality reduction

Counterexample

Discussion prompt

Linear PCA maximises variance in the original feature space. If data lies on a curved manifold, the top linear directions miss the structure.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

We never explicitly compute the lifted coordinates — the kernel function implicitly measures dot products in a (possibly infinite-dimensional) feature space.

7. The Kernel PCA algorithm

Concept

  1. Compute kernel matrix K (n×n), where K[i,j] = κ(xᵢ, xⱼ)
  2. Center K in feature space: K̃ = K − 1ₙK − K1ₙ + 1ₙK1ₙ (1ₙ = ones/n)
  3. Eigendecompose K̃ = VΛVᵀ; keep top d eigenvectors
  4. Project new points via inner products with the training kernel — no explicit Φ(x) needed

This is structurally identical to linear PCA on the Gram matrix — the only change is the matrix you fill.

8. Guess the shape of the answer: Kernel matrix and centering (3-point trace)

Estimation

Predict first

Three 2-D points: x₀=[0,0], x₁=[1,0], x₂=[0,1]. RBF kernel with γ=0.5: κ(xᵢ,xⱼ) = exp(−0.5·‖xᵢ−xⱼ‖²).

Commit before you compute: what does Kernel matrix and centering (3-point trace) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Center K: row_means=[0.7377, 0.6581, 0.6581], grand_mean=0.6847

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean; centering removes the mean in the implicit feature space, matching the 'subtract mean' step of linear PCA.

9. Kernel matrix and centering (3-point trace)

Worked example

Three 2-D points: x₀=[0,0], x₁=[1,0], x₂=[0,1]. RBF kernel with γ=0.5: κ(xᵢ,xⱼ) = exp(−0.5·‖xᵢ−xⱼ‖²).

import numpy as np
X = np.array([[0.,0.],[1.,0.],[0.,1.]])
gamma = 0.5
K = np.array([[np.exp(-gamma*np.sum((X[i]-X[j])**2))
               for j in range(3)] for i in range(3)])
print(K.round(4))
K[i,j]j=0j=1j=2
i=01.00000.60650.6065
i=10.60651.00000.3679
i=20.60650.36791.0000

Center K: row_means=[0.7377, 0.6581, 0.6581], grand_mean=0.6847

Why: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean; centering removes the mean in the implicit feature space, matching the 'subtract mean' step of linear PCA.

K̃[i,j]j=0j=1j=2
i=00.2093−0.1046−0.1046
i=1−0.10460.3684−0.2637
i=2−0.1046−0.26370.3684

10. Fill in: j=0 for Kernel matrix and centering (3-point trace)

Comparison

Comparison matrix

From Kernel matrix and centering (3-point trace): refill the j=0 column from what you know. The rest of the table is as it appeared.

K[i,j]j=0j=1j=2
i=01.00000.60650.6065
i=10.60651.00000.3679
i=20.60650.36791.0000

11. Guess the shape of the answer: Eigendecompose K̃ → top-2 projection (circle…

Estimation

Predict first

50-point noisy double-circle (n=50). Compute K (RBF γ=1), center it, eigendecompose, project onto top 2 eigenvectors.

Commit before you compute: what does Eigendecompose K̃ → top-2 projection (circle data) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Project: zᵢ = K̃·v / √λ (column i of Z)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Normalising by √λ gives unit-variance components in feature space — same convention as linear PCA's scaling by 1/√λ.

12. Eigendecompose K̃ → top-2 projection (circle data)

Worked example

50-point noisy double-circle (n=50). Compute K (RBF γ=1), center it, eigendecompose, project onto top 2 eigenvectors.

import numpy as np
np.random.seed(42)
n = 50; theta = np.linspace(0, 4*np.pi, n)
X = np.c_[np.cos(theta)+0.05*np.random.randn(n),
          np.sin(theta)+0.05*np.random.randn(n)]

def rbf_K(X, g=1.0):
    K = np.zeros((len(X),len(X)))
    for i in range(len(X)):
        for j in range(len(X)):
            K[i,j] = np.exp(-g*np.sum((X[i]-X[j])**2))
    return K

def center_K(K):
    one = np.ones_like(K)/len(K)
    return K - one@K - K@one + one@K@one

K = center_K(rbf_K(X))
vals, vecs = np.linalg.eigh(K)
vals, vecs = vals[::-1], vecs[:,::-1]
Z = np.c_[K@vecs[:,0]/np.sqrt(vals[0]),
          K@vecs[:,1]/np.sqrt(vals[1])]
print('top eigenvalues:', vals[:5].round(4))
print('Z[0:3]:\n', Z[:3].round(4))
quantityvalue
λ₁ (top eigenvalue)11.0288
λ₂10.3913
λ₃4.8680
Z[0] (first projection)[−0.6460, −0.1292]
Z[1][−0.6555, 0.0189]

Project: zᵢ = K̃·v / √λ (column i of Z)

Why: Normalising by √λ gives unit-variance components in feature space — same convention as linear PCA's scaling by 1/√λ.

13. What each one costs: Eigendecompose K̃ → top-2 projection (circle data)

Trade off

Comparison matrix

From Eigendecompose K̃ → top-2 projection (circle data): every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

quantityvalue
λ₁ (top eigenvalue)11.0288
λ₂10.3913
λ₃4.8680
Z[0] (first projection)[−0.6460, −0.1292]
Z[1][−0.6555, 0.0189]

14. Choosing the kernel

Concept

kernelκ(x,z)use when
RBF (Gaussian)exp(−γ‖x−z‖²)local/curved structure; Swiss roll, concentric rings
polynomial (degree d)(xᵀz + c)^ddata lives on a polynomial manifold; image patches
linearxᵀzequivalent to standard PCA — baseline check

The γ hyperparameter of the RBF kernel controls locality: large γ → only very close points have high similarity; small γ → similarity decays slowly (wider neighbourhood).

15. By analogy: Choosing the kernel

Analogy

Discussion prompt

Explain Choosing the kernel by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The γ hyperparameter of the RBF kernel controls locality: large γ → only very close points have high similarity; small γ → similarity decays slowly (wider neighbourhood).

16. Something is wrong here: projecting new points like linear PCA

Anomaly

Predict first

A student writes this, and it looks reasonable:

A new test point x* arrives. Compute its lifted feature vector Φ(x*) and multiply by the eigenvectors V.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This requires knowing Φ(x*) explicitly — but for RBF, Φ maps into infinite dimensions.

Compute the kernel vector k[i] = κ(x, xᵢ) between the new point and all training points, center it, then project.

Why: This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.

17. Trap: projecting new points like linear PCA

Trap

The trap

A new test point x* arrives. Compute its lifted feature vector Φ(x*) and multiply by the eigenvectors V.

z* = Vᵀ Φ(x*)

Why: This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.

The fix

Compute the kernel vector k[i] = κ(x, xᵢ) between the new point and all training points, center it, then project.

k̃* = k* − (1/n)K1 − (1/n)k1 + (1/n²)1ᵀK1; z = k̃*ᵀ (V / √Λ)

Why: The projection uses only kernel evaluations — no explicit Φ. The centering term accounts for the training-set mean in feature space.

18. Break it on purpose: projecting new points like linear PCA

Break the constraint

Discussion prompt

The rule this trap just fixed:

The projection uses only kernel evaluations — no explicit Φ. The centering term accounts for the training-set mean in feature space.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.

19. t-SNE: neighbourhood-preserving embedding

Section

Part 2 of 3

20. t-SNE core idea

Concept

t-SNE converts high-dim pairwise distances into conditional probabilities P[j|i]: how likely is xⱼ to be a neighbour of xᵢ under a Gaussian centred at xᵢ?

\[ P_{j|i} = \frac{\exp(-\|x_i-x_j\|^2 / 2\sigma_i^2)}{\sum_{k\neq i}\exp(-\|x_i-x_k\|^2 / 2\sigma_i^2)} \]

In the 2-D embedding it uses a t-distribution (heavy tail) for Q: this lets dissimilar points spread far apart, preventing cluster collapse.

\[ Q_{ij} = \frac{(1+\|y_i-y_j\|^2)^{-1}}{\sum_{k\neq l}(1+\|y_k-y_l\|^2)^{-1}} \]

21. Teach it back: t-SNE core idea

Explain it

Discussion prompt

Explain t-SNE core idea to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

In the 2-D embedding it uses a t-distribution (heavy tail) for Q: this lets dissimilar points spread far apart, preventing cluster collapse.

22. t-SNE: perplexity and learning_rate

Concept

parameterwhat it controlstypical range
perplexityeffective number of neighbours per point (sets σᵢ)5–50
learning_ratestep size in gradient descent on KL(P‖Q)10–1000 (default 200)
max_itertotal optimisation steps250–2000 (default 1000)

Low perplexity → tight local clusters but distorted global layout. High perplexity → more global, but clusters blur. Perplexity ~30 is the robust default for most visualisation tasks.

23. Teach it back: t-SNE: perplexity and learning_rate

Explain it

Discussion prompt

Explain t-SNE: perplexity and learning_rate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Low perplexity → tight local clusters but distorted global layout. High perplexity → more global, but clusters blur. Perplexity ~30 is the robust default for most visualisation tasks.

24. Guess the shape of the answer: t-SNE on digits[:200] — verified trace

Estimation

Predict first

Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.

Commit before you compute: what does t-SNE on digits[:200] — verified trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. t-SNE minimises KL(P‖Q) by gradient descent on the low-dim coordinates y — the optimised embedding is the result, not a learned mapping.

25. t-SNE on digits[:200] — verified trace

Worked example

Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.

from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler
import numpy as np

digits = load_digits()
X = StandardScaler().fit_transform(digits.data[:200])

tsne = TSNE(n_components=2, perplexity=30,
            learning_rate=200.0, max_iter=1000,
            random_state=42, init='pca')
Z = tsne.fit_transform(X)
print(f'shape: {Z.shape}')
print(f'KL divergence: {tsne.kl_divergence_:.4f}')
metricvalue
output shape(200, 2)
final KL divergence0.3019
embedding x-range[−46.0, 50.8]
embedding y-range[−47.1, 41.6]

KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)

Why: t-SNE minimises KL(P‖Q) by gradient descent on the low-dim coordinates y — the optimised embedding is the result, not a learned mapping.

26. Work backwards from the answer: t-SNE on digits[:200] — verified trace

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.

27. Why t-SNE is transductive — no new-point projection

Concept

t-SNE optimises the embedding coordinates y₁…yₙ directly — it learns no function f(x) = y. The gradient update moves each yᵢ; there is no weight matrix to apply to a new xᵢ.

This is the key practical limitation: t-SNE is for static visualisation, not online or streaming inference.

28. UMAP: graph-based with inductive mapping

Section

Part 3 of 3

29. UMAP's two-stage construction

Concept

  1. Build a fuzzy simplicial set (graph): for each point xᵢ, find k nearest neighbours; assign edge weights via a normalised exponential decaying with distance
  2. Symmetrise the graph: combine directed weights into an undirected fuzzy set (high-dim graph Gₕ)
  3. Optimise a low-dim graph Gₗ to have the same topological structure — minimise cross-entropy between Gₕ and Gₗ edge weights

The cross-entropy objective rewards matching edges (attractive force on near neighbours) and punishes phantom edges (repulsive force on non-neighbours).

30. By analogy: UMAP's two-stage construction

Analogy

Discussion prompt

Explain UMAP's two-stage construction by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The cross-entropy objective rewards matching edges (attractive force on near neighbours) and punishes phantom edges (repulsive force on non-neighbours).

31. UMAP vs t-SNE — inductive property

Concept

propertyt-SNEUMAP
objectiveKL(P‖Q)cross-entropy on fuzzy graph edges
new-point projectionno — transductive onlyyes — reducer.transform(X_new)
global structurepoor (crowding problem)better preserved
speed (n=10000)minutesseconds
interpretable distancesno (arbitrary scale)roughly meaningful

UMAP learns an approximate parametric mapping via the graph — fit on training data, transform on new points. This makes it usable in a production pipeline.

32. Fill in: t-SNE for UMAP vs t-SNE — inductive property

Comparison

Comparison matrix

From UMAP vs t-SNE — inductive property: refill the t-SNE column from what you know. The rest of the table is as it appeared.

propertyt-SNEUMAP
objectiveKL(P‖Q)cross-entropy on fuzzy graph edges
new-point projectionno — transductive onlyyes — reducer.transform(X_new)
global structurepoor (crowding problem)better preserved
speed (n=10000)minutesseconds
interpretable distancesno (arbitrary scale)roughly meaningful

33. Something is wrong here: interpreting t-SNE cluster distances

Anomaly

Predict first

A student writes this, and it looks reasonable:

Cluster A is far from cluster B in the 2-D t-SNE plot, so A and B must be very dissimilar in the original space.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity.

Use t-SNE for local neighbourhood structure and cluster identity — not for measuring how far apart clusters are.

Why: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity. The distances between clusters are meaningless — only within-cluster structure and neighbourhood identity are preserved.

34. Trap: interpreting t-SNE cluster distances

Trap

The trap

Cluster A is far from cluster B in the 2-D t-SNE plot, so A and B must be very dissimilar in the original space.

Compare inter-cluster distances in the t-SNE embedding

Why: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity. The distances between clusters are meaningless — only within-cluster structure and neighbourhood identity are preserved.

The fix

Use t-SNE for local neighbourhood structure and cluster identity — not for measuring how far apart clusters are.

Verify inter-cluster relationships with pairwise cosine similarity or kNN graphs in the original space

Why: t-SNE compresses global distances. If you need distance-meaningful projections, use PCA, kPCA with a global kernel, or UMAP (which better preserves global topology).

35. Which of these survive contact with Lesson 58: Nonlinear Dimensionality…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Linear PCA maximises variance in the original feature space. If data lies on a curved manifold, the top linear directions miss the structure.; The γ hyperparameter of the RBF kernel controls locality: large γ → only very close points have high similarity; small γ → similarity decays slowly (wider neighbourhood).; In the 2-D embedding it uses a t-distribution (heavy tail) for Q: this lets dissimilar points spread far apart, preventing cluster collapse.
Breaks
A new test point x* arrives. Compute its lifted feature vector Φ(x*) and multiply by the eigenvectors V.; Cluster A is far from cluster B in the 2-D t-SNE plot, so A and B must be very dissimilar in the original space.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP) puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

36. Method comparison: Swiss roll

Concept

The Swiss roll is a 2-D surface coiled in 3-D. A good method should unroll it — recovering the two intrinsic dimensions.

methodSwiss roll resultwhy
PCA (linear)still coiled, EVR=0.73can't bend the axes
kPCA (RBF γ=1)partially unrolls local patchescaptures local curvature
t-SNE (perp=30)unrolls into separate colour bandspreserves neighbourhood structure

PCA explained variance sum=0.7282 with 2 components — high, but that's because PCA projects the roll flat rather than unrolling it.

37. Where does each piece belong: Lesson 58: Nonlinear Dimensionality…

Sorting

Sort into buckets

These are the pieces of Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP), out of order. Put each one back under the part of the lesson it belongs to.

Kernel PCA: PCA in implicit space
The kernel trick for dimensionality reduction; The Kernel PCA algorithm; Kernel matrix and centering (3-point trace)
t-SNE: neighbourhood-preserving embedding
t-SNE core idea; t-SNE: perplexity and learning_rate; t-SNE on digits[:200] — verified trace
UMAP: graph-based with inductive mapping
UMAP's two-stage construction; UMAP vs t-SNE — inductive property; Method comparison: Swiss roll
s1
Kernel PCA: PCA in implicit space is where Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP) puts The kernel trick for dimensionality reduction, The Kernel PCA algorithm, Kernel matrix and centering (3-point trace). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
t-SNE: neighbourhood-preserving embedding is where Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP) puts t-SNE core idea, t-SNE: perplexity and learning_rate, t-SNE on digits[:200] — verified trace. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
UMAP: graph-based with inductive mapping is where Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP) puts UMAP's two-stage construction, UMAP vs t-SNE — inductive property, Method comparison: Swiss roll. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

38. Without one step: Nonlinear DR selection recipe

Constraint

Discussion prompt

Run Nonlinear DR selection recipe with this step confiscated:

t-SNE for static 2-D visualisation of cluster structure; set perplexity 5–50, use init='pca', ignore inter-cluster distances

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Linear baseline first (PCA, Lesson 30): check explained-variance ratio; if EVR ≥ 0.9 in 2 components, linear is sufficient
  2. Kernel PCA when you need an algebraically interpretable embedding and may project new points; pick kernel by geometry (RBF for local curves, poly for…
  3. t-SNE for static 2-D visualisation of cluster structure; set perplexity 5–50, use init='pca', ignore inter-cluster distances
  4. UMAP when you need online inference (new-point projection), faster runtime, or better-preserved global topology
  5. Sanity check: visualise with a colour label; if well-separated clusters appear, the method is working

39. Nonlinear DR selection recipe

Pattern

  1. Linear baseline first (PCA, Lesson 30): check explained-variance ratio; if EVR ≥ 0.9 in 2 components, linear is sufficient
  2. Kernel PCA when you need an algebraically interpretable embedding and may project new points; pick kernel by geometry (RBF for local curves, poly for polynomial manifolds)
  3. t-SNE for static 2-D visualisation of cluster structure; set perplexity 5–50, use init='pca', ignore inter-cluster distances
  4. UMAP when you need online inference (new-point projection), faster runtime, or better-preserved global topology
  5. Sanity check: visualise with a colour label; if well-separated clusters appear, the method is working

40. Where does it stop working: Nonlinear DR selection recipe

Edge cases

Discussion prompt

Nonlinear DR selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Linear baseline first (PCA, Lesson 30): check explained-variance ratio; if EVR ≥ 0.9 in 2 components, linear is sufficient
  2. Kernel PCA when you need an algebraically interpretable embedding and may project new points; pick kernel by geometry (RBF for local curves, poly for…
  3. t-SNE for static 2-D visualisation of cluster structure; set perplexity 5–50, use init='pca', ignore inter-cluster distances
  4. UMAP when you need online inference (new-point projection), faster runtime, or better-preserved global topology
  5. Sanity check: visualise with a colour label; if well-separated clusters appear, the method is working

41. Rule out three: Check yourself — kernel centering

Elimination

Eliminate the wrong options

Why must the kernel matrix K be centered before eigendecomposition in Kernel PCA?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. To zero the mean in the implicit feature space Φ, matching the 'subtract mean' step of linear PCA
  • B. To make K positive-definite
  • C. To normalize the kernel values to the range [0, 1]
  • D. To remove the diagonal (self-similarity) terms

Survives elimination: A

Why: Linear PCA subtracts the feature-space mean before computing covariances. Kernel centering does the same thing without ever computing Φ explicitly: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean projects all points to be mean-zero in feature space.

42. Check yourself — kernel centering

Check

Think through the centering formula before clicking.

Check your understanding

Why must the kernel matrix K be centered before eigendecomposition in Kernel PCA?

  • A. To zero the mean in the implicit feature space Φ, matching the 'subtract mean' step of linear PCA (correct)
  • B. To make K positive-definite
  • C. To normalize the kernel values to the range [0, 1]
  • D. To remove the diagonal (self-similarity) terms

Answer: A

Why: Linear PCA subtracts the feature-space mean before computing covariances. Kernel centering does the same thing without ever computing Φ explicitly: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean projects all points to be mean-zero in feature space.

Why B tempts people
Positive-definiteness comes from the choice of Mercer kernel (e.g. RBF), not from centering. A centered kernel can still have tiny numerical negatives.
Why C tempts people
Centering does not bound values to [0,1]; the centered matrix contains both positive and negative entries (as the 3-point trace showed: K̃[1,2] = −0.2637).
Why D tempts people
The diagonal terms K[i,i]=1 (for RBF) are reduced but not removed; centering is a mean-subtraction, not a zeroing of the diagonal.

43. Answer it before you see the options: Check yourself — t-SNE vs UMAP

Prediction

Predict first

Which method supports projecting a new unseen point x* without re-running on the full training set?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: UMAP (reducer.transform(x*))

Why: UMAP learns an approximate parametric mapping during fit; transform applies it to new points in O(k·d) time. t-SNE has no learned mapping — it only produces coordinates for the training set; adding any new point requires reoptimising the full embedding.

44. Check yourself — t-SNE vs UMAP

Check

A production system scores incoming images in real time. The team wants a 2-D embedding refreshed for each new image.

Check your understanding

Which method supports projecting a new unseen point x* without re-running on the full training set?

  • A. UMAP (reducer.transform(x*)) (correct)
  • B. t-SNE — just append x* and run one more gradient step
  • C. t-SNE with n_iter=1 for each new point
  • D. Neither — nonlinear methods never support new-point projection

Answer: A

Why: UMAP learns an approximate parametric mapping during fit; transform applies it to new points in O(k·d) time. t-SNE has no learned mapping — it only produces coordinates for the training set; adding any new point requires reoptimising the full embedding.

Why B tempts people
t-SNE has no fixed mapping to apply one step to. The gradient moves all existing yᵢ relative to each other; a single new step would corrupt the existing embedding, not extend it.
Why C tempts people
Reducing max_iter to 1 gives an unconverged, meaningless embedding — t-SNE needs hundreds of iterations to reach a usable state, and still has no mechanism to add just one new point.
Why D tempts people
Kernel PCA and UMAP both support new-point projection via a kernel vector or the learned mapping respectively. 'Neither' is false.

45. Rule out three: Check yourself — perplexity

Elimination

Eliminate the wrong options

Which perplexity setting is most likely to reveal the fine cluster structure?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. perplexity = 5 (a few neighbours per point)
  • B. perplexity = 100 (≈ 20% of all points)
  • C. perplexity must equal n − 1 = 499 for best results
  • D. perplexity does not affect cluster resolution, only speed

Survives elimination: A

Why: Perplexity is roughly the number of effective neighbours. Setting it near the cluster size (5–10 for 8-point clusters) tunes σᵢ so intra-cluster similarities dominate. A perplexity of 100 blurs boundaries by pulling in inter-cluster neighbours.

46. Check yourself — perplexity

Check

A dataset has n=500, with tight local clusters of ~8 points each.

Check your understanding

Which perplexity setting is most likely to reveal the fine cluster structure?

  • A. perplexity = 5 (a few neighbours per point) (correct)
  • B. perplexity = 100 (≈ 20% of all points)
  • C. perplexity must equal n − 1 = 499 for best results
  • D. perplexity does not affect cluster resolution, only speed

Answer: A

Why: Perplexity is roughly the number of effective neighbours. Setting it near the cluster size (5–10 for 8-point clusters) tunes σᵢ so intra-cluster similarities dominate. A perplexity of 100 blurs boundaries by pulling in inter-cluster neighbours.

Why B tempts people
perplexity=100 treats 100 points as neighbours, far exceeding the cluster size; the Gaussian spreads across multiple clusters and washes out the fine structure.
Why C tempts people
perplexity must be strictly less than n; n−1 would make every point equally a neighbour of every other, giving a uniform P and no useful structure.
Why D tempts people
Perplexity is the most impactful hyperparameter for cluster resolution — it sets σᵢ per point, directly controlling which neighbours dominate P[j|i].

47. Your turn: implement Kernel PCA

Section

Project

48. Project: Kernel PCA from scratch + DR comparison

Concept

Build Kernel PCA from scratch for the RBF kernel, then compare PCA, kPCA, and t-SNE on the Swiss roll dataset.

#milestonemethod
1RBF kernel matrix + centering on 3-point toynumpy only
2Eigendecompose K̃ and project onto 2 componentsnp.linalg.eigh
3Compare PCA vs kPCA vs t-SNE on Swiss rollsklearn

Build rules: implement rbf_kernel, center_kernel, and kpca_project as separate functions; the full program calls all three in sequence.

49. Break it if you can: Project: Kernel PCA from scratch + DR comparison

Counterexample

Discussion prompt

Build Kernel PCA from scratch for the RBF kernel, then compare PCA, kPCA, and t-SNE on the Swiss roll dataset.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: implement rbf_kernel, center_kernel, and kpca_project as separate functions; the full program calls all three in sequence.

50. Milestone 1 — kernel matrix and centering

Worked example

Your turn: compute K (RBF γ=0.5) and K̃ for X = [[0,0],[1,0],[0,1]]. Predict the grand mean before running.

Hint: K[i,j] = exp(−0.5·‖xᵢ−xⱼ‖²); centering: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean.

import numpy as np
X = np.array([[0.,0.],[1.,0.],[0.,1.]])
gamma = 0.5
K = np.array([[np.exp(-gamma*np.sum((X[i]-X[j])**2))
               for j in range(3)] for i in range(3)])
row_m = K.mean(1); col_m = K.mean(0); gm = K.mean()
K_c = K - row_m[:,None] - col_m[None,:] + gm
print('K:\n', K.round(4))
print('K_c:\n', K_c.round(4))
quantityvalue
K[0,1] = K[0,2]0.6065
K[1,2]0.3679
grand mean0.6847
K̃[1,1] = K̃[2,2]0.3684
K̃[1,2]−0.2637

51. Milestone 2 — eigendecompose and project

Worked example

Your turn: eigendecompose K̃ from Milestone 1. Predict the number of non-zero eigenvalues.

Hint: np.linalg.eigh returns ascending order — flip to descending. Project: Z[:,i] = K_c @ v[:,i] / sqrt(λᵢ).

vals, vecs = np.linalg.eigh(K_c)
vals, vecs = vals[::-1], vecs[:,::-1]   # descending
print('eigenvalues:', vals.round(4))

# Project onto top 2
Z = np.zeros((3, 2))
for i in range(2):
    Z[:,i] = K_c @ vecs[:,i] / np.sqrt(vals[i])
print('Z (top-2 embedding):\n', Z.round(4))
quantityvalue
λ₁0.6321
λ₂0.3139
λ₃ (zero — centering)≈ 0.0000
Z[0][−0.4124, −0.5824]
Z[1][0.7275, −0.0000]

52. Milestone 3 — Swiss roll: PCA vs kPCA vs t-SNE

Worked example

Your turn: apply all three methods to the Swiss roll (n=300). Predict which method unrolls it best.

Hint: make_swiss_roll(300, noise=0.1, random_state=42); StandardScaler; kPCA RBF γ=1; t-SNE perplexity=30.

from sklearn.datasets import make_swiss_roll
from sklearn.decomposition import PCA, KernelPCA
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler

X_sw, _ = make_swiss_roll(300, noise=0.1, random_state=42)
X_sc = StandardScaler().fit_transform(X_sw)

pca = PCA(2).fit_transform(X_sc)
kp  = KernelPCA(2,'rbf',gamma=1.0).fit_transform(X_sc)
ts  = TSNE(2,perplexity=30,learning_rate=200,
           max_iter=1000,random_state=42,init='pca').fit_transform(X_sc)
print('PCA EVR sum:', PCA(2).fit(X_sc).explained_variance_ratio_.sum().round(4))
methodunrolls?note
PCAno — EVR=0.7282projects coil flat
kPCA RBF γ=1partiallylocal patches aligned
t-SNE perp=30yes — colour bandsneighbourhood structure preserved

53. What each one costs: Milestone 3 — Swiss roll: PCA vs kPCA vs t-SNE

Trade off

Comparison matrix

From Milestone 3 — Swiss roll: PCA vs kPCA vs t-SNE: every row here is a choice with a cost. Fill the unrolls? column, then say which row you would actually pick and what you give up for it.

methodunrolls?note
PCAno — EVR=0.7282projects coil flat
kPCA RBF γ=1partiallylocal patches aligned
t-SNE perp=30yes — colour bandsneighbourhood structure preserved

54. The full program

Concept

import numpy as np
from sklearn.datasets import make_swiss_roll
from sklearn.decomposition import PCA, KernelPCA
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler

# 1. Kernel PCA from scratch (toy)
X3 = np.array([[0.,0.],[1.,0.],[0.,1.]]); g=0.5
K = np.array([[np.exp(-g*np.sum((X3[i]-X3[j])**2)) for j in range(3)] for i in range(3)])
row_m=K.mean(1); col_m=K.mean(0); gm=K.mean()
K_c = K - row_m[:,None] - col_m[None,:] + gm
vals,vecs = np.linalg.eigh(K_c); vals,vecs=vals[::-1],vecs[:,::-1]
print('Eigenvalues:', vals.round(4))

# 2. Swiss roll comparison
X_sw,_ = make_swiss_roll(300, noise=0.1, random_state=42)
X_sc = StandardScaler().fit_transform(X_sw)
print('PCA EVR sum:', PCA(2).fit(X_sc).explained_variance_ratio_.sum().round(4))
ts = TSNE(2,perplexity=30,learning_rate=200,max_iter=1000,
          random_state=42,init='pca').fit_transform(X_sc)
print('t-SNE KL:', ts.shape, '(KL logged during fit)')
outputvalue
3-pt eigenvalues[0.6321, 0.3139, ≈0]
Swiss PCA EVR sum0.7282
t-SNE output shape(300, 2)

If your scratch kPCA eigenvalues match and t-SNE runs — you've built the full pipeline.

55. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
3-pt eigenvalues[0.6321, 0.3139, ≈0]
Swiss PCA EVR sum0.7282
t-SNE output shape(300, 2)

56. Show it off

Concept

Out loud, slides closed: (1) walk through all four Kernel PCA steps without looking at notes, (2) explain why t-SNE cluster distances are meaningless, (3) give one reason UMAP is preferred for production pipelines.

Stretch (homework): implement kPCA new-point projection (the kernel vector centering formula); apply t-SNE with perplexity ∈ {5, 30, 50} to digits and explain what changes; analyse why you can't use t-SNE for new data but UMAP handles it.

57. Connect it up: Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP)

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Kernel PCA: PCA in implicit space · t-SNE: neighbourhood-preserving embedding · UMAP: graph-based with inductive mapping · Your turn: implement Kernel PCA. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

58. What you can do now

Recap

ideathe one thing to remember
Kernel PCAcenter K before eigh — or PCs are shifted
new-point projectionkPCA: kernel vector; t-SNE: impossible; UMAP: transform()
t-SNE perplexity≈ cluster size; inter-cluster distances are meaningless
UMAP vs t-SNEUMAP is inductive; t-SNE is transductive

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 58 — Kernel PCA, t-SNE, UMAP — Barron · USAAIO Round 2 Preparation, 2026
  2. Kernel PCA from scratch (RBF gamma=0.5, 3-point trace; circle data eigenvalues), t-SNE KL=0.3019 on digits[:200], Swiss roll PCA/kPCA/t-SNE — all verified — sklearn 1.x + numpy 2.2.6 + scipy, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108