USAAIO Lesson 58, from Phase 3. It builds Kernel PCA from scratch - the kernel matrix, centering, eigendecomposition, and projection - and covers the choice between an RBF and a polynomial kernel. It then covers the mechanics of t-SNE, including perplexity, the learning rate, the KL divergence, and why projecting a new point fails, and the theory of UMAP, with its fuzzy simplicial sets, cross-entropy optimization, and inductive mapping. It ends by comparing the methods on the Swiss roll. All the numbers were verified with sklearn 1.x and numpy 2.2.6. The lesson runs to 31 slides.
Subject: Machine Learning · 58 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 58 · Phase 3
PCA lives in the original feature space — but real manifolds are curved. Today: three methods that work in the implicit or neighbourhood space instead.
Objectives
perplexity, learning_rate, KL divergence, and why it is transductiveWarm-up
Discussion prompt
Before we open Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP): without looking back, what was the main idea of Transfer Learning & Fine-Tuning, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
feature extraction vs fine-tuning, freezing pretrained layers, the 4-quadrant decision rule (data size x domain similarity), progressive layer unfreezing, differential learning rates (10x lower for pretrained layers), and data augmentation strategies (mixup, cutmix). Build a frozen-backbone classifier and progressively unfreeze it.
Section
Part 1 of 3
Concept
Linear PCA maximises variance in the original feature space. If data lies on a curved manifold, the top linear directions miss the structure.
\[ \text{Kernel PCA: replace }X^\top X \text{ with the kernel matrix }K,\; K_{ij} = \kappa(x_i, x_j) \]
We never explicitly compute the lifted coordinates — the kernel function implicitly measures dot products in a (possibly infinite-dimensional) feature space.
Counterexample
Discussion prompt
Linear PCA maximises variance in the original feature space. If data lies on a curved manifold, the top linear directions miss the structure.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
We never explicitly compute the lifted coordinates — the kernel function implicitly measures dot products in a (possibly infinite-dimensional) feature space.
Concept
This is structurally identical to linear PCA on the Gram matrix — the only change is the matrix you fill.
Estimation
Predict first
Three 2-D points: x₀=[0,0], x₁=[1,0], x₂=[0,1]. RBF kernel with γ=0.5: κ(xᵢ,xⱼ) = exp(−0.5·‖xᵢ−xⱼ‖²).
Commit before you compute: what does Kernel matrix and centering (3-point trace) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Center K: row_means=[0.7377, 0.6581, 0.6581], grand_mean=0.6847
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean; centering removes the mean in the implicit feature space, matching the 'subtract mean' step of linear PCA.
Worked example
Three 2-D points: x₀=[0,0], x₁=[1,0], x₂=[0,1]. RBF kernel with γ=0.5: κ(xᵢ,xⱼ) = exp(−0.5·‖xᵢ−xⱼ‖²).
import numpy as np
X = np.array([[0.,0.],[1.,0.],[0.,1.]])
gamma = 0.5
K = np.array([[np.exp(-gamma*np.sum((X[i]-X[j])**2))
for j in range(3)] for i in range(3)])
print(K.round(4))| K[i,j] | j=0 | j=1 | j=2 |
|---|---|---|---|
| i=0 | 1.0000 | 0.6065 | 0.6065 |
| i=1 | 0.6065 | 1.0000 | 0.3679 |
| i=2 | 0.6065 | 0.3679 | 1.0000 |
Center K: row_means=[0.7377, 0.6581, 0.6581], grand_mean=0.6847
Why: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean; centering removes the mean in the implicit feature space, matching the 'subtract mean' step of linear PCA.
| K̃[i,j] | j=0 | j=1 | j=2 |
|---|---|---|---|
| i=0 | 0.2093 | −0.1046 | −0.1046 |
| i=1 | −0.1046 | 0.3684 | −0.2637 |
| i=2 | −0.1046 | −0.2637 | 0.3684 |
Comparison
Comparison matrix
From Kernel matrix and centering (3-point trace): refill the j=0 column from what you know. The rest of the table is as it appeared.
| K[i,j] | j=0 | j=1 | j=2 |
|---|---|---|---|
| i=0 | 1.0000 | 0.6065 | 0.6065 |
| i=1 | 0.6065 | 1.0000 | 0.3679 |
| i=2 | 0.6065 | 0.3679 | 1.0000 |
Estimation
Predict first
50-point noisy double-circle (n=50). Compute K (RBF γ=1), center it, eigendecompose, project onto top 2 eigenvectors.
Commit before you compute: what does Eigendecompose K̃ → top-2 projection (circle data) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Project: zᵢ = K̃·v / √λ (column i of Z)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Normalising by √λ gives unit-variance components in feature space — same convention as linear PCA's scaling by 1/√λ.
Worked example
50-point noisy double-circle (n=50). Compute K (RBF γ=1), center it, eigendecompose, project onto top 2 eigenvectors.
import numpy as np
np.random.seed(42)
n = 50; theta = np.linspace(0, 4*np.pi, n)
X = np.c_[np.cos(theta)+0.05*np.random.randn(n),
np.sin(theta)+0.05*np.random.randn(n)]
def rbf_K(X, g=1.0):
K = np.zeros((len(X),len(X)))
for i in range(len(X)):
for j in range(len(X)):
K[i,j] = np.exp(-g*np.sum((X[i]-X[j])**2))
return K
def center_K(K):
one = np.ones_like(K)/len(K)
return K - one@K - K@one + one@K@one
K = center_K(rbf_K(X))
vals, vecs = np.linalg.eigh(K)
vals, vecs = vals[::-1], vecs[:,::-1]
Z = np.c_[K@vecs[:,0]/np.sqrt(vals[0]),
K@vecs[:,1]/np.sqrt(vals[1])]
print('top eigenvalues:', vals[:5].round(4))
print('Z[0:3]:\n', Z[:3].round(4))| quantity | value |
|---|---|
| λ₁ (top eigenvalue) | 11.0288 |
| λ₂ | 10.3913 |
| λ₃ | 4.8680 |
| Z[0] (first projection) | [−0.6460, −0.1292] |
| Z[1] | [−0.6555, 0.0189] |
Project: zᵢ = K̃·v / √λ (column i of Z)
Why: Normalising by √λ gives unit-variance components in feature space — same convention as linear PCA's scaling by 1/√λ.
Trade off
Comparison matrix
From Eigendecompose K̃ → top-2 projection (circle data): every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| quantity | value |
|---|---|
| λ₁ (top eigenvalue) | 11.0288 |
| λ₂ | 10.3913 |
| λ₃ | 4.8680 |
| Z[0] (first projection) | [−0.6460, −0.1292] |
| Z[1] | [−0.6555, 0.0189] |
Concept
| kernel | κ(x,z) | use when |
|---|---|---|
| RBF (Gaussian) | exp(−γ‖x−z‖²) | local/curved structure; Swiss roll, concentric rings |
| polynomial (degree d) | (xᵀz + c)^d | data lives on a polynomial manifold; image patches |
| linear | xᵀz | equivalent to standard PCA — baseline check |
The γ hyperparameter of the RBF kernel controls locality: large γ → only very close points have high similarity; small γ → similarity decays slowly (wider neighbourhood).
Analogy
Discussion prompt
Explain Choosing the kernel by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The γ hyperparameter of the RBF kernel controls locality: large γ → only very close points have high similarity; small γ → similarity decays slowly (wider neighbourhood).
Anomaly
Predict first
A student writes this, and it looks reasonable:
A new test point x* arrives. Compute its lifted feature vector Φ(x*) and multiply by the eigenvectors V.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This requires knowing Φ(x*) explicitly — but for RBF, Φ maps into infinite dimensions.
Compute the kernel vector k[i] = κ(x, xᵢ) between the new point and all training points, center it, then project.
Why: This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.
Trap
A new test point x* arrives. Compute its lifted feature vector Φ(x*) and multiply by the eigenvectors V.
z* = Vᵀ Φ(x*)
Why: This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.
Compute the kernel vector k[i] = κ(x, xᵢ) between the new point and all training points, center it, then project.
k̃* = k* − (1/n)K1 − (1/n)k1 + (1/n²)1ᵀK1; z = k̃*ᵀ (V / √Λ)
Why: The projection uses only kernel evaluations — no explicit Φ. The centering term accounts for the training-set mean in feature space.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The projection uses only kernel evaluations — no explicit Φ. The centering term accounts for the training-set mean in feature space.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This requires knowing Φ(x) explicitly — but for RBF, Φ maps into infinite dimensions. We never have Φ(x) as a vector.
Section
Part 2 of 3
Concept
t-SNE converts high-dim pairwise distances into conditional probabilities P[j|i]: how likely is xⱼ to be a neighbour of xᵢ under a Gaussian centred at xᵢ?
\[ P_{j|i} = \frac{\exp(-\|x_i-x_j\|^2 / 2\sigma_i^2)}{\sum_{k\neq i}\exp(-\|x_i-x_k\|^2 / 2\sigma_i^2)} \]
In the 2-D embedding it uses a t-distribution (heavy tail) for Q: this lets dissimilar points spread far apart, preventing cluster collapse.
\[ Q_{ij} = \frac{(1+\|y_i-y_j\|^2)^{-1}}{\sum_{k\neq l}(1+\|y_k-y_l\|^2)^{-1}} \]
Explain it
Discussion prompt
Explain t-SNE core idea to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
In the 2-D embedding it uses a t-distribution (heavy tail) for Q: this lets dissimilar points spread far apart, preventing cluster collapse.
Concept
| parameter | what it controls | typical range |
|---|---|---|
| perplexity | effective number of neighbours per point (sets σᵢ) | 5–50 |
| learning_rate | step size in gradient descent on KL(P‖Q) | 10–1000 (default 200) |
| max_iter | total optimisation steps | 250–2000 (default 1000) |
Low perplexity → tight local clusters but distorted global layout. High perplexity → more global, but clusters blur. Perplexity ~30 is the robust default for most visualisation tasks.
Explain it
Discussion prompt
Explain t-SNE: perplexity and learning_rate to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Low perplexity → tight local clusters but distorted global layout. High perplexity → more global, but clusters blur. Perplexity ~30 is the robust default for most visualisation tasks.
Estimation
Predict first
Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.
Commit before you compute: what does t-SNE on digits[:200] — verified trace come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. t-SNE minimises KL(P‖Q) by gradient descent on the low-dim coordinates y — the optimised embedding is the result, not a learned mapping.
Worked example
Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.
from sklearn.datasets import load_digits
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler
import numpy as np
digits = load_digits()
X = StandardScaler().fit_transform(digits.data[:200])
tsne = TSNE(n_components=2, perplexity=30,
learning_rate=200.0, max_iter=1000,
random_state=42, init='pca')
Z = tsne.fit_transform(X)
print(f'shape: {Z.shape}')
print(f'KL divergence: {tsne.kl_divergence_:.4f}')| metric | value |
|---|---|
| output shape | (200, 2) |
| final KL divergence | 0.3019 |
| embedding x-range | [−46.0, 50.8] |
| embedding y-range | [−47.1, 41.6] |
KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)
Why: t-SNE minimises KL(P‖Q) by gradient descent on the low-dim coordinates y — the optimised embedding is the result, not a learned mapping.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
KL divergence = 0.3019: lower is better, 0 is perfect (P = Q everywhere)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Run t-SNE (perplexity=30, learning_rate=200, max_iter=1000) on 200 digit images. Measure final KL divergence.
Concept
t-SNE optimises the embedding coordinates y₁…yₙ directly — it learns no function f(x) = y. The gradient update moves each yᵢ; there is no weight matrix to apply to a new xᵢ.
This is the key practical limitation: t-SNE is for static visualisation, not online or streaming inference.
Section
Part 3 of 3
Concept
The cross-entropy objective rewards matching edges (attractive force on near neighbours) and punishes phantom edges (repulsive force on non-neighbours).
Analogy
Discussion prompt
Explain UMAP's two-stage construction by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The cross-entropy objective rewards matching edges (attractive force on near neighbours) and punishes phantom edges (repulsive force on non-neighbours).
Concept
| property | t-SNE | UMAP |
|---|---|---|
| objective | KL(P‖Q) | cross-entropy on fuzzy graph edges |
| new-point projection | no — transductive only | yes — reducer.transform(X_new) |
| global structure | poor (crowding problem) | better preserved |
| speed (n=10000) | minutes | seconds |
| interpretable distances | no (arbitrary scale) | roughly meaningful |
UMAP learns an approximate parametric mapping via the graph — fit on training data, transform on new points. This makes it usable in a production pipeline.
Comparison
Comparison matrix
From UMAP vs t-SNE — inductive property: refill the t-SNE column from what you know. The rest of the table is as it appeared.
| property | t-SNE | UMAP |
|---|---|---|
| objective | KL(P‖Q) | cross-entropy on fuzzy graph edges |
| new-point projection | no — transductive only | yes — reducer.transform(X_new) |
| global structure | poor (crowding problem) | better preserved |
| speed (n=10000) | minutes | seconds |
| interpretable distances | no (arbitrary scale) | roughly meaningful |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Cluster A is far from cluster B in the 2-D t-SNE plot, so A and B must be very dissimilar in the original space.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity.
Use t-SNE for local neighbourhood structure and cluster identity — not for measuring how far apart clusters are.
Why: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity. The distances between clusters are meaningless — only within-cluster structure and neighbourhood identity are preserved.
Trap
Cluster A is far from cluster B in the 2-D t-SNE plot, so A and B must be very dissimilar in the original space.
Compare inter-cluster distances in the t-SNE embedding
Why: t-SNE's t-distribution repulsion spreads all clusters apart regardless of true similarity. The distances between clusters are meaningless — only within-cluster structure and neighbourhood identity are preserved.
Use t-SNE for local neighbourhood structure and cluster identity — not for measuring how far apart clusters are.
Verify inter-cluster relationships with pairwise cosine similarity or kNN graphs in the original space
Why: t-SNE compresses global distances. If you need distance-meaningful projections, use PCA, kPCA with a global kernel, or UMAP (which better preserves global topology).
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Concept
The Swiss roll is a 2-D surface coiled in 3-D. A good method should unroll it — recovering the two intrinsic dimensions.
| method | Swiss roll result | why |
|---|---|---|
| PCA (linear) | still coiled, EVR=0.73 | can't bend the axes |
| kPCA (RBF γ=1) | partially unrolls local patches | captures local curvature |
| t-SNE (perp=30) | unrolls into separate colour bands | preserves neighbourhood structure |
PCA explained variance sum=0.7282 with 2 components — high, but that's because PCA projects the roll flat rather than unrolling it.
Sorting
Sort into buckets
These are the pieces of Lesson 58: Nonlinear Dimensionality Reduction (Kernel PCA, t-SNE, UMAP), out of order. Put each one back under the part of the lesson it belongs to.
Constraint
Discussion prompt
Run Nonlinear DR selection recipe with this step confiscated:
t-SNE for static 2-D visualisation of cluster structure; set perplexity 5–50, use init='pca', ignore inter-cluster distances
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
init='pca', ignore inter-cluster distancesPattern
init='pca', ignore inter-cluster distancesEdge cases
Discussion prompt
Nonlinear DR selection recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
init='pca', ignore inter-cluster distancesElimination
Eliminate the wrong options
Why must the kernel matrix K be centered before eigendecomposition in Kernel PCA?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Linear PCA subtracts the feature-space mean before computing covariances. Kernel centering does the same thing without ever computing Φ explicitly: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean projects all points to be mean-zero in feature space.
Check
Think through the centering formula before clicking.
Check your understanding
Why must the kernel matrix K be centered before eigendecomposition in Kernel PCA?
Answer: A
Why: Linear PCA subtracts the feature-space mean before computing covariances. Kernel centering does the same thing without ever computing Φ explicitly: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean projects all points to be mean-zero in feature space.
Prediction
Predict first
Which method supports projecting a new unseen point x* without re-running on the full training set?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: UMAP (reducer.transform(x*))
Why: UMAP learns an approximate parametric mapping during fit; transform applies it to new points in O(k·d) time. t-SNE has no learned mapping — it only produces coordinates for the training set; adding any new point requires reoptimising the full embedding.
Check
A production system scores incoming images in real time. The team wants a 2-D embedding refreshed for each new image.
Check your understanding
Which method supports projecting a new unseen point x* without re-running on the full training set?
Answer: A
Why: UMAP learns an approximate parametric mapping during fit; transform applies it to new points in O(k·d) time. t-SNE has no learned mapping — it only produces coordinates for the training set; adding any new point requires reoptimising the full embedding.
Elimination
Eliminate the wrong options
Which perplexity setting is most likely to reveal the fine cluster structure?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Perplexity is roughly the number of effective neighbours. Setting it near the cluster size (5–10 for 8-point clusters) tunes σᵢ so intra-cluster similarities dominate. A perplexity of 100 blurs boundaries by pulling in inter-cluster neighbours.
Check
A dataset has n=500, with tight local clusters of ~8 points each.
Check your understanding
Which perplexity setting is most likely to reveal the fine cluster structure?
Answer: A
Why: Perplexity is roughly the number of effective neighbours. Setting it near the cluster size (5–10 for 8-point clusters) tunes σᵢ so intra-cluster similarities dominate. A perplexity of 100 blurs boundaries by pulling in inter-cluster neighbours.
Section
Project
Concept
Build Kernel PCA from scratch for the RBF kernel, then compare PCA, kPCA, and t-SNE on the Swiss roll dataset.
| # | milestone | method |
|---|---|---|
| 1 | RBF kernel matrix + centering on 3-point toy | numpy only |
| 2 | Eigendecompose K̃ and project onto 2 components | np.linalg.eigh |
| 3 | Compare PCA vs kPCA vs t-SNE on Swiss roll | sklearn |
Build rules: implement rbf_kernel, center_kernel, and kpca_project as separate functions; the full program calls all three in sequence.
Counterexample
Discussion prompt
Build Kernel PCA from scratch for the RBF kernel, then compare PCA, kPCA, and t-SNE on the Swiss roll dataset.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: implement rbf_kernel, center_kernel, and kpca_project as separate functions; the full program calls all three in sequence.
Worked example
Your turn: compute K (RBF γ=0.5) and K̃ for X = [[0,0],[1,0],[0,1]]. Predict the grand mean before running.
Hint: K[i,j] = exp(−0.5·‖xᵢ−xⱼ‖²); centering: K̃[i,j] = K[i,j] − row_mean_i − col_mean_j + grand_mean.
import numpy as np
X = np.array([[0.,0.],[1.,0.],[0.,1.]])
gamma = 0.5
K = np.array([[np.exp(-gamma*np.sum((X[i]-X[j])**2))
for j in range(3)] for i in range(3)])
row_m = K.mean(1); col_m = K.mean(0); gm = K.mean()
K_c = K - row_m[:,None] - col_m[None,:] + gm
print('K:\n', K.round(4))
print('K_c:\n', K_c.round(4))| quantity | value |
|---|---|
| K[0,1] = K[0,2] | 0.6065 |
| K[1,2] | 0.3679 |
| grand mean | 0.6847 |
| K̃[1,1] = K̃[2,2] | 0.3684 |
| K̃[1,2] | −0.2637 |
Worked example
Your turn: eigendecompose K̃ from Milestone 1. Predict the number of non-zero eigenvalues.
Hint: np.linalg.eigh returns ascending order — flip to descending. Project: Z[:,i] = K_c @ v[:,i] / sqrt(λᵢ).
vals, vecs = np.linalg.eigh(K_c)
vals, vecs = vals[::-1], vecs[:,::-1] # descending
print('eigenvalues:', vals.round(4))
# Project onto top 2
Z = np.zeros((3, 2))
for i in range(2):
Z[:,i] = K_c @ vecs[:,i] / np.sqrt(vals[i])
print('Z (top-2 embedding):\n', Z.round(4))| quantity | value |
|---|---|
| λ₁ | 0.6321 |
| λ₂ | 0.3139 |
| λ₃ (zero — centering) | ≈ 0.0000 |
| Z[0] | [−0.4124, −0.5824] |
| Z[1] | [0.7275, −0.0000] |
Worked example
Your turn: apply all three methods to the Swiss roll (n=300). Predict which method unrolls it best.
Hint: make_swiss_roll(300, noise=0.1, random_state=42); StandardScaler; kPCA RBF γ=1; t-SNE perplexity=30.
from sklearn.datasets import make_swiss_roll
from sklearn.decomposition import PCA, KernelPCA
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler
X_sw, _ = make_swiss_roll(300, noise=0.1, random_state=42)
X_sc = StandardScaler().fit_transform(X_sw)
pca = PCA(2).fit_transform(X_sc)
kp = KernelPCA(2,'rbf',gamma=1.0).fit_transform(X_sc)
ts = TSNE(2,perplexity=30,learning_rate=200,
max_iter=1000,random_state=42,init='pca').fit_transform(X_sc)
print('PCA EVR sum:', PCA(2).fit(X_sc).explained_variance_ratio_.sum().round(4))| method | unrolls? | note |
|---|---|---|
| PCA | no — EVR=0.7282 | projects coil flat |
| kPCA RBF γ=1 | partially | local patches aligned |
| t-SNE perp=30 | yes — colour bands | neighbourhood structure preserved |
Trade off
Comparison matrix
From Milestone 3 — Swiss roll: PCA vs kPCA vs t-SNE: every row here is a choice with a cost. Fill the unrolls? column, then say which row you would actually pick and what you give up for it.
| method | unrolls? | note |
|---|---|---|
| PCA | no — EVR=0.7282 | projects coil flat |
| kPCA RBF γ=1 | partially | local patches aligned |
| t-SNE perp=30 | yes — colour bands | neighbourhood structure preserved |
Concept
import numpy as np
from sklearn.datasets import make_swiss_roll
from sklearn.decomposition import PCA, KernelPCA
from sklearn.manifold import TSNE
from sklearn.preprocessing import StandardScaler
# 1. Kernel PCA from scratch (toy)
X3 = np.array([[0.,0.],[1.,0.],[0.,1.]]); g=0.5
K = np.array([[np.exp(-g*np.sum((X3[i]-X3[j])**2)) for j in range(3)] for i in range(3)])
row_m=K.mean(1); col_m=K.mean(0); gm=K.mean()
K_c = K - row_m[:,None] - col_m[None,:] + gm
vals,vecs = np.linalg.eigh(K_c); vals,vecs=vals[::-1],vecs[:,::-1]
print('Eigenvalues:', vals.round(4))
# 2. Swiss roll comparison
X_sw,_ = make_swiss_roll(300, noise=0.1, random_state=42)
X_sc = StandardScaler().fit_transform(X_sw)
print('PCA EVR sum:', PCA(2).fit(X_sc).explained_variance_ratio_.sum().round(4))
ts = TSNE(2,perplexity=30,learning_rate=200,max_iter=1000,
random_state=42,init='pca').fit_transform(X_sc)
print('t-SNE KL:', ts.shape, '(KL logged during fit)')| output | value |
|---|---|
| 3-pt eigenvalues | [0.6321, 0.3139, ≈0] |
| Swiss PCA EVR sum | 0.7282 |
| t-SNE output shape | (300, 2) |
If your scratch kPCA eigenvalues match and t-SNE runs — you've built the full pipeline.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| 3-pt eigenvalues | [0.6321, 0.3139, ≈0] |
| Swiss PCA EVR sum | 0.7282 |
| t-SNE output shape | (300, 2) |
Concept
Out loud, slides closed: (1) walk through all four Kernel PCA steps without looking at notes, (2) explain why t-SNE cluster distances are meaningless, (3) give one reason UMAP is preferred for production pipelines.
Stretch (homework): implement kPCA new-point projection (the kernel vector centering formula); apply t-SNE with perplexity ∈ {5, 30, 50} to digits and explain what changes; analyse why you can't use t-SNE for new data but UMAP handles it.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Kernel PCA: PCA in implicit space · t-SNE: neighbourhood-preserving embedding · UMAP: graph-based with inductive mapping · Your turn: implement Kernel PCA. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| Kernel PCA | center K before eigh — or PCs are shifted |
| new-point projection | kPCA: kernel vector; t-SNE: impossible; UMAP: transform() |
| t-SNE perplexity | ≈ cluster size; inter-cluster distances are meaningless |
| UMAP vs t-SNE | UMAP is inductive; t-SNE is transductive |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.