USAAIO Lesson 28, from Week 10, fully worked. It builds feature maps out on XOR, then gives the kernel trick k(x,z)=φ(x)·φ(z), deriving the polynomial kernel (x·z+c)² term by term and matching it to an explicit six-dimensional φ. It applies Mercer's theorem - a kernel is valid exactly when its Gram matrix is PSD - to accept the RBF kernel and reject a fake one by its negative eigenvalue, then shows that the RBF kernel corresponds to an infinite-dimensional feature space, with the series verified numerically. It covers the composition rules for sums, products, and positive scalings, and closes with kernelized ridge regression, where the dual equals the primal, and an RBF SVM that cracks XOR. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.
Subject: Machine Learning · 111 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 28 · Week 10 (Kernel Methods)
Reach into a high- (even infinite-) dimensional feature space without ever going there. We derive the kernel trick k(x,z) = φ(x)·φ(z) from the feature map up, prove which similarities are legal kernels via a PSD Gram matrix, show the RBF kernel is an infinite feature space, and kernelize both ridge and an SVM — every algebra step and every printed number verified by real execution.
Objectives
φ(x) out in full and say exactly why forming it can be expensive or impossiblek(x,z) = (x·z + c)² and match it, entry by entry, to an explicit 6-dimensional φk is a valid kernel iff every Gram matrix is PSD — and reject a fake kernel by finding a negative eigenvalueexp(−‖x−z‖²/2σ²) as an infinite-dimensional feature space and verify the series numericallyφWarm-up
Discussion prompt
Before we open Lesson 28: Kernel Methods: without looking back, what was the main idea of Numerical Stability, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
IEEE-754 finite precision (overflow, underflow, cancellation), the log-sum-exp trick and numerically stable softmax, the condition number and ill-conditioned solves, and mixed-precision training. Build a stable softmax and log-sum-exp that survive inputs the naive versions overflow on.
Section
Part 1 of 8 — why lift the data
Concept
Four points, the classic XOR labeling. The two + points sit on one diagonal, the two − points on the other. We carry this dataset through the entire lesson.
| point | x₁ | x₂ | label y |
|---|---|---|---|
| A | 0 | 0 | 0 (−) |
| B | 0 | 1 | 1 (+) |
| C | 1 | 0 | 1 (+) |
| D | 1 | 1 | 0 (−) |
No straight line separates the + from the −: any line that puts B and C together also traps A or D with them. XOR is not linearly separable in the input plane.
Comparison
Comparison matrix
From The running example: XOR: refill the x₁ column from what you know. The rest of the table is as it appeared.
| point | x₁ | x₂ | label y |
|---|---|---|---|
| A | 0 | 0 | 0 (−) |
| B | 0 | 1 | 1 (+) |
| C | 1 | 0 | 1 (+) |
| D | 1 | 1 | 0 (−) |
Intuition
A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.
The trick every kernel method rests on: don't change the classifier, change the space. Map each point x to φ(x) in a higher-dimensional space, run the linear method there, and the decision boundary looks curved back in the original plane.
The coordinate that untangles XOR is the product x₁·x₂ — it alone tells the two diagonals apart. Let's build that map.
Counterexample
Discussion prompt
A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The coordinate that untangles XOR is the product x₁·x₂ — it alone tells the two diagonals apart. Let's build that map.
Picture it
Figure (svg): The four XOR points lifted so that D rises above the plane on a new axis, with a horizontal plane separating the two plus points from the two minus points.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
Picture the lift as lifting the plane into 3-D by the height x₁·x₂. The two − points (A at height 0, D at height 1) split apart in height from the two + points (both at height 0) — a flat sheet can now slide between them.
Intuition
Picture the lift as lifting the plane into 3-D by the height x₁·x₂. The two − points (A at height 0, D at height 1) split apart in height from the two + points (both at height 0) — a flat sheet can now slide between them.
Figure (svg): The four XOR points lifted so that D rises above the plane on a new axis, with a horizontal plane separating the two plus points from the two minus points.
The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.
Analogy
Discussion prompt
Explain Linear there = nonlinear here by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.
Concept
feature map φ — A function φ: ℝᵈ → ℝᴰ that sends each input x to a longer feature vector φ(x). D (the feature dimension) is usually much larger than d, and can be infinite.
For XOR, take the quadratic map that appends the cross term and squares:
\[ \varphi(x) = \big(1,\; x_1,\; x_2,\; x_1 x_2\big) \]
The new fourth coordinate x₁x₂ is 0 for A, B, C but 1 for D — exactly the signal a line in the lifted space can exploit.
Explain it
Discussion prompt
Explain A feature map φ to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
For XOR, take the quadratic map that appends the cross term and squares:
Estimation
Predict first
Apply φ(x) = (1, x₁, x₂, x₁x₂) to all four points, then check that the single rule x₁ + x₂ − 2x₁x₂ > 0.5 splits the labels:
Commit before you compute: what does φ makes XOR linearly separable come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Lifted scores are [0, 1, 1, 0]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. In the 4-D feature space the plane w=(0,1,1,−2) gives +1 to the two + points and 0 to the two − points — a clean linear split that was impossible in the 2-D input.
Worked example
Apply φ(x) = (1, x₁, x₂, x₁x₂) to all four points, then check that the single rule x₁ + x₂ − 2x₁x₂ > 0.5 splits the labels:
import numpy as np
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
def phi(x):
x1, x2 = x
return np.array([1.0, x1, x2, x1*x2])
Phi = np.array([phi(p) for p in X]) # (4,4) lifted points
w = np.array([0.0, 1.0, 1.0, -2.0]) # a separating plane
scores = Phi @ w
print(scores) # [0, 1, 1, 0]
print((scores > 0.5).astype(int)) # matches yLifted scores are [0, 1, 1, 0]
Why: In the 4-D feature space the plane w=(0,1,1,−2) gives +1 to the two + points and 0 to the two − points — a clean linear split that was impossible in the 2-D input.
| point | φ(x) = (1, x₁, x₂, x₁x₂) | score = φ·w | label |
|---|---|---|---|
| A (0,0) | (1, 0, 0, 0) | 0 | 0 |
| B (0,1) | (1, 0, 1, 0) | 1 | 1 |
| C (1,0) | (1, 1, 0, 0) | 1 | 1 |
| D (1,1) | (1, 1, 1, 1) | 0 | 0 |
Discrimination
Sort into buckets
Sort these by score = φ·w, from memory, without looking back at φ makes XOR linearly separable. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Lifting works, but the feature dimension grows fast. All degree-≤2 monomials of a d-dimensional input already number on the order of d²/2; degree p gives on the order of dᵖ.
\[ \text{degree } p,\ \dim d:\quad \dim \varphi \;=\; \binom{d+p}{p} \;\sim\; d^{\,p} \]
For images (d in the thousands) even p = 3 is astronomically large, and the RBF map is infinite-dimensional. Building φ(x) explicitly is hopeless. We need its dot products without ever forming it.
Section
Part 2 of 8 — dot products for free
Intuition
Here is the escape hatch. A whole family of algorithms — the SVM dual, ridge regression, PCA, nearest-mean — never look at a feature vector alone. They only ever combine features through inner products φ(xᵢ)·φ(xⱼ).
So we do not need φ(x) itself. We only need a fast way to compute the number φ(x)·φ(z) for any pair. If a cheap function of x and z returns that number directly, we can run the algorithm in the huge feature space at input-space cost.
Concept
kernel — A function k(x, z) that equals the inner product of two feature vectors, k(x,z) = φ(x)·φ(z), for some feature map φ — but is computed directly from x and z, without building φ.
\[ k(x, z) \;=\; \varphi(x) \cdot \varphi(z) \]
The kernel trick: replace every φ(xᵢ)·φ(xⱼ) in the algorithm by k(xᵢ, xⱼ). The feature map is used implicitly and never materialized. The rest of the lesson is (1) which k are legal and (2) what you can do with them.
Definition probe
Sort into buckets
Every line below is part of the definition of feature map φ or of kernel — one or the other, never both. Put each where it belongs.
Missing information
Discussion prompt
Take the simplest interesting kernel, k(x,z) = (x·z + c)² in 2-D, and find the φ hiding inside it. Start by writing the dot product out and squaring — nothing skipped.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
x = (x₁, x₂), z = (z₁, z₂), so the dot product is x₁z₁ + x₂z₂, plus the constant c.
Worked example
Take the simplest interesting kernel, k(x,z) = (x·z + c)² in 2-D, and find the φ hiding inside it. Start by writing the dot product out and squaring — nothing skipped.
Write x·z + c for 2-D vectors
Why: x = (x₁, x₂), z = (z₁, z₂), so the dot product is x₁z₁ + x₂z₂, plus the constant c.
\[ x\cdot z + c = x_1 z_1 + x_2 z_2 + c \]
Square the trinomial
Why: Use (p+q+r)² = p² + q² + r² + 2pq + 2pr + 2qr with p=x₁z₁, q=x₂z₂, r=c.
\[ (x_1 z_1 + x_2 z_2 + c)^2 \]
Notation
Annotate
From Derive the polynomial kernel — set up — read this one piece at a time. What is each part doing?
On: \( x\cdot z + c = x_1 z_1 + x_2 z_2 + c \)
Worked example
Expand every term of the square
Why: Square each of the three parts, then add twice each pairwise product.
\[ \begin{aligned} (x\cdot z + c)^2 =\;& x_1^2 z_1^2 + x_2^2 z_2^2 + c^2 \\ &+ 2\,x_1 z_1\, x_2 z_2 + 2c\,x_1 z_1 + 2c\,x_2 z_2 \end{aligned} \]
Regroup so each term is (something in x)·(same thing in z)
Why: Every term factors into an x-part times the matching z-part — that pairing is exactly what an inner product looks like.
\[ \begin{aligned} =\;& (x_1^2)(z_1^2) + (x_2^2)(z_2^2) + (\sqrt2\,x_1 x_2)(\sqrt2\,z_1 z_2) \\ &+ (\sqrt{2c}\,x_1)(\sqrt{2c}\,z_1) + (\sqrt{2c}\,x_2)(\sqrt{2c}\,z_2) + (c)(c) \end{aligned} \]
Translation
\( \begin{aligned} =\;& (x_1^2)(z_1^2) + (x_2^2)(z_2^2) + (\sqrt2\,x_1 x_2)(\sqrt2\,z_1 z_2) \\ &+ (\sqrt{2c}\,x_1)(\sqrt{2c}\,z_1) + (\sqrt{2c}\,x_2)(\sqrt{2c}\,z_2) + (c)(c) \end{aligned} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Fill the middle
Fill in the blanks
From Derive the polynomial kernel — read off φ — finish the line. Write what belongs on the right of the equals sign before you look.
\boxed\varphi(x)\cdot\varphi(z)\,}}
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Each parenthesized x-piece becomes a coordinate of the feature map.
Worked example
Collect the six x-parts into one vector φ(x)
Why: Each parenthesized x-piece becomes a coordinate of the feature map. The √2 factors are what make the cross terms line up as a clean dot product.
\[ \varphi(x) = \big(x_1^2,\; x_2^2,\; \sqrt2\,x_1 x_2,\; \sqrt{2c}\,x_1,\; \sqrt{2c}\,x_2,\; c\big) \]
Therefore (x·z + c)² = φ(x)·φ(z)
Why: The kernel — one dot product plus a square, computed on the raw 2-vectors — equals the dot product of these 6-dimensional feature vectors. Left side: O(d) work. Right side: O(D) work. Same number.
\[ \boxed{\,(x\cdot z + c)^2 = \varphi(x)\cdot\varphi(z)\,} \]
The kernel gave us the entire quadratic feature space — squares, the cross term, and scaled linears — for the price of a dot product. Now verify it on real numbers.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Therefore (x·z + c)² = φ(x)·φ(z)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The kernel gave us the entire quadratic feature space — squares, the cross term, and scaled linears — for the price of a dot product. Now verify it on real numbers.
Pattern
Predict first
The table runs: a·b | 1.0 · kernel (a·b + c)² | 4.0 · φ(a)·φ(b) explicit | 3.9999999999999987
In Verify the polynomial kernel in code, given the rows so far: what is the next one — the row where quantity is np.isclose(...)?
Correct: np.isclose(...) | True
| quantity | value (verified) |
|---|---|
| a·b | 1.0 |
| kernel (a·b + c)² | 4.0 |
| φ(a)·φ(b) explicit | 3.9999999999999987 |
| np.isclose(...) | True |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The cheap kernel gives exactly 4.0.
Worked example
Take a = (1, 2), b = (3, −1), c = 1. Compute the kernel the cheap way and the explicit-feature way; they must agree. Runnable as written:
import numpy as np
def phi(x, c=1.0):
x1, x2 = x
return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.]); c = 1.0
k_direct = (a @ b + c)**2 # 1 dot + 1 square
k_feat = phi(a, c) @ phi(b, c) # 6-dim dot product
print(k_direct) # 4.0
print(k_feat) # 3.9999999999999987
print(np.isclose(k_direct, k_feat))a·b = 1(3) + 2(−1) = 1, so (a·b + c)² = (1 + 1)² = 4
Why: The cheap kernel gives exactly 4.0. The explicit map gives 3.9999999999999987 — identical up to floating-point round-off, so np.isclose returns True.
| quantity | value (verified) |
|---|---|
| a·b | 1.0 |
| kernel (a·b + c)² | 4.0 |
| φ(a)·φ(b) explicit | 3.9999999999999987 |
| np.isclose(...) | True |
Trade off
Comparison matrix
From Verify the polynomial kernel in code: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | value (verified) |
|---|---|
| a·b | 1.0 |
| kernel (a·b + c)² | 4.0 |
| φ(a)·φ(b) explicit | 3.9999999999999987 |
| np.isclose(...) | True |
Worked example
To see there is no magic, print each feature vector and multiply component-wise. The six products must sum to the kernel value 4.
import numpy as np
def phi(x, c=1.0):
x1, x2 = x
return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
pa, pb = phi(a), phi(b)
print(np.round(pa, 4)) # phi(a)
print(np.round(pb, 4)) # phi(b)
print(np.round(pa * pb, 4)) # elementwise products
print(np.sum(pa * pb)) # 3.9999999999999982 (== 4 up to round-off)Elementwise products are [9, 4, −12, 6, −4, 1]
Why: Sum: 9 + 4 − 12 + 6 − 4 + 1 = 4. Each coordinate contributes; the cross term (−12) and scaled-linear terms (6, −4) are exactly the parts a plain (x·z) miss.
| coord | φ(a) | φ(b) | product |
|---|---|---|---|
| x₁² | 1.0 | 9.0 | 9.0 |
| x₂² | 4.0 | 1.0 | 4.0 |
| √2·x₁x₂ | 2.8284 | −4.2426 | −12.0 |
| √2·x₁ | 1.4142 | 4.2426 | 6.0 |
| √2·x₂ | 2.8284 | −1.4142 | −4.0 |
| c | 1.0 | 1.0 | 1.0 |
Pattern
Step through it
Step through Term-by-term: the dot product lands on 4 one row at a time. What is driving the change, and what would the row after the last one be?
Concept
The degree-2 case generalizes. The kernel k(x,z) = (x·z + c)ᵖ corresponds to the feature map of all monomials up to degree p, with the constant c balancing lower- against higher-degree terms.
\[ k(x,z) = (x\cdot z + c)^p \;=\; \varphi_p(x)\cdot \varphi_p(z),\qquad \dim \varphi_p = \binom{d+p}{p} \]
For d = 2, p = 2 that is C(4,2) = 6 features — exactly the 6-vector we just derived. But the kernel is always a single dot-plus-power on the raw inputs: the feature count never enters the computation.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The feature map has 6 dimensions, so computing k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.
The kernel is computed on the raw inputs: one d-dimensional dot product x·z, add c, square. Cost scales with the input dimension d, not the feature dimension D.
Why: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.
Trap
The feature map has 6 dimensions, so computing k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.
Cost of k ≈ dim φ = 6 (grows with the feature space)
Why: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.
The kernel is computed on the raw inputs: one d-dimensional dot product x·z, add c, square. Cost scales with the input dimension d, not the feature dimension D.
Cost of k = O(d) — independent of dim φ
Why: (x·z + c)² needs 2 multiplies + 1 add + 1 square regardless of whether φ is 6-dimensional or infinite. That O(d)-not-O(D) gap IS the kernel trick; it is why RBF's infinite φ is free.
Section
Part 3 of 8 — which k are legal
Concept
You could invent any symmetric similarity(x, z) and hope it behaves like a kernel. But k is only legal if some feature map φ actually produces it — otherwise k(x,z) = φ(x)·φ(z) is a lie and the algorithms built on it lose their guarantees.
Mercer's theorem gives a checkable condition, no φ required: look at the Gram matrix of k on any dataset.
Concept
Gram matrix — For points x₁,…,xₙ, the n×n matrix K with Kᵢⱼ = k(xᵢ, xⱼ) — every pairwise kernel value.
Mercer's theorem: k is a valid kernel iff its Gram matrix K is positive semidefinite (PSD) for every finite dataset:
\[ K_{ij} = k(x_i, x_j),\qquad z^\top K z \ge 0 \ \text{ for all } z \iff K \succeq 0 \]
Why PSD is exactly right: if K = ΦΦᵀ (features stacked in rows of Φ), then zᵀKz = ‖Φᵀz‖² ≥ 0 automatically. PSD is precisely the fingerprint of 'these numbers came from real inner products.' In practice: check min eigenvalue ≥ 0 (Lesson 13).
Worked example
One direction of Mercer is a two-line proof. If k(xᵢ,xⱼ) = φ(xᵢ)·φ(xⱼ), stack the features as rows of Φ; then K = ΦΦᵀ. Show zᵀKz ≥ 0 for any z:
Write the quadratic form with K = ΦΦᵀ
Why: Substitute K = ΦΦᵀ into zᵀKz. Matrix multiplication is associative, so we can regroup the product.
\[ z^\top K z = z^\top \Phi \Phi^\top z = (\Phi^\top z)^\top (\Phi^\top z) \]
That is a squared norm, hence ≥ 0
Why: Any vector dotted with itself is its squared length, which can never be negative. So every Gram matrix of real inner products is PSD — no assumption on the data needed.
\[ z^\top K z = \lVert \Phi^\top z\rVert^2 \;\ge\; 0 \quad\Longrightarrow\quad K \succeq 0 \]
| step | expression | why |
|---|---|---|
| substitute | zᵀΦΦᵀz | K = ΦΦᵀ |
| regroup | (Φᵀz)ᵀ(Φᵀz) | associativity |
| conclude | ‖Φᵀz‖² ≥ 0 | norms are non-negative |
Notation
Annotate
From Why a real kernel's Gram is always PSD — read this one piece at a time. What is each part doing?
On: \( z^\top K z = \lVert \Phi^\top z\rVert^2 \;\ge\; 0 \quad\Longrightarrow\quad K \succeq 0 \)
Fill the middle
Fill in the blanks
From The RBF Gram matrix is PSD — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X2, 1)[:, None]
+ np.sum(X2, 1)[None, :]
- 2 * X @ X.T) # ||xi - xj||^2 for all pairs
K = np.exp(-sq / 2.0) # RBF, sigma = 1
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4)) # 0.0263 >= 0
Why: sq is what everything below it consumes, so the wrong expression here fails later and somewhere else. Every eigenvalue is positive, so K is PSD on this dataset — consistent with Mercer.
Worked example
The RBF (Gaussian) kernel k(x,z) = exp(−‖x−z‖²/2σ²) is the workhorse. Build its Gram matrix on 6 random points (seeded) and confirm the smallest eigenvalue is ≥ 0:
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X**2, 1)[:, None]
+ np.sum(X**2, 1)[None, :]
- 2 * X @ X.T) # ||xi - xj||^2 for all pairs
K = np.exp(-sq / 2.0) # RBF, sigma = 1
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4)) # 0.0263 >= 0min eigenvalue = 0.0263 ≥ 0 → Gram is PSD → valid kernel
Why: Every eigenvalue is positive, so K is PSD on this dataset — consistent with Mercer. The diagonal is all 1s because exp(−0/2σ²)=1 (a point is maximally similar to itself).
| eigenvalue # | value |
|---|---|
| 1 (smallest) | 0.0263 |
| 2 | 0.0571 |
| 3 | 0.4790 |
| 6 (largest) | 3.3872 |
Pattern
Step through it
Step through The RBF Gram matrix is PSD one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Test the plausible-looking candidate k(x,z) = x·z − ‖x−z‖². Build its Gram matrix on the same 6 points and look for a negative eigenvalue:
Commit before you compute: what does A fake kernel fails Mercer come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: min eigenvalue = −14.1257 < 0 → NOT PSD → NOT a valid kernel
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. One negative eigenvalue is a counterexample to Mercer's condition on this very dataset, so no feature map φ could ever produce this k.
Worked example
Test the plausible-looking candidate k(x,z) = x·z − ‖x−z‖². Build its Gram matrix on the same 6 points and look for a negative eigenvalue:
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X**2, 1)[:, None]
+ np.sum(X**2, 1)[None, :]
- 2 * X @ X.T)
K = X @ X.T - sq # candidate k = x.z - ||x-z||^2
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4)) # -14.1257 < 0min eigenvalue = −14.1257 < 0 → NOT PSD → NOT a valid kernel
Why: One negative eigenvalue is a counterexample to Mercer's condition on this very dataset, so no feature map φ could ever produce this k. It only takes a single negative eigenvalue to disqualify a candidate.
| check | value |
|---|---|
| eigenvalues (rounded) | [−14.13, 0, 0, 1.54, 3.66, 14.88] |
| min eigenvalue | −14.1257 |
| PSD? | no |
| valid kernel? | no |
Faded example
Fill in the blanks
The linear kernel passes (the honest baseline), with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
K = X @ X.T # linear-kernel Gram = X X^T
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 6)) # -0.0 (round-off), i.e. >= 0
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. K = XXᵀ is PSD by the ΦΦᵀ proof.
Worked example
For contrast, the linear kernel k(x,z) = x·z has Gram K = XXᵀ — literally ΦΦᵀ with φ = identity. Its min eigenvalue should be ≥ 0 (allowing tiny negative round-off) on the same 6 points:
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
K = X @ X.T # linear-kernel Gram = X X^T
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 6)) # -0.0 (round-off), i.e. >= 0min eigenvalue = −0.0 (numerically 0) → PSD → valid
Why: K = XXᵀ is PSD by the ΦΦᵀ proof. With only 2 columns in X, K has rank 2, so four eigenvalues are ~0 (the tiny −0.0 is floating-point noise, not a real negative) and two are positive. The linear kernel is always valid.
| check | value |
|---|---|
| eigenvalues (rounded) | [−0, −0, 0, 0, 0.9956, 4.9605] |
| min eigenvalue | −0.0 (≈ 0) |
| PSD? / valid kernel? | yes / yes |
Anomaly
Predict first
A student writes this, and it looks reasonable:
k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Symmetry is necessary but NOT sufficient.
Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.
Why: Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.
Trap
k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.
Feed an ad-hoc symmetric similarity to the solver
Why: Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.
Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.
Verify Gram PSD (or use a proven kernel) before fitting
Why: The candidate above has min eigenvalue −14.13, so it is rejected. RBF (min eig 0.0263) and polynomial pass. One eigvalsh call is cheap insurance against a silently broken model.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.
Section
Part 4 of 8 — the Gaussian kernel
Intuition
The RBF kernel exp(−‖x−z‖²/2σ²) reads as pure geometry: it is 1 when x = z (zero distance) and decays smoothly toward 0 as the points move apart. It is a similarity that only cares about distance.
The bandwidth σ sets the reach: small σ means similarity drops off fast (each point influences only its close neighbors — a wiggly boundary); large σ means slow decay (far points still count — a smooth boundary).
Worked example
Compute the RBF between our running pair a=(1,2), b=(3,-1) (far apart) and between two nearby points, at two bandwidths:
import numpy as np
def rbf(x, z, sigma=1.0):
return np.exp(-np.sum((x - z)**2) / (2 * sigma**2))
a = np.array([1., 2.]); b = np.array([3., -1.])
p = np.array([1., 2.]); q = np.array([1.1, 2.1])
print(np.sum((a - b)**2)) # 13.0 (far)
print(rbf(a, a)) # 1.0 (self)
print(round(rbf(a, b, 1.0), 6)) # 0.001503
print(round(rbf(a, b, 3.0), 6)) # 0.485672
print(round(rbf(p, q, 1.0), 6)) # 0.99005 (near)Far pair: σ=1 → 0.0015 (almost dissimilar); σ=3 → 0.4857
Why: ‖a−b‖²=13. At σ=1 that distance kills the similarity; widening to σ=3 lets the same pair count as ~half-similar. A near pair (distance 0.02) stays at 0.990 either way.
| pair | ‖·‖² | σ | RBF value |
|---|---|---|---|
| a with itself | 0 | 1 | 1.0 |
| a, b (far) | 13 | 1 | 0.001503 |
| a, b (far) | 13 | 3 | 0.485672 |
| p, q (near) | 0.02 | 1 | 0.99005 |
Scale up
Step through it
Step through RBF values: near vs far and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Concept
sklearn writes the RBF kernel with a single parameter gamma instead of the bandwidth σ. They are the same dial, reciprocally related:
\[ \exp\!\Big(-\frac{\lVert x-z\rVert^2}{2\sigma^2}\Big) = \exp\big(-\gamma\,\lVert x-z\rVert^2\big),\qquad \gamma = \frac{1}{2\sigma^2} \]
So large γ = small σ = fast decay = wiggly, local boundary (risk of overfitting); small γ = large σ = smooth, global boundary (risk of underfitting). When we pass gamma=1.0 to SVC shortly, that is σ = 1/√2 — a tight, local kernel, which is exactly what four well-separated XOR points want.
Intuition
Think of each training point as placing a soft Gaussian bump of influence around itself, of width σ. The decision boundary threads between the summed bumps of the two classes.
Wide bumps (large σ, small γ) overlap and blur into one smooth boundary — simple, may underfit. Narrow bumps (small σ, large γ) stay isolated islands around each point — flexible, may memorize noise (overfit). Tuning σ/γ on validation data is the RBF SVM's core bias–variance knob, the twin of λ in ridge.
Concept
Expand the exponential of an inner product as its power series. Because exp(t) = Σ tⁿ/n!, the kernel exp(x·z) splits into a sum over all polynomial degrees at once:
\[ \exp(x\cdot z) = \sum_{n=0}^{\infty} \frac{(x\cdot z)^n}{n!} \]
Each term (x·z)ⁿ/n! is itself a polynomial kernel of degree n, carrying its own block of monomial features. Summing over n = 0, 1, 2, … stacks infinitely many feature coordinates. RBF is exp(x·z) dressed with normalizing factors, so it too has an infinite φ — one you could never write down, yet dot in O(d).
Fill the middle
Fill in the blanks
From Verify the infinite series converges — one line has had its right-hand side removed. Put it back.
import numpy as np
from math import factorial
def feat(x, N):
return np.array([xn / np.sqrt(factorial(n)) for n in range(N)])
x, z = 0.7, 0.4
for N in [3, 6, 12]:
approx = feat(x, N) @ feat(z, N)**
print(N, round(approx, 8), round(np.exp(x*z), 8))
Why: approx is what everything below it consumes, so the wrong expression here fails later and somewhere else. With only 3 features the dot is 1.31920 (off in the 3rd decimal); 6 features give 1.32312912; 12 features match exp(0.28)=1.32312981 exactly.
Worked example
In 1-D the feature map is φ(x)ₙ = xⁿ/√(n!), so φ(x)·φ(z) = Σ (xz)ⁿ/n! = exp(xz). Truncate at N terms and watch it converge to the true kernel for x=0.7, z=0.4:
import numpy as np
from math import factorial
def feat(x, N):
return np.array([x**n / np.sqrt(factorial(n)) for n in range(N)])
x, z = 0.7, 0.4
for N in [3, 6, 12]:
approx = feat(x, N) @ feat(z, N)
print(N, round(approx, 8), round(np.exp(x*z), 8))Truncated dot product → exp(xz) = 1.32312981 as N grows
Why: With only 3 features the dot is 1.31920 (off in the 3rd decimal); 6 features give 1.32312912; 12 features match exp(0.28)=1.32312981 exactly. The infinite feature vector is real — we just don't need all of it.
| N (features kept) | feat·feat | exp(xz) |
|---|---|---|
| 3 | 1.31920000 | 1.32312981 |
| 6 | 1.32312912 | 1.32312981 |
| 12 | 1.32312981 | 1.32312981 |
Invariant
Step through it
Step through Verify the infinite series converges one row at a time. One of these columns never changes — find it, and say why it cannot.
Section
Part 5 of 8 — build new from old
Concept
Valid kernels are closed under a few operations, so you can assemble complex kernels from simple proven ones and stay inside Mercer's guarantee:
k₁ + k₂ is a kernel (PSD + PSD = PSD)k₁ · k₂ is a kernel (elementwise product of PSD Grams stays PSD — the Schur product theorem)c·k for c > 0 is a kernelNot guaranteed: a difference k₁ − k₂ — subtracting PSD matrices can create a negative eigenvalue. Let's check all four on the Gram matrices.
Intuition
The rules let you design a kernel to match your data's structure without ever proving PSD from scratch. Start from proven blocks and glue them:
\[ k(x,z) = \underbrace{x\cdot z}_{\text{linear trend}} \;+\; \underbrace{\exp\!\big(-\gamma\lVert x-z\rVert^2\big)}_{\text{local bumps}} \]
This sum captures a broad linear trend and fine local structure at once — and because sum-of-kernels is a kernel, it is automatically valid. Products let you combine kernels on different feature groups (e.g. an RBF on the numeric columns times a kernel on the categorical ones).
Fill the middle
Fill in the blanks
From Composition preserves PSD — one line has had its right-hand side removed. Put it back.
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X2,1)[:,None] + np.sum(X2,1)[None,:] - 2*X@X.T
K_rbf = np.exp(-sq/2.0)
K_lin = X @ X.T
m = lambda M: round(np.linalg.eigvalsh(M).min(), 6)
print(m(K_lin + K_rbf)) # 0.027333 sum
print(m(K_lin * K_rbf)) # 0.012704 product (elementwise)
print(m(3.0 * K_rbf)) # 0.079021 scaling
print(m(K_lin - K_rbf)) # -3.377703 difference: BREAKS PSD
Why: K_rbf is what everything below it consumes, so the wrong expression here fails later and somewhere else. Sum 0.0273, product 0.0127, 3×RBF 0.0790 are all ≥ 0 — still valid kernels.
Worked example
Take a linear Gram and an RBF Gram on the same 6 points. Check the min eigenvalue of the sum, product, a positive scaling, and the (illegal) difference:
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X**2,1)[:,None] + np.sum(X**2,1)[None,:] - 2*X@X.T
K_rbf = np.exp(-sq/2.0)
K_lin = X @ X.T
m = lambda M: round(np.linalg.eigvalsh(M).min(), 6)
print(m(K_lin + K_rbf)) # 0.027333 sum
print(m(K_lin * K_rbf)) # 0.012704 product (elementwise)
print(m(3.0 * K_rbf)) # 0.079021 scaling
print(m(K_lin - K_rbf)) # -3.377703 difference: BREAKS PSDSum, product, scaling stay PSD; the difference goes negative
Why: Sum 0.0273, product 0.0127, 3×RBF 0.0790 are all ≥ 0 — still valid kernels. The difference K_lin − K_rbf hits −3.3777, so it is NOT a kernel. This is why the rules list sum/product/scaling but never subtraction.
| combination | min eigenvalue | valid kernel? |
|---|---|---|
| k_lin + k_rbf | 0.027333 | yes |
| k_lin · k_rbf | 0.012704 | yes |
| 3 · k_rbf | 0.079021 | yes |
| k_lin − k_rbf | −3.377703 | no |
Comparison
Comparison matrix
From Composition preserves PSD: refill the min eigenvalue column from what you know. The rest of the table is as it appeared.
| combination | min eigenvalue | valid kernel? |
|---|---|---|
| k_lin + k_rbf | 0.027333 | yes |
| k_lin · k_rbf | 0.012704 | yes |
| 3 · k_rbf | 0.079021 | yes |
| k_lin − k_rbf | −3.377703 | no |
Section
Part 6 of 8 — the dual view
Concept
Kernelizing an algorithm means rewriting it so features appear only inside dot products, which then become kernel calls. For ridge regression the key fact (the representer theorem) is that the optimal weight vector is a combination of the training features:
\[ w^\star = \sum_{i} \alpha_i\, \varphi(x_i) = \Phi^\top \alpha \]
Substituting this back turns the problem from solving for w in feature space into solving for the n dual coefficients α — and every φ cancels into a Gram matrix K = ΦΦᵀ of kernel values.
Concept
Primal ridge solves a D × D system (D = feature dimension); the kernel dual solves an n × n system (n = number of training points). You pick whichever matrix is smaller.
| view | solve | cost driver | best when |
|---|---|---|---|
| primal | (ΦᵀΦ + λI) w = Φᵀy | D × D | few features (D ≪ n) |
| dual (kernel) | (K + λI) α = y | n × n | huge/∞ features (D ≫ n) |
This is why kernels shine with the RBF map: D = ∞, so the primal is impossible, but the dual is just an n × n solve. The feature dimension has been traded away for the dataset size.
Pattern
Predict first
The table runs: 0 | 0.55081 | 0.55081 · 1 | 1.952872 | 1.952872 · 2 | 5.063328 | 5.063328
In Primal ridge vs the kernel dual, given the rows so far: what is the next one — the row where point x is 3?
Correct: 3 | 9.88218 | 9.88218
| point x | primal ŷ | dual ŷ |
|---|---|---|
| 0 | 0.55081 | 0.55081 |
| 1 | 1.952872 | 1.952872 |
| 2 | 5.063328 | 5.063328 |
| 3 | 9.88218 | 9.88218 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Both give [0.55081, 1.952872, 5.063328, 9.88218] and np.allclose returns True.
Worked example
Primal ridge in explicit features solves (ΦᵀΦ + λI)w = Φᵀy. The dual solves (K + λI)α = y with K = ΦΦᵀ, then predicts ŷ = Kα. They give identical predictions — verify on 4 one-dimensional points with the degree-2 kernel k(x,z)=(xz+1)²:
import numpy as np
X = np.array([[0.],[1.],[2.],[3.]]); y = np.array([1., 2., 5., 10.])
lam = 1.0
def phi(x):
x = x[0]
return np.array([1.0, np.sqrt(2)*x, x**2])
Phi = np.array([phi(xi) for xi in X]) # explicit features
w = np.linalg.solve(Phi.T@Phi + lam*np.eye(3), Phi.T@y)
pred_primal = Phi @ w
K = (X @ X.T + 1.0)**2 # Gram, no phi
alpha = np.linalg.solve(K + lam*np.eye(4), y) # dual solve
pred_dual = K @ alpha
print(np.round(pred_primal, 6))
print(np.round(pred_dual, 6))
print(np.allclose(pred_primal, pred_dual)) # TruePrimal and dual predictions match to machine precision
Why: Both give [0.55081, 1.952872, 5.063328, 9.88218] and np.allclose returns True. The dual never builds the 3-dim φ beyond checking — it works entirely through the 4×4 kernel matrix K.
| point x | primal ŷ | dual ŷ |
|---|---|---|
| 0 | 0.55081 | 0.55081 |
| 1 | 1.952872 | 1.952872 |
| 2 | 5.063328 | 5.063328 |
| 3 | 9.88218 | 9.88218 |
Pattern
Step through it
Step through Primal ridge vs the kernel dual one row at a time. What is driving the change, and what would the row after the last one be?
Concept
The SVM's dual objective (Lesson 24's Lagrangian) contains data only as inner products φ(xᵢ)·φ(xⱼ). Replace each by k(xᵢ, xⱼ) and the SVM runs in the feature space with a decision function that is a weighted sum of kernels:
\[ f(x) = \operatorname{sign}\!\Big(\sum_{i} \alpha_i\, y_i\, k(x_i, x) + b\Big) \]
Only the support vectors have αᵢ ≠ 0. With an RBF kernel this becomes a smooth nonlinear boundary — enough to finally crack XOR.
Explain it
Discussion prompt
Explain Kernelizing an SVM to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Only the support vectors have αᵢ ≠ 0. With an RBF kernel this becomes a smooth nonlinear boundary — enough to finally crack XOR.
Missing information
Discussion prompt
Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The linear kernel = no feature lift = a straight boundary, stuck at chance on XOR. The RBF kernel implicitly lifts into an infinite space where XOR is separable — the kernel trick doing exactly what the explicit φ = (1,x₁,x₂,x₁x₂) did in Part 1, but for free.
Worked example
Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
lin = SVC(kernel='linear').fit(X, y).score(X, y)
rbf = SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y).score(X, y)
print(lin) # 0.5 (fails)
print(rbf) # 1.0 (perfect)linear accuracy 0.5 (fails), RBF accuracy 1.0 (perfect)
Why: The linear kernel = no feature lift = a straight boundary, stuck at chance on XOR. The RBF kernel implicitly lifts into an infinite space where XOR is separable — the kernel trick doing exactly what the explicit φ = (1,x₁,x₂,x₁x₂) did in Part 1, but for free.
| kernel | XOR accuracy | boundary |
|---|---|---|
| linear | 0.5 (fails) | straight line |
| RBF | 1.0 | curved (nonlinear) |
| poly deg-2 | 1.0 | curved (nonlinear) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
linear accuracy 0.5 (fails), RBF accuracy 1.0 (perfect)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:
Estimation
Predict first
Look at what the RBF SVM actually decided. Print its per-point predictions and how many support vectors it kept — for four tightly-packed points, expect all of them:
Commit before you compute: what does Inside the RBF fit: every point is a support vector come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: predictions [0 1 1 0] match the labels; 4 support vectors
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every prediction is correct. All four points are support vectors (α ≠ 0) because with only four points and a tight kernel, each one is needed to shape the boundary.
Worked example
Look at what the RBF SVM actually decided. Print its per-point predictions and how many support vectors it kept — for four tightly-packed points, expect all of them:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
clf = SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y)
print(clf.predict(X)) # [0 1 1 0]
print(y) # [0 1 1 0]
print(int(clf.n_support_.sum())) # 4predictions [0 1 1 0] match the labels; 4 support vectors
Why: Every prediction is correct. All four points are support vectors (α ≠ 0) because with only four points and a tight kernel, each one is needed to shape the boundary. The decision function is Σ αᵢ yᵢ k(xᵢ, x) + b — pure kernel evaluations.
| point | true y | RBF prediction |
|---|---|---|
| A (0,0) | 0 | 0 |
| B (0,1) | 1 | 1 |
| C (1,0) | 1 | 1 |
| D (1,1) | 0 | 0 |
Discrimination
Sort into buckets
Sort these by true y, from memory, without looking back at Inside the RBF fit: every point is a support…. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
An SVM is a powerful margin classifier, so a linear-kernel SVM will surely handle XOR.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A linear kernel does NO feature lift — its decision boundary is a single straight line.
Match the kernel to the data: use RBF or polynomial whenever the structure is nonlinear.
Why: A linear kernel does NO feature lift — its decision boundary is a single straight line. XOR is not linearly separable, so the SVM is stuck at 50% no matter how you tune C. The power of an SVM is the margin, not nonlinearity.
Trap
An SVM is a powerful margin classifier, so a linear-kernel SVM will surely handle XOR.
SVC(kernel='linear') on XOR → 0.5
Why: A linear kernel does NO feature lift — its decision boundary is a single straight line. XOR is not linearly separable, so the SVM is stuck at 50% no matter how you tune C. The power of an SVM is the margin, not nonlinearity.
Match the kernel to the data: use RBF or polynomial whenever the structure is nonlinear.
SVC(kernel='rbf') on XOR → 1.0
Why: The nonlinear kernel supplies the implicit feature lift that makes XOR separable — 100% accuracy. Choosing the kernel is where you inject the nonlinearity; the SVM machinery around it is unchanged.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
+ points sit on one diagonal, the two − points on the other. We carry this dataset through the entire lesson.; A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.; The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.; k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.Concept
The trick generalizes far beyond the SVM. Anywhere an algorithm touches data only through dot products, you can kernelize it:
Master the three moves — dot-products-only, Mercer-check, solve-the-dual — and every one of these is the same idea in new clothes.
Analogy
Discussion prompt
Explain Where kernels go from here by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The trick generalizes far beyond the SVM. Anywhere an algorithm touches data only through dot products, you can kernelize it:
Section
Part 7 of 8 — pattern & checks
Ranking
Put in order
These are the steps of The kernel-method recipe, scrambled. Put them back in order before the next slide shows you.
φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)(x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reachK must be PSD (min eig ≥ 0) — verify, or compose from proven kernels(K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)(x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reachK must be PSD (min eig ≥ 0) — verify, or compose from proven kernels(K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λEdge cases
Discussion prompt
The kernel-method recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)(x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reachK must be PSD (min eig ≥ 0) — verify, or compose from proven kernels(K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λElimination
Eliminate the wrong options
The kernel trick lets an algorithm work in a feature space φ without ever forming φ(x). It does this by:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: A kernel returns the exact inner product φ(x)·φ(z) as a cheap function of the raw inputs. Algorithms that use features only through dot products (SVM, ridge, PCA) then run in the feature space at O(d) cost per pair.
Check
What is a kernel actually computing?
Check your understanding
The kernel trick lets an algorithm work in a feature space φ without ever forming φ(x). It does this by:
Answer: A
Why: A kernel returns the exact inner product φ(x)·φ(z) as a cheap function of the raw inputs. Algorithms that use features only through dot products (SVM, ridge, PCA) then run in the feature space at O(d) cost per pair.
Prediction
Predict first
By Mercer's theorem, k(x,z) is a valid kernel if and only if:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: its Gram matrix is positive semidefinite for every finite dataset
Why: Validity means k = φ(x)·φ(z) for some φ, which holds exactly when every Gram matrix Kᵢⱼ = k(xᵢ,xⱼ) is PSD (min eigenvalue ≥ 0). That is Mercer's theorem.
Check
When is a candidate k a legal kernel?
Check your understanding
By Mercer's theorem, k(x,z) is a valid kernel if and only if:
Answer: A
Why: Validity means k = φ(x)·φ(z) for some φ, which holds exactly when every Gram matrix Kᵢⱼ = k(xᵢ,xⱼ) is PSD (min eigenvalue ≥ 0). That is Mercer's theorem.
Prediction
Predict first
If k₁ and k₂ are valid kernels, which of these is guaranteed to be a valid kernel?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: k₁ + k₂ (and also k₁·k₂, and c·k₁ for c > 0)
Why: Sums, elementwise products, and positive scalings of valid kernels keep the Gram matrix PSD, so they stay valid. This is what lets you compose kernels freely.
Check
Which combinations are guaranteed to stay valid?
Check your understanding
If k₁ and k₂ are valid kernels, which of these is guaranteed to be a valid kernel?
Answer: A
Why: Sums, elementwise products, and positive scalings of valid kernels keep the Gram matrix PSD, so they stay valid. This is what lets you compose kernels freely.
Elimination
Eliminate the wrong options
You want RBF-kernel ridge regression on n = 500 points. Why solve the dual (K + λI)α = y instead of the primal (ΦᵀΦ + λI)w = Φᵀy?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: RBF's φ is infinite-dimensional, so ΦᵀΦ (D×D) is infinite and unbuildable. The dual replaces every φ(xᵢ)·φ(xⱼ) with k(xᵢ,xⱼ), giving a finite n×n Gram matrix K — a 500×500 solve. Kernelizing trades feature dimension for dataset size.
Check
Think about the RBF feature dimension.
Check your understanding
You want RBF-kernel ridge regression on n = 500 points. Why solve the dual (K + λI)α = y instead of the primal (ΦᵀΦ + λI)w = Φᵀy?
Answer: A
Why: RBF's φ is infinite-dimensional, so ΦᵀΦ (D×D) is infinite and unbuildable. The dual replaces every φ(xᵢ)·φ(xⱼ) with k(xᵢ,xⱼ), giving a finite n×n Gram matrix K — a 500×500 solve. Kernelizing trades feature dimension for dataset size.
Section
Part 8 of 8 — the project
Concept
Three milestones, each proving one pillar of the lesson yourself: the polynomial kernel equals its features, the RBF Gram is PSD, and the RBF SVM separates XOR while a linear one can't.
| # | requirement | tool |
|---|---|---|
| 1 | (x·z+c)² == φ(x)·φ(z) | explicit φ, np.isclose |
| 2 | RBF Gram matrix is PSD | np.linalg.eigvalsh |
| 3 | RBF SVM solves XOR, linear fails | sklearn SVC |
Build rules: type every line yourself, run after each milestone, and when the Gram's min eigenvalue looks negative, re-check the sign convention in your distance formula before blaming the kernel.
Counterexample
Discussion prompt
Three milestones, each proving one pillar of the lesson yourself: the polynomial kernel equals its features, the RBF Gram is PSD, and the RBF SVM separates XOR while a linear one can't.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, run after each milestone, and when the Gram's min eigenvalue looks negative, re-check the sign convention in your distance formula before blaming the kernel.
Worked example
Your turn: confirm (a·b+1)² equals φ(a)·φ(b) for the quadratic feature map. Predict out loud: exact match, or off by round-off?
Hint: φ(x) = [x₁², x₂², √2·x₁x₂, √2c·x₁, √2c·x₂, c]. Use np.isclose, not ==, because √2 introduces floating-point dust.
import numpy as np
def phi(x, c=1.0):
x1, x2 = x
return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
print((a@b+1)**2, phi(a)@phi(b))| computation | value |
|---|---|
| (a·b+1)² | 4.0 |
| φ(a)·φ(b) | 3.9999999999999987 |
| np.isclose | True |
Worked example
Your turn: build an RBF Gram matrix on 6 seeded points and confirm its minimum eigenvalue is ≥ 0. Predict: valid kernel?
Hint: squared distances via ‖x‖² + ‖z‖² − 2x·z, then exp(−sq/2σ²); read np.linalg.eigvalsh(K).min().
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X**2,1)[:,None] + np.sum(X**2,1)[None,:] - 2*X@X.T
K = np.exp(-sq/2.0)
print(round(np.linalg.eigvalsh(K).min(), 6))| check | value |
|---|---|
| min eigenvalue | 0.02634 |
| ≥ 0 ? | yes |
| valid kernel (PSD)? | yes |
Worked example
Your turn: fit a linear and an RBF SVM on XOR and score each. Predict both accuracies before you print.
Hint: SVC(kernel='linear') vs SVC(kernel='rbf', gamma=1.0, C=10), then .score(X, y).
import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
print(SVC(kernel='linear').fit(X, y).score(X, y),
SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y).score(X, y))| kernel | accuracy |
|---|---|
| linear | 0.5 |
| RBF | 1.0 |
Trade off
Comparison matrix
From Milestone 3 — RBF SVM on XOR: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.
| kernel | accuracy |
|---|---|
| linear | 0.5 |
| RBF | 1.0 |
Concept
import numpy as np
from sklearn.svm import SVC
# 1. polynomial kernel == explicit quadratic features
def phi(x, c=1.0):
x1, x2 = x
return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
print(np.isclose((a@b+1)**2, phi(a)@phi(b)))
# 2. RBF Gram is PSD
rng = np.random.default_rng(0)
G = rng.normal(size=(6, 2))
sq = np.sum(G**2,1)[:,None] + np.sum(G**2,1)[None,:] - 2*G@G.T
K = np.exp(-sq/2.0)
print(round(np.linalg.eigvalsh(K).min(), 4), '>= 0')
# 3. RBF SVM cracks XOR, linear fails
X = np.array([[0,0],[0,1],[1,0],[1,1]], float); y = np.array([0,1,1,0])
print(SVC(kernel='linear').fit(X,y).score(X,y),
SVC(kernel='rbf', gamma=1.0, C=10).fit(X,y).score(X,y))| printed line | value |
|---|---|
| kernel == features | True |
| RBF min eig | 0.0263 >= 0 |
| linear / RBF on XOR | 0.5 1.0 |
If the kernel matches its features, the RBF Gram is PSD, and the RBF SVM cracks XOR while linear can't — you have used an infinite feature space without ever visiting it.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| kernel == features | True |
| RBF min eig | 0.0263 >= 0 |
| linear / RBF on XOR | 0.5 1.0 |
Concept
Slides closed, out loud: explain (1) what a kernel avoids computing and why that matters for RBF, (2) why Mercer's PSD condition is the exact right test, and (3) how the RBF SVM separates XOR that a linear one cannot.
Stretch: prove the composition product rule keeps the Gram PSD (Schur product theorem), then re-derive the kernel ridge dual (K+λI)α = y and confirm its predictions match explicit-feature ridge. Kernels return as the SVM dual (Week 25) and as the covariance function of a Gaussian process.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The problem: feature maps · The kernel trick · Mercer's theorem · RBF = infinite dimensions · Composing kernels · Kernelizing ridge. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
φ that makes XOR linearly separable, and explain why forming φ explicitly is expensive or impossible(x·z+c)² = φ(x)·φ(z) term by term and verify it in code (value 4.0)⟺ PSD Gram — accepting RBF (min eig 0.0263) and rejecting a fake kernel (min eig −14.13)exp(xz)1.0| idea | the one thing to remember |
|---|---|
| kernel trick | k(x,z) = φ(x)·φ(z), computed in O(d) |
| polynomial kernel | (x·z+c)² carries the whole quadratic φ |
| Mercer | valid ⟺ Gram matrix PSD (min eig ≥ 0) |
| RBF | exp(−‖x−z‖²/2σ²), infinite-dimensional φ |
| composition | sum, product, positive scaling stay valid; difference does not |
| kernelize | features only via dot products → replace with k; solve the dual |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.