Lesson 28: Kernel Methods

USAAIO Lesson 28, from Week 10, fully worked. It builds feature maps out on XOR, then gives the kernel trick k(x,z)=φ(x)·φ(z), deriving the polynomial kernel (x·z+c)² term by term and matching it to an explicit six-dimensional φ. It applies Mercer's theorem - a kernel is valid exactly when its Gram matrix is PSD - to accept the RBF kernel and reject a fake one by its negative eigenvalue, then shows that the RBF kernel corresponds to an infinite-dimensional feature space, with the series verified numerically. It covers the composition rules for sums, products, and positive scalings, and closes with kernelized ridge regression, where the dual equals the primal, and an RBF SVM that cracks XOR. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 61 slides.

Subject: Machine Learning · 111 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Kernel Methods

Title

USAAIO · Lesson 28 · Week 10 (Kernel Methods)

Reach into a high- (even infinite-) dimensional feature space without ever going there. We derive the kernel trick k(x,z) = φ(x)·φ(z) from the feature map up, prove which similarities are legal kernels via a PSD Gram matrix, show the RBF kernel is an infinite feature space, and kernelize both ridge and an SVM — every algebra step and every printed number verified by real execution.

2. By the end of this lesson you can

Objectives

  1. Write a feature map φ(x) out in full and say exactly why forming it can be expensive or impossible
  2. Derive the polynomial kernel k(x,z) = (x·z + c)² and match it, entry by entry, to an explicit 6-dimensional φ
  3. State and apply Mercer's theorem — k is a valid kernel iff every Gram matrix is PSD — and reject a fake kernel by finding a negative eigenvalue
  4. Explain the RBF kernel exp(−‖x−z‖²/2σ²) as an infinite-dimensional feature space and verify the series numerically
  5. Use the composition rules (sum, product, positive scaling) and kernelize ridge regression and an SVM — solving XOR without ever building φ

3. What survived from Numerical Stability?

Warm-up

Discussion prompt

Before we open Lesson 28: Kernel Methods: without looking back, what was the main idea of Numerical Stability, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

IEEE-754 finite precision (overflow, underflow, cancellation), the log-sum-exp trick and numerically stable softmax, the condition number and ill-conditioned solves, and mixed-precision training. Build a stable softmax and log-sum-exp that survive inputs the naive versions overflow on.

4. The problem: feature maps

Section

Part 1 of 8 — why lift the data

5. The running example: XOR

Concept

Four points, the classic XOR labeling. The two + points sit on one diagonal, the two − points on the other. We carry this dataset through the entire lesson.

pointx₁x₂label y
A000 (−)
B011 (+)
C101 (+)
D110 (−)

No straight line separates the + from the −: any line that puts B and C together also traps A or D with them. XOR is not linearly separable in the input plane.

6. Fill in: x₁ for The running example: XOR

Comparison

Comparison matrix

From The running example: XOR: refill the x₁ column from what you know. The rest of the table is as it appeared.

pointx₁x₂label y
A000 (−)
B011 (+)
C101 (+)
D110 (−)

7. Lift it into a richer space

Intuition

A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.

The trick every kernel method rests on: don't change the classifier, change the space. Map each point x to φ(x) in a higher-dimensional space, run the linear method there, and the decision boundary looks curved back in the original plane.

The coordinate that untangles XOR is the product x₁·x₂ — it alone tells the two diagonals apart. Let's build that map.

8. Break it if you can: Lift it into a richer space

Counterexample

Discussion prompt

A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The coordinate that untangles XOR is the product x₁·x₂ — it alone tells the two diagonals apart. Let's build that map.

9. Picture it first: Linear there = nonlinear here

Picture it

Figure (svg): The four XOR points lifted so that D rises above the plane on a new axis, with a horizontal plane separating the two plus points from the two minus points.

Add the coordinate x₁x₂: point D lifts off the floor, and a flat plane separates the classes.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Picture the lift as lifting the plane into 3-D by the height x₁·x₂. The two − points (A at height 0, D at height 1) split apart in height from the two + points (both at height 0) — a flat sheet can now slide between them.

10. Linear there = nonlinear here

Intuition

Picture the lift as lifting the plane into 3-D by the height x₁·x₂. The two − points (A at height 0, D at height 1) split apart in height from the two + points (both at height 0) — a flat sheet can now slide between them.

Figure (svg): The four XOR points lifted so that D rises above the plane on a new axis, with a horizontal plane separating the two plus points from the two minus points.

Add the coordinate x₁x₂: point D lifts off the floor, and a flat plane separates the classes.

The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.

11. By analogy: Linear there = nonlinear here

Analogy

Discussion prompt

Explain Linear there = nonlinear here by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.

12. A feature map φ

Concept

feature map φ — A function φ: ℝᵈ → ℝᴰ that sends each input x to a longer feature vector φ(x). D (the feature dimension) is usually much larger than d, and can be infinite.

For XOR, take the quadratic map that appends the cross term and squares:

\[ \varphi(x) = \big(1,\; x_1,\; x_2,\; x_1 x_2\big) \]

The new fourth coordinate x₁x₂ is 0 for A, B, C but 1 for D — exactly the signal a line in the lifted space can exploit.

13. Teach it back: A feature map φ

Explain it

Discussion prompt

Explain A feature map φ to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For XOR, take the quadratic map that appends the cross term and squares:

14. Guess the shape of the answer: φ makes XOR linearly separable

Estimation

Predict first

Apply φ(x) = (1, x₁, x₂, x₁x₂) to all four points, then check that the single rule x₁ + x₂ − 2x₁x₂ > 0.5 splits the labels:

Commit before you compute: what does φ makes XOR linearly separable come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Lifted scores are [0, 1, 1, 0]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. In the 4-D feature space the plane w=(0,1,1,−2) gives +1 to the two + points and 0 to the two − points — a clean linear split that was impossible in the 2-D input.

15. φ makes XOR linearly separable

Worked example

Apply φ(x) = (1, x₁, x₂, x₁x₂) to all four points, then check that the single rule x₁ + x₂ − 2x₁x₂ > 0.5 splits the labels:

import numpy as np
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
def phi(x):
    x1, x2 = x
    return np.array([1.0, x1, x2, x1*x2])
Phi = np.array([phi(p) for p in X])          # (4,4) lifted points
w = np.array([0.0, 1.0, 1.0, -2.0])          # a separating plane
scores = Phi @ w
print(scores)                                # [0, 1, 1, 0]
print((scores > 0.5).astype(int))            # matches y

Lifted scores are [0, 1, 1, 0]

Why: In the 4-D feature space the plane w=(0,1,1,−2) gives +1 to the two + points and 0 to the two − points — a clean linear split that was impossible in the 2-D input.

pointφ(x) = (1, x₁, x₂, x₁x₂)score = φ·wlabel
A (0,0)(1, 0, 0, 0)00
B (0,1)(1, 0, 1, 0)11
C (1,0)(1, 1, 0, 0)11
D (1,1)(1, 1, 1, 1)00

16. Which is which, by score = φ·w

Discrimination

Sort into buckets

Sort these by score = φ·w, from memory, without looking back at φ makes XOR linearly separable. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0
A (0,0); D (1,1)
1
B (0,1); C (1,0)
g1
score = φ·w is "0" for A (0,0), D (1,1) — that is what the table on "φ makes XOR linearly separable" records, and it is the single property separating this group from the rest.
g2
score = φ·w is "1" for B (0,1), C (1,0) — that is what the table on "φ makes XOR linearly separable" records, and it is the single property separating this group from the rest.

17. The catch: φ explodes

Concept

Lifting works, but the feature dimension grows fast. All degree-≤2 monomials of a d-dimensional input already number on the order of d²/2; degree p gives on the order of dᵖ.

\[ \text{degree } p,\ \dim d:\quad \dim \varphi \;=\; \binom{d+p}{p} \;\sim\; d^{\,p} \]

For images (d in the thousands) even p = 3 is astronomically large, and the RBF map is infinite-dimensional. Building φ(x) explicitly is hopeless. We need its dot products without ever forming it.

18. The kernel trick

Section

Part 2 of 8 — dot products for free

19. Learners only touch dot products

Intuition

Here is the escape hatch. A whole family of algorithms — the SVM dual, ridge regression, PCA, nearest-mean — never look at a feature vector alone. They only ever combine features through inner products φ(xᵢ)·φ(xⱼ).

So we do not need φ(x) itself. We only need a fast way to compute the number φ(x)·φ(z) for any pair. If a cheap function of x and z returns that number directly, we can run the algorithm in the huge feature space at input-space cost.

20. Definition: a kernel

Concept

kernel — A function k(x, z) that equals the inner product of two feature vectors, k(x,z) = φ(x)·φ(z), for some feature map φ — but is computed directly from x and z, without building φ.

\[ k(x, z) \;=\; \varphi(x) \cdot \varphi(z) \]

The kernel trick: replace every φ(xᵢ)·φ(xⱼ) in the algorithm by k(xᵢ, xⱼ). The feature map is used implicitly and never materialized. The rest of the lesson is (1) which k are legal and (2) what you can do with them.

21. Take the definitions apart: feature map φ vs kernel

Definition probe

Sort into buckets

Every line below is part of the definition of feature map φ or of kernel — one or the other, never both. Put each where it belongs.

feature map φ
ℝᵈ → ℝᴰ that sends each input x to a longer feature vector φ(x).; D (the feature dimension) is usually much larger than d, and can be infinite.
kernel
A function k(x, z) that equals the inner product of two feature vectors, k(x,z) = φ(x)·φ(z), for some feature map φ; but is computed directly from x and z, without building φ.
b1
A function φ: ℝᵈ → ℝᴰ that sends each input x to a longer feature vector φ(x). D (the feature dimension) is usually much larger than d, and can be infinite.
b2
A function k(x, z) that equals the inner product of two feature vectors, k(x,z) = φ(x)·φ(z), for some feature map φ — but is computed directly from x and z, without building φ.

22. What has to be given first: Derive the polynomial kernel — set up

Missing information

Discussion prompt

Take the simplest interesting kernel, k(x,z) = (x·z + c)² in 2-D, and find the φ hiding inside it. Start by writing the dot product out and squaring — nothing skipped.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

x = (x₁, x₂), z = (z₁, z₂), so the dot product is x₁z₁ + x₂z₂, plus the constant c.

23. Derive the polynomial kernel — set up

Worked example

Take the simplest interesting kernel, k(x,z) = (x·z + c)² in 2-D, and find the φ hiding inside it. Start by writing the dot product out and squaring — nothing skipped.

Write x·z + c for 2-D vectors

Why: x = (x₁, x₂), z = (z₁, z₂), so the dot product is x₁z₁ + x₂z₂, plus the constant c.

\[ x\cdot z + c = x_1 z_1 + x_2 z_2 + c \]

Square the trinomial

Why: Use (p+q+r)² = p² + q² + r² + 2pq + 2pr + 2qr with p=x₁z₁, q=x₂z₂, r=c.

\[ (x_1 z_1 + x_2 z_2 + c)^2 \]

24. Decode the notation: Derive the polynomial kernel — set up

Notation

Annotate

From Derive the polynomial kernel — set up — read this one piece at a time. What is each part doing?

On: \( x\cdot z + c = x_1 z_1 + x_2 z_2 + c \)

  • x = (x₁, x₂), z = (z₁, z₂), so the dot product is x₁z₁ + x₂z₂, plus the constant c.
  • Use (p+q+r)² = p² + q² + r² + 2pq + 2pr + 2qr with p=x₁z₁, q=x₂z₂, r=c.

25. Derive the polynomial kernel — expand

Worked example

Expand every term of the square

Why: Square each of the three parts, then add twice each pairwise product.

\[ \begin{aligned} (x\cdot z + c)^2 =\;& x_1^2 z_1^2 + x_2^2 z_2^2 + c^2 \\ &+ 2\,x_1 z_1\, x_2 z_2 + 2c\,x_1 z_1 + 2c\,x_2 z_2 \end{aligned} \]

Regroup so each term is (something in x)·(same thing in z)

Why: Every term factors into an x-part times the matching z-part — that pairing is exactly what an inner product looks like.

\[ \begin{aligned} =\;& (x_1^2)(z_1^2) + (x_2^2)(z_2^2) + (\sqrt2\,x_1 x_2)(\sqrt2\,z_1 z_2) \\ &+ (\sqrt{2c}\,x_1)(\sqrt{2c}\,z_1) + (\sqrt{2c}\,x_2)(\sqrt{2c}\,z_2) + (c)(c) \end{aligned} \]

26. Say it in words: Derive the polynomial kernel — expand

Translation

\( \begin{aligned} =\;& (x_1^2)(z_1^2) + (x_2^2)(z_2^2) + (\sqrt2\,x_1 x_2)(\sqrt2\,z_1 z_2) \\ &+ (\sqrt{2c}\,x_1)(\sqrt{2c}\,z_1) + (\sqrt{2c}\,x_2)(\sqrt{2c}\,z_2) + (c)(c) \end{aligned} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

27. Complete the line: Derive the polynomial kernel — read off φ

Fill the middle

Fill in the blanks

From Derive the polynomial kernel — read off φ — finish the line. Write what belongs on the right of the equals sign before you look.

\boxed\varphi(x)\cdot\varphi(z)\,}}

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Each parenthesized x-piece becomes a coordinate of the feature map.

28. Derive the polynomial kernel — read off φ

Worked example

Collect the six x-parts into one vector φ(x)

Why: Each parenthesized x-piece becomes a coordinate of the feature map. The √2 factors are what make the cross terms line up as a clean dot product.

\[ \varphi(x) = \big(x_1^2,\; x_2^2,\; \sqrt2\,x_1 x_2,\; \sqrt{2c}\,x_1,\; \sqrt{2c}\,x_2,\; c\big) \]

Therefore (x·z + c)² = φ(x)·φ(z)

Why: The kernel — one dot product plus a square, computed on the raw 2-vectors — equals the dot product of these 6-dimensional feature vectors. Left side: O(d) work. Right side: O(D) work. Same number.

\[ \boxed{\,(x\cdot z + c)^2 = \varphi(x)\cdot\varphi(z)\,} \]

The kernel gave us the entire quadratic feature space — squares, the cross term, and scaled linears — for the price of a dot product. Now verify it on real numbers.

29. Work backwards from the answer: Derive the polynomial kernel — read off φ

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Therefore (x·z + c)² = φ(x)·φ(z)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The kernel gave us the entire quadratic feature space — squares, the cross term, and scaled linears — for the price of a dot product. Now verify it on real numbers.

30. Predict the next row: Verify the polynomial kernel in code

Pattern

Predict first

The table runs: a·b | 1.0 · kernel (a·b + c)² | 4.0 · φ(a)·φ(b) explicit | 3.9999999999999987

In Verify the polynomial kernel in code, given the rows so far: what is the next one — the row where quantity is np.isclose(...)?

Correct: np.isclose(...) | True

quantityvalue (verified)
a·b1.0
kernel (a·b + c)²4.0
φ(a)·φ(b) explicit3.9999999999999987
np.isclose(...)True

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The cheap kernel gives exactly 4.0.

31. Verify the polynomial kernel in code

Worked example

Take a = (1, 2), b = (3, −1), c = 1. Compute the kernel the cheap way and the explicit-feature way; they must agree. Runnable as written:

import numpy as np
def phi(x, c=1.0):
    x1, x2 = x
    return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
                     np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.]); c = 1.0
k_direct = (a @ b + c)**2         # 1 dot + 1 square
k_feat   = phi(a, c) @ phi(b, c)  # 6-dim dot product
print(k_direct)                   # 4.0
print(k_feat)                     # 3.9999999999999987
print(np.isclose(k_direct, k_feat))

a·b = 1(3) + 2(−1) = 1, so (a·b + c)² = (1 + 1)² = 4

Why: The cheap kernel gives exactly 4.0. The explicit map gives 3.9999999999999987 — identical up to floating-point round-off, so np.isclose returns True.

quantityvalue (verified)
a·b1.0
kernel (a·b + c)²4.0
φ(a)·φ(b) explicit3.9999999999999987
np.isclose(...)True

32. What each one costs: Verify the polynomial kernel in code

Trade off

Comparison matrix

From Verify the polynomial kernel in code: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityvalue (verified)
a·b1.0
kernel (a·b + c)²4.0
φ(a)·φ(b) explicit3.9999999999999987
np.isclose(...)True

33. Term-by-term: the dot product lands on 4

Worked example

To see there is no magic, print each feature vector and multiply component-wise. The six products must sum to the kernel value 4.

import numpy as np
def phi(x, c=1.0):
    x1, x2 = x
    return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
                     np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
pa, pb = phi(a), phi(b)
print(np.round(pa, 4))            # phi(a)
print(np.round(pb, 4))            # phi(b)
print(np.round(pa * pb, 4))       # elementwise products
print(np.sum(pa * pb))            # 3.9999999999999982 (== 4 up to round-off)

Elementwise products are [9, 4, −12, 6, −4, 1]

Why: Sum: 9 + 4 − 12 + 6 − 4 + 1 = 4. Each coordinate contributes; the cross term (−12) and scaled-linear terms (6, −4) are exactly the parts a plain (x·z) miss.

coordφ(a)φ(b)product
x₁²1.09.09.0
x₂²4.01.04.0
√2·x₁x₂2.8284−4.2426−12.0
√2·x₁1.41424.24266.0
√2·x₂2.8284−1.4142−4.0
c1.01.01.0

34. Watch it run: Term-by-term: the dot product lands on 4

Pattern

Step through it

Step through Term-by-term: the dot product lands on 4 one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: coord is x₁²
  2. Step 2: coord is x₂²
  3. Step 3: coord is √2·x₁x₂
  4. Step 4: coord is √2·x₁
  5. Step 5: coord is √2·x₂
  6. Step 6: coord is c

35. The general polynomial kernel

Concept

The degree-2 case generalizes. The kernel k(x,z) = (x·z + c)ᵖ corresponds to the feature map of all monomials up to degree p, with the constant c balancing lower- against higher-degree terms.

\[ k(x,z) = (x\cdot z + c)^p \;=\; \varphi_p(x)\cdot \varphi_p(z),\qquad \dim \varphi_p = \binom{d+p}{p} \]

For d = 2, p = 2 that is C(4,2) = 6 features — exactly the 6-vector we just derived. But the kernel is always a single dot-plus-power on the raw inputs: the feature count never enters the computation.

36. Something is wrong here: a kernel costs feature-space work

Anomaly

Predict first

A student writes this, and it looks reasonable:

The feature map has 6 dimensions, so computing k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.

The kernel is computed on the raw inputs: one d-dimensional dot product x·z, add c, square. Cost scales with the input dimension d, not the feature dimension D.

Why: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.

37. Trap: a kernel costs feature-space work

Trap

The trap

The feature map has 6 dimensions, so computing k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.

Cost of k ≈ dim φ = 6 (grows with the feature space)

Why: This misses the whole point. If the kernel cost scaled with the feature dimension, the trick would buy nothing — and for the infinite-dimensional RBF it would be impossible.

The fix

The kernel is computed on the raw inputs: one d-dimensional dot product x·z, add c, square. Cost scales with the input dimension d, not the feature dimension D.

Cost of k = O(d) — independent of dim φ

Why: (x·z + c)² needs 2 multiplies + 1 add + 1 square regardless of whether φ is 6-dimensional or infinite. That O(d)-not-O(D) gap IS the kernel trick; it is why RBF's infinite φ is free.

38. Mercer's theorem

Section

Part 3 of 8 — which k are legal

39. Not every similarity is a kernel

Concept

You could invent any symmetric similarity(x, z) and hope it behaves like a kernel. But k is only legal if some feature map φ actually produces it — otherwise k(x,z) = φ(x)·φ(z) is a lie and the algorithms built on it lose their guarantees.

Mercer's theorem gives a checkable condition, no φ required: look at the Gram matrix of k on any dataset.

40. The Gram matrix and the PSD test

Concept

Gram matrix — For points x₁,…,xₙ, the n×n matrix K with Kᵢⱼ = k(xᵢ, xⱼ) — every pairwise kernel value.

Mercer's theorem: k is a valid kernel iff its Gram matrix K is positive semidefinite (PSD) for every finite dataset:

\[ K_{ij} = k(x_i, x_j),\qquad z^\top K z \ge 0 \ \text{ for all } z \iff K \succeq 0 \]

Why PSD is exactly right: if K = ΦΦᵀ (features stacked in rows of Φ), then zᵀKz = ‖Φᵀz‖² ≥ 0 automatically. PSD is precisely the fingerprint of 'these numbers came from real inner products.' In practice: check min eigenvalue ≥ 0 (Lesson 13).

41. Why a real kernel's Gram is always PSD

Worked example

One direction of Mercer is a two-line proof. If k(xᵢ,xⱼ) = φ(xᵢ)·φ(xⱼ), stack the features as rows of Φ; then K = ΦΦᵀ. Show zᵀKz ≥ 0 for any z:

Write the quadratic form with K = ΦΦᵀ

Why: Substitute K = ΦΦᵀ into zᵀKz. Matrix multiplication is associative, so we can regroup the product.

\[ z^\top K z = z^\top \Phi \Phi^\top z = (\Phi^\top z)^\top (\Phi^\top z) \]

That is a squared norm, hence ≥ 0

Why: Any vector dotted with itself is its squared length, which can never be negative. So every Gram matrix of real inner products is PSD — no assumption on the data needed.

\[ z^\top K z = \lVert \Phi^\top z\rVert^2 \;\ge\; 0 \quad\Longrightarrow\quad K \succeq 0 \]

stepexpressionwhy
substitutezᵀΦΦᵀzK = ΦΦᵀ
regroup(Φᵀz)ᵀ(Φᵀz)associativity
conclude‖Φᵀz‖² ≥ 0norms are non-negative

42. Decode the notation: Why a real kernel's Gram is always PSD

Notation

Annotate

From Why a real kernel's Gram is always PSD — read this one piece at a time. What is each part doing?

On: \( z^\top K z = \lVert \Phi^\top z\rVert^2 \;\ge\; 0 \quad\Longrightarrow\quad K \succeq 0 \)

  • Substitute K = ΦΦᵀ into zᵀKz. Matrix multiplication is associative, so we can regroup the product.
  • Any vector dotted with itself is its squared length, which can never be negative. So every Gram matrix of real inner products is PSD — no assumption on the data needed.

43. Restore the missing line: The RBF Gram matrix is PSD

Fill the middle

Fill in the blanks

From The RBF Gram matrix is PSD — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X2, 1)[:, None]
+ np.sum(X
2, 1)[None, :]
- 2 * X @ X.T) # ||xi - xj||^2 for all pairs
K = np.exp(-sq / 2.0) # RBF, sigma = 1
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4)) # 0.0263 >= 0

Why: sq is what everything below it consumes, so the wrong expression here fails later and somewhere else. Every eigenvalue is positive, so K is PSD on this dataset — consistent with Mercer.

44. The RBF Gram matrix is PSD

Worked example

The RBF (Gaussian) kernel k(x,z) = exp(−‖x−z‖²/2σ²) is the workhorse. Build its Gram matrix on 6 random points (seeded) and confirm the smallest eigenvalue is ≥ 0:

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X**2, 1)[:, None]
      + np.sum(X**2, 1)[None, :]
      - 2 * X @ X.T)            # ||xi - xj||^2 for all pairs
K = np.exp(-sq / 2.0)          # RBF, sigma = 1
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4))     # 0.0263  >= 0

min eigenvalue = 0.0263 ≥ 0 → Gram is PSD → valid kernel

Why: Every eigenvalue is positive, so K is PSD on this dataset — consistent with Mercer. The diagonal is all 1s because exp(−0/2σ²)=1 (a point is maximally similar to itself).

eigenvalue #value
1 (smallest)0.0263
20.0571
30.4790
6 (largest)3.3872

45. Watch it run: The RBF Gram matrix is PSD

Pattern

Step through it

Step through The RBF Gram matrix is PSD one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: eigenvalue # is 1 (smallest)
  2. Step 2: eigenvalue # is 2
  3. Step 3: eigenvalue # is 3
  4. Step 4: eigenvalue # is 6 (largest)

46. Guess the shape of the answer: A fake kernel fails Mercer

Estimation

Predict first

Test the plausible-looking candidate k(x,z) = x·z − ‖x−z‖². Build its Gram matrix on the same 6 points and look for a negative eigenvalue:

Commit before you compute: what does A fake kernel fails Mercer come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: min eigenvalue = −14.1257 < 0 → NOT PSD → NOT a valid kernel

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. One negative eigenvalue is a counterexample to Mercer's condition on this very dataset, so no feature map φ could ever produce this k.

47. A fake kernel fails Mercer

Worked example

Test the plausible-looking candidate k(x,z) = x·z − ‖x−z‖². Build its Gram matrix on the same 6 points and look for a negative eigenvalue:

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = (np.sum(X**2, 1)[:, None]
      + np.sum(X**2, 1)[None, :]
      - 2 * X @ X.T)
K = X @ X.T - sq               # candidate k = x.z - ||x-z||^2
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 4))     # -14.1257  < 0

min eigenvalue = −14.1257 < 0 → NOT PSD → NOT a valid kernel

Why: One negative eigenvalue is a counterexample to Mercer's condition on this very dataset, so no feature map φ could ever produce this k. It only takes a single negative eigenvalue to disqualify a candidate.

checkvalue
eigenvalues (rounded)[−14.13, 0, 0, 1.54, 3.66, 14.88]
min eigenvalue−14.1257
PSD?no
valid kernel?no

48. Finish it with less help: The linear kernel passes (the honest…

Faded example

Fill in the blanks

The linear kernel passes (the honest baseline), with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
K = X @ X.T # linear-kernel Gram = X X^T
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 6)) # -0.0 (round-off), i.e. >= 0

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. K = XXᵀ is PSD by the ΦΦᵀ proof.

49. The linear kernel passes (the honest baseline)

Worked example

For contrast, the linear kernel k(x,z) = x·z has Gram K = XXᵀ — literally ΦΦᵀ with φ = identity. Its min eigenvalue should be ≥ 0 (allowing tiny negative round-off) on the same 6 points:

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
K = X @ X.T                    # linear-kernel Gram = X X^T
eig = np.linalg.eigvalsh(K)
print(np.round(eig, 4))
print(round(eig.min(), 6))     # -0.0 (round-off), i.e. >= 0

min eigenvalue = −0.0 (numerically 0) → PSD → valid

Why: K = XXᵀ is PSD by the ΦΦᵀ proof. With only 2 columns in X, K has rank 2, so four eigenvalues are ~0 (the tiny −0.0 is floating-point noise, not a real negative) and two are positive. The linear kernel is always valid.

checkvalue
eigenvalues (rounded)[−0, −0, 0, 0, 0.9956, 4.9605]
min eigenvalue−0.0 (≈ 0)
PSD? / valid kernel?yes / yes

50. Something is wrong here: any similarity function is a kernel

Anomaly

Predict first

A student writes this, and it looks reasonable:

k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Symmetry is necessary but NOT sufficient.

Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.

Why: Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.

51. Trap: any similarity function is a kernel

Trap

The trap

k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.

Feed an ad-hoc symmetric similarity to the solver

Why: Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.

The fix

Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.

Verify Gram PSD (or use a proven kernel) before fitting

Why: The candidate above has min eigenvalue −14.13, so it is rejected. RBF (min eig 0.0263) and polynomial pass. One eigvalsh call is cheap insurance against a silently broken model.

52. Break it on purpose: any similarity function is a kernel

Break the constraint

Discussion prompt

The rule this trap just fixed:

Check Mercer's condition first: the Gram matrix must be PSD (min eigenvalue ≥ 0). Use kernels already proven valid — linear, polynomial, RBF.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Symmetry is necessary but NOT sufficient. Without a PSD Gram matrix there is no underlying inner-product space, the SVM's dual is no longer convex, and the optimizer can diverge or return nonsense.

53. RBF = infinite dimensions

Section

Part 4 of 8 — the Gaussian kernel

54. Distance-based similarity

Intuition

The RBF kernel exp(−‖x−z‖²/2σ²) reads as pure geometry: it is 1 when x = z (zero distance) and decays smoothly toward 0 as the points move apart. It is a similarity that only cares about distance.

The bandwidth σ sets the reach: small σ means similarity drops off fast (each point influences only its close neighbors — a wiggly boundary); large σ means slow decay (far points still count — a smooth boundary).

55. RBF values: near vs far

Worked example

Compute the RBF between our running pair a=(1,2), b=(3,-1) (far apart) and between two nearby points, at two bandwidths:

import numpy as np
def rbf(x, z, sigma=1.0):
    return np.exp(-np.sum((x - z)**2) / (2 * sigma**2))
a = np.array([1., 2.]); b = np.array([3., -1.])
p = np.array([1., 2.]); q = np.array([1.1, 2.1])
print(np.sum((a - b)**2))       # 13.0  (far)
print(rbf(a, a))                # 1.0   (self)
print(round(rbf(a, b, 1.0), 6)) # 0.001503
print(round(rbf(a, b, 3.0), 6)) # 0.485672
print(round(rbf(p, q, 1.0), 6)) # 0.99005 (near)

Far pair: σ=1 → 0.0015 (almost dissimilar); σ=3 → 0.4857

Why: ‖a−b‖²=13. At σ=1 that distance kills the similarity; widening to σ=3 lets the same pair count as ~half-similar. A near pair (distance 0.02) stays at 0.990 either way.

pair‖·‖²σRBF value
a with itself011.0
a, b (far)1310.001503
a, b (far)1330.485672
p, q (near)0.0210.99005

56. What happens as it grows: RBF values: near vs far

Scale up

Step through it

Step through RBF values: near vs far and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: pair is a with itself
  2. Step 2: pair is a, b (far)
  3. Step 3: pair is a, b (far)
  4. Step 4: pair is p, q (near)

57. σ and sklearn's γ are the same knob

Concept

sklearn writes the RBF kernel with a single parameter gamma instead of the bandwidth σ. They are the same dial, reciprocally related:

\[ \exp\!\Big(-\frac{\lVert x-z\rVert^2}{2\sigma^2}\Big) = \exp\big(-\gamma\,\lVert x-z\rVert^2\big),\qquad \gamma = \frac{1}{2\sigma^2} \]

So large γ = small σ = fast decay = wiggly, local boundary (risk of overfitting); small γ = large σ = smooth, global boundary (risk of underfitting). When we pass gamma=1.0 to SVC shortly, that is σ = 1/√2 — a tight, local kernel, which is exactly what four well-separated XOR points want.

58. Bandwidth shapes the boundary

Intuition

Think of each training point as placing a soft Gaussian bump of influence around itself, of width σ. The decision boundary threads between the summed bumps of the two classes.

Wide bumps (large σ, small γ) overlap and blur into one smooth boundary — simple, may underfit. Narrow bumps (small σ, large γ) stay isolated islands around each point — flexible, may memorize noise (overfit). Tuning σ/γ on validation data is the RBF SVM's core bias–variance knob, the twin of λ in ridge.

59. Why RBF is infinite-dimensional

Concept

Expand the exponential of an inner product as its power series. Because exp(t) = Σ tⁿ/n!, the kernel exp(x·z) splits into a sum over all polynomial degrees at once:

\[ \exp(x\cdot z) = \sum_{n=0}^{\infty} \frac{(x\cdot z)^n}{n!} \]

Each term (x·z)ⁿ/n! is itself a polynomial kernel of degree n, carrying its own block of monomial features. Summing over n = 0, 1, 2, … stacks infinitely many feature coordinates. RBF is exp(x·z) dressed with normalizing factors, so it too has an infinite φ — one you could never write down, yet dot in O(d).

60. Restore the missing line: Verify the infinite series converges

Fill the middle

Fill in the blanks

From Verify the infinite series converges — one line has had its right-hand side removed. Put it back.

import numpy as np
from math import factorial
def feat(x, N):
return np.array([xn / np.sqrt(factorial(n)) for n in range(N)])
x, z = 0.7, 0.4
for N in [3, 6, 12]:
approx =
feat(x, N) @ feat(z, N)**
print(N, round(approx, 8), round(np.exp(x*z), 8))

Why: approx is what everything below it consumes, so the wrong expression here fails later and somewhere else. With only 3 features the dot is 1.31920 (off in the 3rd decimal); 6 features give 1.32312912; 12 features match exp(0.28)=1.32312981 exactly.

61. Verify the infinite series converges

Worked example

In 1-D the feature map is φ(x)ₙ = xⁿ/√(n!), so φ(x)·φ(z) = Σ (xz)ⁿ/n! = exp(xz). Truncate at N terms and watch it converge to the true kernel for x=0.7, z=0.4:

import numpy as np
from math import factorial
def feat(x, N):
    return np.array([x**n / np.sqrt(factorial(n)) for n in range(N)])
x, z = 0.7, 0.4
for N in [3, 6, 12]:
    approx = feat(x, N) @ feat(z, N)
    print(N, round(approx, 8), round(np.exp(x*z), 8))

Truncated dot product → exp(xz) = 1.32312981 as N grows

Why: With only 3 features the dot is 1.31920 (off in the 3rd decimal); 6 features give 1.32312912; 12 features match exp(0.28)=1.32312981 exactly. The infinite feature vector is real — we just don't need all of it.

N (features kept)feat·featexp(xz)
31.319200001.32312981
61.323129121.32312981
121.323129811.32312981

62. What stays fixed: Verify the infinite series converges

Invariant

Step through it

Step through Verify the infinite series converges one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: N (features kept) is 3
  2. Step 2: N (features kept) is 6
  3. Step 3: N (features kept) is 12

63. Composing kernels

Section

Part 5 of 8 — build new from old

64. The composition rules

Concept

Valid kernels are closed under a few operations, so you can assemble complex kernels from simple proven ones and stay inside Mercer's guarantee:

Not guaranteed: a difference k₁ − k₂ — subtracting PSD matrices can create a negative eigenvalue. Let's check all four on the Gram matrices.

65. What composition buys you

Intuition

The rules let you design a kernel to match your data's structure without ever proving PSD from scratch. Start from proven blocks and glue them:

\[ k(x,z) = \underbrace{x\cdot z}_{\text{linear trend}} \;+\; \underbrace{\exp\!\big(-\gamma\lVert x-z\rVert^2\big)}_{\text{local bumps}} \]

This sum captures a broad linear trend and fine local structure at once — and because sum-of-kernels is a kernel, it is automatically valid. Products let you combine kernels on different feature groups (e.g. an RBF on the numeric columns times a kernel on the categorical ones).

66. Restore the missing line: Composition preserves PSD

Fill the middle

Fill in the blanks

From Composition preserves PSD — one line has had its right-hand side removed. Put it back.

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X2,1)[:,None] + np.sum(X2,1)[None,:] - 2*X@X.T
K_rbf = np.exp(-sq/2.0)
K_lin = X @ X.T
m = lambda M: round(np.linalg.eigvalsh(M).min(), 6)
print(m(K_lin + K_rbf)) # 0.027333 sum
print(m(K_lin * K_rbf)) # 0.012704 product (elementwise)
print(m(3.0 * K_rbf)) # 0.079021 scaling
print(m(K_lin - K_rbf)) # -3.377703 difference: BREAKS PSD

Why: K_rbf is what everything below it consumes, so the wrong expression here fails later and somewhere else. Sum 0.0273, product 0.0127, 3×RBF 0.0790 are all ≥ 0 — still valid kernels.

67. Composition preserves PSD

Worked example

Take a linear Gram and an RBF Gram on the same 6 points. Check the min eigenvalue of the sum, product, a positive scaling, and the (illegal) difference:

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X**2,1)[:,None] + np.sum(X**2,1)[None,:] - 2*X@X.T
K_rbf = np.exp(-sq/2.0)
K_lin = X @ X.T
m = lambda M: round(np.linalg.eigvalsh(M).min(), 6)
print(m(K_lin + K_rbf))    # 0.027333  sum
print(m(K_lin * K_rbf))    # 0.012704  product (elementwise)
print(m(3.0 * K_rbf))      # 0.079021  scaling
print(m(K_lin - K_rbf))    # -3.377703 difference: BREAKS PSD

Sum, product, scaling stay PSD; the difference goes negative

Why: Sum 0.0273, product 0.0127, 3×RBF 0.0790 are all ≥ 0 — still valid kernels. The difference K_lin − K_rbf hits −3.3777, so it is NOT a kernel. This is why the rules list sum/product/scaling but never subtraction.

combinationmin eigenvaluevalid kernel?
k_lin + k_rbf0.027333yes
k_lin · k_rbf0.012704yes
3 · k_rbf0.079021yes
k_lin − k_rbf−3.377703no

68. Fill in: min eigenvalue for Composition preserves PSD

Comparison

Comparison matrix

From Composition preserves PSD: refill the min eigenvalue column from what you know. The rest of the table is as it appeared.

combinationmin eigenvaluevalid kernel?
k_lin + k_rbf0.027333yes
k_lin · k_rbf0.012704yes
3 · k_rbf0.079021yes
k_lin − k_rbf−3.377703no

69. Kernelizing ridge

Section

Part 6 of 8 — the dual view

70. The representer idea

Concept

Kernelizing an algorithm means rewriting it so features appear only inside dot products, which then become kernel calls. For ridge regression the key fact (the representer theorem) is that the optimal weight vector is a combination of the training features:

\[ w^\star = \sum_{i} \alpha_i\, \varphi(x_i) = \Phi^\top \alpha \]

Substituting this back turns the problem from solving for w in feature space into solving for the n dual coefficients α — and every φ cancels into a Gram matrix K = ΦΦᵀ of kernel values.

71. When kernelizing pays off

Concept

Primal ridge solves a D × D system (D = feature dimension); the kernel dual solves an n × n system (n = number of training points). You pick whichever matrix is smaller.

viewsolvecost driverbest when
primal(ΦᵀΦ + λI) w = ΦᵀyD × Dfew features (D ≪ n)
dual (kernel)(K + λI) α = yn × nhuge/∞ features (D ≫ n)

This is why kernels shine with the RBF map: D = ∞, so the primal is impossible, but the dual is just an n × n solve. The feature dimension has been traded away for the dataset size.

72. Predict the next row: Primal ridge vs the kernel dual

Pattern

Predict first

The table runs: 0 | 0.55081 | 0.55081 · 1 | 1.952872 | 1.952872 · 2 | 5.063328 | 5.063328

In Primal ridge vs the kernel dual, given the rows so far: what is the next one — the row where point x is 3?

Correct: 3 | 9.88218 | 9.88218

point xprimal ŷdual ŷ
00.550810.55081
11.9528721.952872
25.0633285.063328
39.882189.88218

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Both give [0.55081, 1.952872, 5.063328, 9.88218] and np.allclose returns True.

73. Primal ridge vs the kernel dual

Worked example

Primal ridge in explicit features solves (ΦᵀΦ + λI)w = Φᵀy. The dual solves (K + λI)α = y with K = ΦΦᵀ, then predicts ŷ = Kα. They give identical predictions — verify on 4 one-dimensional points with the degree-2 kernel k(x,z)=(xz+1)²:

import numpy as np
X = np.array([[0.],[1.],[2.],[3.]]); y = np.array([1., 2., 5., 10.])
lam = 1.0
def phi(x):
    x = x[0]
    return np.array([1.0, np.sqrt(2)*x, x**2])
Phi = np.array([phi(xi) for xi in X])              # explicit features
w = np.linalg.solve(Phi.T@Phi + lam*np.eye(3), Phi.T@y)
pred_primal = Phi @ w
K = (X @ X.T + 1.0)**2                              # Gram, no phi
alpha = np.linalg.solve(K + lam*np.eye(4), y)       # dual solve
pred_dual = K @ alpha
print(np.round(pred_primal, 6))
print(np.round(pred_dual, 6))
print(np.allclose(pred_primal, pred_dual))          # True

Primal and dual predictions match to machine precision

Why: Both give [0.55081, 1.952872, 5.063328, 9.88218] and np.allclose returns True. The dual never builds the 3-dim φ beyond checking — it works entirely through the 4×4 kernel matrix K.

point xprimal ŷdual ŷ
00.550810.55081
11.9528721.952872
25.0633285.063328
39.882189.88218

74. Watch it run: Primal ridge vs the kernel dual

Pattern

Step through it

Step through Primal ridge vs the kernel dual one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: point x is 0
  2. Step 2: point x is 1
  3. Step 3: point x is 2
  4. Step 4: point x is 3

75. Kernelizing an SVM

Concept

The SVM's dual objective (Lesson 24's Lagrangian) contains data only as inner products φ(xᵢ)·φ(xⱼ). Replace each by k(xᵢ, xⱼ) and the SVM runs in the feature space with a decision function that is a weighted sum of kernels:

\[ f(x) = \operatorname{sign}\!\Big(\sum_{i} \alpha_i\, y_i\, k(x_i, x) + b\Big) \]

Only the support vectors have αᵢ ≠ 0. With an RBF kernel this becomes a smooth nonlinear boundary — enough to finally crack XOR.

76. Teach it back: Kernelizing an SVM

Explain it

Discussion prompt

Explain Kernelizing an SVM to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Only the support vectors have αᵢ ≠ 0. With an RBF kernel this becomes a smooth nonlinear boundary — enough to finally crack XOR.

77. What has to be given first: Kernel SVM solves XOR

Missing information

Discussion prompt

Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The linear kernel = no feature lift = a straight boundary, stuck at chance on XOR. The RBF kernel implicitly lifts into an infinite space where XOR is separable — the kernel trick doing exactly what the explicit φ = (1,x₁,x₂,x₁x₂) did in Part 1, but for free.

78. Kernel SVM solves XOR

Worked example

Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
lin = SVC(kernel='linear').fit(X, y).score(X, y)
rbf = SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y).score(X, y)
print(lin)   # 0.5  (fails)
print(rbf)   # 1.0  (perfect)

linear accuracy 0.5 (fails), RBF accuracy 1.0 (perfect)

Why: The linear kernel = no feature lift = a straight boundary, stuck at chance on XOR. The RBF kernel implicitly lifts into an infinite space where XOR is separable — the kernel trick doing exactly what the explicit φ = (1,x₁,x₂,x₁x₂) did in Part 1, but for free.

kernelXOR accuracyboundary
linear0.5 (fails)straight line
RBF1.0curved (nonlinear)
poly deg-21.0curved (nonlinear)

79. Work backwards from the answer: Kernel SVM solves XOR

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

linear accuracy 0.5 (fails), RBF accuracy 1.0 (perfect)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Fit sklearn's SVM on our XOR data with a linear kernel and an RBF kernel; score both. The linear kernel cannot bend, the RBF can:

80. Guess the shape of the answer: Inside the RBF fit: every point is a support…

Estimation

Predict first

Look at what the RBF SVM actually decided. Print its per-point predictions and how many support vectors it kept — for four tightly-packed points, expect all of them:

Commit before you compute: what does Inside the RBF fit: every point is a support vector come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: predictions [0 1 1 0] match the labels; 4 support vectors

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every prediction is correct. All four points are support vectors (α ≠ 0) because with only four points and a tight kernel, each one is needed to shape the boundary.

81. Inside the RBF fit: every point is a support vector

Worked example

Look at what the RBF SVM actually decided. Print its per-point predictions and how many support vectors it kept — for four tightly-packed points, expect all of them:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
clf = SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y)
print(clf.predict(X))          # [0 1 1 0]
print(y)                       # [0 1 1 0]
print(int(clf.n_support_.sum()))  # 4

predictions [0 1 1 0] match the labels; 4 support vectors

Why: Every prediction is correct. All four points are support vectors (α ≠ 0) because with only four points and a tight kernel, each one is needed to shape the boundary. The decision function is Σ αᵢ yᵢ k(xᵢ, x) + b — pure kernel evaluations.

pointtrue yRBF prediction
A (0,0)00
B (0,1)11
C (1,0)11
D (1,1)00

82. Which is which, by true y

Discrimination

Sort into buckets

Sort these by true y, from memory, without looking back at Inside the RBF fit: every point is a support…. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0
A (0,0); D (1,1)
1
B (0,1); C (1,0)
g1
true y is "0" for A (0,0), D (1,1) — that is what the table on "Inside the RBF fit: every point is a…" records, and it is the single property separating this group from the rest.
g2
true y is "1" for B (0,1), C (1,0) — that is what the table on "Inside the RBF fit: every point is a…" records, and it is the single property separating this group from the rest.

83. Something is wrong here: a linear kernel for nonlinear data

Anomaly

Predict first

A student writes this, and it looks reasonable:

An SVM is a powerful margin classifier, so a linear-kernel SVM will surely handle XOR.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A linear kernel does NO feature lift — its decision boundary is a single straight line.

Match the kernel to the data: use RBF or polynomial whenever the structure is nonlinear.

Why: A linear kernel does NO feature lift — its decision boundary is a single straight line. XOR is not linearly separable, so the SVM is stuck at 50% no matter how you tune C. The power of an SVM is the margin, not nonlinearity.

84. Trap: a linear kernel for nonlinear data

Trap

The trap

An SVM is a powerful margin classifier, so a linear-kernel SVM will surely handle XOR.

SVC(kernel='linear') on XOR → 0.5

Why: A linear kernel does NO feature lift — its decision boundary is a single straight line. XOR is not linearly separable, so the SVM is stuck at 50% no matter how you tune C. The power of an SVM is the margin, not nonlinearity.

The fix

Match the kernel to the data: use RBF or polynomial whenever the structure is nonlinear.

SVC(kernel='rbf') on XOR → 1.0

Why: The nonlinear kernel supplies the implicit feature lift that makes XOR separable — 100% accuracy. Choosing the kernel is where you inject the nonlinearity; the SVM machinery around it is unchanged.

85. Which of these survive contact with Lesson 28: Kernel Methods?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Four points, the classic XOR labeling. The two + points sit on one diagonal, the two − points on the other. We carry this dataset through the entire lesson.; A line fails in 2-D — but if we add a new coordinate built from the old ones, the four points can spread into a space where a flat plane does separate them.; The separating plane in the lifted space, viewed back in the original plane, is a curve. Same linear machinery, curved boundary — that is the entire payoff of lifting.
Breaks
The feature map has 6 dimensions, so computing k(x,z) = (x·z+c)² must cost about 6 multiplies — the same as building φ and dotting.; k(x,z) = x·z − ‖x−z‖² is symmetric and looks like a sensible similarity, so drop it straight into a kernel SVM.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 28: Kernel Methods puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

86. Where kernels go from here

Concept

The trick generalizes far beyond the SVM. Anywhere an algorithm touches data only through dot products, you can kernelize it:

Master the three moves — dot-products-only, Mercer-check, solve-the-dual — and every one of these is the same idea in new clothes.

87. By analogy: Where kernels go from here

Analogy

Discussion prompt

Explain Where kernels go from here by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The trick generalizes far beyond the SVM. Anywhere an algorithm touches data only through dot products, you can kernelize it:

88. The kernel toolkit

Section

Part 7 of 8 — pattern & checks

89. Rebuild the recipe: The kernel-method recipe

Ranking

Put in order

These are the steps of The kernel-method recipe, scrambled. Put them back in order before the next slide shows you.

  1. Kernelize: rewrite the algorithm so features appear only as dot products φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)
  2. Pick a kernel: linear (no lift), polynomial (x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reach
  3. Validate (Mercer): the Gram matrix K must be PSD (min eig ≥ 0) — verify, or compose from proven kernels
  4. Compose if needed: sum, product, and positive scaling of valid kernels stay valid; differences do not
  5. Fit & tune: solve the dual ((K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λ

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

90. The kernel-method recipe

Pattern

  1. Kernelize: rewrite the algorithm so features appear only as dot products φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)
  2. Pick a kernel: linear (no lift), polynomial (x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reach
  3. Validate (Mercer): the Gram matrix K must be PSD (min eig ≥ 0) — verify, or compose from proven kernels
  4. Compose if needed: sum, product, and positive scaling of valid kernels stay valid; differences do not
  5. Fit & tune: solve the dual ((K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λ

91. Where does it stop working: The kernel-method recipe

Edge cases

Discussion prompt

The kernel-method recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Kernelize: rewrite the algorithm so features appear only as dot products φ(xᵢ)·φ(xⱼ), then replace each with k(xᵢ, xⱼ)
  2. Pick a kernel: linear (no lift), polynomial (x·z+c)ᵖ, or RBF exp(−‖x−z‖²/2σ²) for infinite-dimensional reach
  3. Validate (Mercer): the Gram matrix K must be PSD (min eig ≥ 0) — verify, or compose from proven kernels
  4. Compose if needed: sum, product, and positive scaling of valid kernels stay valid; differences do not
  5. Fit & tune: solve the dual ((K+λI)α = y for ridge; the SVM dual for classification), then tune σ/degree and C/λ

92. Rule out three: Check yourself — the kernel trick

Elimination

Eliminate the wrong options

The kernel trick lets an algorithm work in a feature space φ without ever forming φ(x). It does this by:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. computing k(x,z) = φ(x)·φ(z) directly from x and z, in input-space cost
  • B. approximating φ(x) with a small neural network
  • C. projecting φ(x) onto its top principal component
  • D. discarding the feature space and using raw x·z instead

Survives elimination: A

Why: A kernel returns the exact inner product φ(x)·φ(z) as a cheap function of the raw inputs. Algorithms that use features only through dot products (SVM, ridge, PCA) then run in the feature space at O(d) cost per pair.

93. Check yourself — the kernel trick

Check

What is a kernel actually computing?

Check your understanding

The kernel trick lets an algorithm work in a feature space φ without ever forming φ(x). It does this by:

  • A. computing k(x,z) = φ(x)·φ(z) directly from x and z, in input-space cost (correct)
  • B. approximating φ(x) with a small neural network
  • C. projecting φ(x) onto its top principal component
  • D. discarding the feature space and using raw x·z instead

Answer: A

Why: A kernel returns the exact inner product φ(x)·φ(z) as a cheap function of the raw inputs. Algorithms that use features only through dot products (SVM, ridge, PCA) then run in the feature space at O(d) cost per pair.

Why B tempts people
No approximation is involved — (x·z+c)² equals φ(x)·φ(z) exactly, and the RBF series converges to the true value. A kernel is an identity, not an estimate.
Why C tempts people
That is PCA-style compression to one direction, unrelated. The kernel preserves the full inner product, keeping all (even infinitely many) feature coordinates.
Why D tempts people
Using raw x·z is just the linear kernel — no lift, so XOR still fails. The trick keeps the rich feature space implicitly, it does not discard it.

94. Answer it before you see the options: Check yourself — Mercer's condition

Prediction

Predict first

By Mercer's theorem, k(x,z) is a valid kernel if and only if:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: its Gram matrix is positive semidefinite for every finite dataset

Why: Validity means k = φ(x)·φ(z) for some φ, which holds exactly when every Gram matrix Kᵢⱼ = k(xᵢ,xⱼ) is PSD (min eigenvalue ≥ 0). That is Mercer's theorem.

95. Check yourself — Mercer's condition

Check

When is a candidate k a legal kernel?

Check your understanding

By Mercer's theorem, k(x,z) is a valid kernel if and only if:

  • A. its Gram matrix is positive semidefinite for every finite dataset (correct)
  • B. it is symmetric, i.e. k(x,z) = k(z,x)
  • C. it is non-negative for all x, z
  • D. it equals the Euclidean distance ‖x − z‖

Answer: A

Why: Validity means k = φ(x)·φ(z) for some φ, which holds exactly when every Gram matrix Kᵢⱼ = k(xᵢ,xⱼ) is PSD (min eigenvalue ≥ 0). That is Mercer's theorem.

Why B tempts people
Symmetry is necessary but not sufficient. k(x,z)=x·z−‖x−z‖² is symmetric yet its Gram has eigenvalue −14.13, so it fails Mercer.
Why C tempts people
Kernel values can be negative — the linear kernel x·z is often negative — so non-negativity is neither required nor sufficient.
Why D tempts people
Distance is not an inner product (a point has zero distance to itself but maximal self-similarity), so it is not a kernel. RBF puts distance INSIDE an exponential to build a valid one.

96. Answer it before you see the options: Check yourself — composition

Prediction

Predict first

If k₁ and k₂ are valid kernels, which of these is guaranteed to be a valid kernel?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: k₁ + k₂ (and also k₁·k₂, and c·k₁ for c > 0)

Why: Sums, elementwise products, and positive scalings of valid kernels keep the Gram matrix PSD, so they stay valid. This is what lets you compose kernels freely.

97. Check yourself — composition

Check

Which combinations are guaranteed to stay valid?

Check your understanding

If k₁ and k₂ are valid kernels, which of these is guaranteed to be a valid kernel?

  • A. k₁ + k₂ (and also k₁·k₂, and c·k₁ for c > 0) (correct)
  • B. k₁ − k₂
  • C. 1 / k₁
  • D. −k₁

Answer: A

Why: Sums, elementwise products, and positive scalings of valid kernels keep the Gram matrix PSD, so they stay valid. This is what lets you compose kernels freely.

Why B tempts people
A difference of PSD matrices need not be PSD — we measured min eig −3.3777 for k_lin − k_rbf, so it can produce a negative eigenvalue and is not guaranteed valid.
Why C tempts people
The reciprocal of a kernel has no PSD guarantee — there is no such composition rule, and it typically breaks Mercer.
Why D tempts people
Negating flips a PSD matrix to negative semidefinite, so −k₁ fails the min-eigenvalue-≥-0 test immediately.

98. Rule out three: Check yourself — why kernelize

Elimination

Eliminate the wrong options

You want RBF-kernel ridge regression on n = 500 points. Why solve the dual (K + λI)α = y instead of the primal (ΦᵀΦ + λI)w = Φᵀy?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The RBF feature space is infinite-dimensional, so the primal D×D system can't be formed, but the dual is just a 500×500 solve
  • B. The dual gives a different, more accurate answer than the primal
  • C. The dual automatically regularizes, so no λ is needed
  • D. The primal is only valid for linear kernels

Survives elimination: A

Why: RBF's φ is infinite-dimensional, so ΦᵀΦ (D×D) is infinite and unbuildable. The dual replaces every φ(xᵢ)·φ(xⱼ) with k(xᵢ,xⱼ), giving a finite n×n Gram matrix K — a 500×500 solve. Kernelizing trades feature dimension for dataset size.

99. Check yourself — why kernelize

Check

Think about the RBF feature dimension.

Check your understanding

You want RBF-kernel ridge regression on n = 500 points. Why solve the dual (K + λI)α = y instead of the primal (ΦᵀΦ + λI)w = Φᵀy?

  • A. The RBF feature space is infinite-dimensional, so the primal D×D system can't be formed, but the dual is just a 500×500 solve (correct)
  • B. The dual gives a different, more accurate answer than the primal
  • C. The dual automatically regularizes, so no λ is needed
  • D. The primal is only valid for linear kernels

Answer: A

Why: RBF's φ is infinite-dimensional, so ΦᵀΦ (D×D) is infinite and unbuildable. The dual replaces every φ(xᵢ)·φ(xⱼ) with k(xᵢ,xⱼ), giving a finite n×n Gram matrix K — a 500×500 solve. Kernelizing trades feature dimension for dataset size.

Why B tempts people
When both are solvable they give identical predictions (we verified allclose = True). The dual isn't more accurate — it's just tractable when D is huge.
Why C tempts people
The dual still needs λ — it appears explicitly as (K + λI). Kernelizing changes which matrix you invert, not whether you regularize.
Why D tempts people
The primal works for any explicit feature map, including polynomial. Its only failure is when D is infinite (RBF) or ≫ n, which is precisely when you switch to the dual.

100. Your turn: kernels in code

Section

Part 8 of 8 — the project

101. Project: verify a kernel & crack XOR

Concept

Three milestones, each proving one pillar of the lesson yourself: the polynomial kernel equals its features, the RBF Gram is PSD, and the RBF SVM separates XOR while a linear one can't.

#requirementtool
1(x·z+c)² == φ(x)·φ(z)explicit φ, np.isclose
2RBF Gram matrix is PSDnp.linalg.eigvalsh
3RBF SVM solves XOR, linear failssklearn SVC

Build rules: type every line yourself, run after each milestone, and when the Gram's min eigenvalue looks negative, re-check the sign convention in your distance formula before blaming the kernel.

102. Break it if you can: Project: verify a kernel & crack XOR

Counterexample

Discussion prompt

Three milestones, each proving one pillar of the lesson yourself: the polynomial kernel equals its features, the RBF Gram is PSD, and the RBF SVM separates XOR while a linear one can't.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, run after each milestone, and when the Gram's min eigenvalue looks negative, re-check the sign convention in your distance formula before blaming the kernel.

103. Milestone 1 — kernel equals features

Worked example

Your turn: confirm (a·b+1)² equals φ(a)·φ(b) for the quadratic feature map. Predict out loud: exact match, or off by round-off?

Hint: φ(x) = [x₁², x₂², √2·x₁x₂, √2c·x₁, √2c·x₂, c]. Use np.isclose, not ==, because √2 introduces floating-point dust.

import numpy as np
def phi(x, c=1.0):
    x1, x2 = x
    return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
                     np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
print((a@b+1)**2, phi(a)@phi(b))
computationvalue
(a·b+1)²4.0
φ(a)·φ(b)3.9999999999999987
np.iscloseTrue

104. Milestone 2 — RBF Gram is PSD

Worked example

Your turn: build an RBF Gram matrix on 6 seeded points and confirm its minimum eigenvalue is ≥ 0. Predict: valid kernel?

Hint: squared distances via ‖x‖² + ‖z‖² − 2x·z, then exp(−sq/2σ²); read np.linalg.eigvalsh(K).min().

import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(6, 2))
sq = np.sum(X**2,1)[:,None] + np.sum(X**2,1)[None,:] - 2*X@X.T
K = np.exp(-sq/2.0)
print(round(np.linalg.eigvalsh(K).min(), 6))
checkvalue
min eigenvalue0.02634
≥ 0 ?yes
valid kernel (PSD)?yes

105. Milestone 3 — RBF SVM on XOR

Worked example

Your turn: fit a linear and an RBF SVM on XOR and score each. Predict both accuracies before you print.

Hint: SVC(kernel='linear') vs SVC(kernel='rbf', gamma=1.0, C=10), then .score(X, y).

import numpy as np
from sklearn.svm import SVC
X = np.array([[0,0],[0,1],[1,0],[1,1]], float)
y = np.array([0, 1, 1, 0])
print(SVC(kernel='linear').fit(X, y).score(X, y),
      SVC(kernel='rbf', gamma=1.0, C=10).fit(X, y).score(X, y))
kernelaccuracy
linear0.5
RBF1.0

106. What each one costs: Milestone 3 — RBF SVM on XOR

Trade off

Comparison matrix

From Milestone 3 — RBF SVM on XOR: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.

kernelaccuracy
linear0.5
RBF1.0

107. The full program

Concept

import numpy as np
from sklearn.svm import SVC

# 1. polynomial kernel == explicit quadratic features
def phi(x, c=1.0):
    x1, x2 = x
    return np.array([x1**2, x2**2, np.sqrt(2)*x1*x2,
                     np.sqrt(2*c)*x1, np.sqrt(2*c)*x2, c])
a = np.array([1., 2.]); b = np.array([3., -1.])
print(np.isclose((a@b+1)**2, phi(a)@phi(b)))

# 2. RBF Gram is PSD
rng = np.random.default_rng(0)
G = rng.normal(size=(6, 2))
sq = np.sum(G**2,1)[:,None] + np.sum(G**2,1)[None,:] - 2*G@G.T
K = np.exp(-sq/2.0)
print(round(np.linalg.eigvalsh(K).min(), 4), '>= 0')

# 3. RBF SVM cracks XOR, linear fails
X = np.array([[0,0],[0,1],[1,0],[1,1]], float); y = np.array([0,1,1,0])
print(SVC(kernel='linear').fit(X,y).score(X,y),
      SVC(kernel='rbf', gamma=1.0, C=10).fit(X,y).score(X,y))
printed linevalue
kernel == featuresTrue
RBF min eig0.0263 >= 0
linear / RBF on XOR0.5 1.0

If the kernel matches its features, the RBF Gram is PSD, and the RBF SVM cracks XOR while linear can't — you have used an infinite feature space without ever visiting it.

108. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
kernel == featuresTrue
RBF min eig0.0263 >= 0
linear / RBF on XOR0.5 1.0

109. Show it off

Concept

Slides closed, out loud: explain (1) what a kernel avoids computing and why that matters for RBF, (2) why Mercer's PSD condition is the exact right test, and (3) how the RBF SVM separates XOR that a linear one cannot.

Stretch: prove the composition product rule keeps the Gram PSD (Schur product theorem), then re-derive the kernel ridge dual (K+λI)α = y and confirm its predictions match explicit-feature ridge. Kernels return as the SVM dual (Week 25) and as the covariance function of a Gaussian process.

110. Connect it up: Lesson 28: Kernel Methods

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The problem: feature maps · The kernel trick · Mercer's theorem · RBF = infinite dimensions · Composing kernels · Kernelizing ridge. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

111. What you can do now

Recap

ideathe one thing to remember
kernel trickk(x,z) = φ(x)·φ(z), computed in O(d)
polynomial kernel(x·z+c)² carries the whole quadratic φ
Mercervalid ⟺ Gram matrix PSD (min eig ≥ 0)
RBFexp(−‖x−z‖²/2σ²), infinite-dimensional φ
compositionsum, product, positive scaling stay valid; difference does not
kernelizefeatures only via dot products → replace with k; solve the dual

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 28 (Week 10 — Kernel Methods) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn SVC (kernels)
  3. Bishop, Pattern Recognition and Machine Learning, Ch. 6 (Kernel Methods) — Springer, 2006
  4. Polynomial-kernel identity, RBF Gram PSD, fake-kernel negative eigenvalue, RBF series convergence, kernel-ridge dual == primal, and XOR SVM all produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108