USAAIO Lesson 30, from Week 10, fully worked. It derives the geometric margin as 2/‖w‖, sets up the primal max-margin QP, and writes out the Lagrangian and the KKT conditions. It then derives the dual QP in α term by term and solves that dual BY HAND on a four-point set, giving α=[0,¼,¼,0], explains why only the support vectors survive, and rebuilds the decision function from the duals. It measures the soft-margin C trade-off and closes with the kernel trick on XOR. One running four-point dataset carries the whole deck; every snippet runs standalone, and every number came from real execution. The lesson runs to 63 slides.
Subject: Machine Learning · 118 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 30 · Week 10
The maximum-margin classifier, where only a handful of points matter. We derive the margin 2/‖w‖, build the primal QP, pass to the dual through the Lagrangian, then solve that dual by hand on four points and rebuild f(x) from the multipliers alone.
Objectives
2/‖w‖ and state the primal max-margin QP with slack ξᵢα step by stepα = [0, ¼, ¼, 0], w = [½, ½], b = −2αᵢ > 0) enter f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and rebuild f from the dualsC as a bias–variance knob and kernelize (RBF) to bend the boundary — all matched against sklearnWarm-up
Discussion prompt
Before we open Lesson 30: Support Vector Machines: without looking back, what was the main idea of Mutual Information & Contrastive Learning, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
mutual information I(X;Y) as H(X)+H(Y)−H(X,Y) and as KL(joint‖product), the independence test I=0, feature selection by mutual information, the InfoNCE contrastive loss (CLIP), and the link to compression. Compute MI three equivalent ways and rank features with sklearn.
Section
Part 1 of 6 — the primal
Concept
Four points in the plane, two per class. Class −1 sits low-left, class +1 high-right. We want the boundary line that separates them with the widest gap.
| point | x₁ | x₂ | label y |
|---|---|---|---|
| p₀ | 0 | 0 | −1 |
| p₁ | 1 | 1 | −1 |
| p₂ | 3 | 3 | +1 |
| p₃ | 4 | 4 | +1 |
Keep this set in your head — every derivation, hand-solve, and code block in the lesson uses these exact four points.
Comparison
Comparison matrix
From The data: four labelled points: refill the x₁ column from what you know. The rest of the table is as it appeared.
| point | x₁ | x₂ | label y |
|---|---|---|---|
| p₀ | 0 | 0 | −1 |
| p₁ | 1 | 1 | −1 |
| p₂ | 3 | 3 | +1 |
| p₃ | 4 | 4 | +1 |
Intuition
Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.
The closest opposing pair here is p₁ = (1,1) and p₂ = (3,3). The street's centerline is their perpendicular bisector; the curbs just touch those two points.
The far points p₀ and p₃ sit well back from the curb. Intuition to test later: they should have no say in where the street goes.
Counterexample
Discussion prompt
Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The closest opposing pair here is p₁ = (1,1) and p₂ = (3,3). The street's centerline is their perpendicular bisector; the curbs just touch those two points.
Concept
A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.
\[ \hat y = \operatorname{sign}(w^\top x + b), \qquad \text{boundary: } w^\top x + b = 0 \]
w is perpendicular to the boundary (it is the boundary's normal). b shifts the boundary off the origin. Our job is to choose w and b so the gap around the boundary is as wide as possible.
Analogy
Discussion prompt
Explain A linear classifier and its boundary by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.
Intuition
Take any two points xₐ, x_b that both sit on the boundary: wᵀxₐ + b = 0 and wᵀx_b + b = 0. Subtract to get wᵀ(xₐ − x_b) = 0.
xₐ − x_b is a direction along the boundary, and its dot with w is zero — so w is perpendicular to every in-boundary direction. w points straight across the street.
That is why the perpendicular distance formula uses w/‖w‖: moving along w is the fastest way off the boundary, and the margin is measured in exactly that direction.
Explain it
Discussion prompt
Explain Why w is perpendicular to the boundary to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Take any two points xₐ, x_b that both sit on the boundary: wᵀxₐ + b = 0 and wᵀx_b + b = 0. Subtract to get wᵀ(xₐ − x_b) = 0.
Concept
Call yᵢ(wᵀxᵢ + b) the functional margin of point i. It is positive when the point is on the correct side, and larger when the point is more confidently classified.
But it has a flaw: scaling (w, b) → (2w, 2b) doubles every functional margin without moving the boundary an inch. Size alone is meaningless until we pin the scale.
Fix the scale: force the closest points to have functional margin exactly 1
Why: This is the canonical normalization. It removes the ambiguity by definition, so 'nearest points' means yᵢ(wᵀxᵢ+b) = 1. Everything downstream depends on this convention.
Ranking
Put in order
Put the moves of Derive the geometric margin = 1/‖w‖ into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Projecting the offset onto the unit normal w/‖w‖ gives the signed perpendicular distance.
Worked example
The geometric margin is the true perpendicular distance from a point to the boundary. Distance from x to the hyperplane wᵀx + b = 0 is the standard point-to-plane formula:
Distance from a point to the hyperplane
Why: Projecting the offset onto the unit normal w/‖w‖ gives the signed perpendicular distance.
\[ \text{dist}(x_i) = \frac{\lvert w^\top x_i + b\rvert}{\lVert w\rVert} \]
For a nearest point, the numerator is exactly 1
Why: Our canonical scaling forces yᵢ(wᵀxᵢ+b) = 1 at the closest points, so |wᵀxᵢ+b| = 1 there.
\[ \gamma = \frac{1}{\lVert w\rVert} \]
The full street width (curb to curb) is 2/‖w‖
Why: One margin γ = 1/‖w‖ on each side of the centerline, so the total separation between the +1 and −1 margin lines is 2/‖w‖. Verified on our fit: ‖w‖ = 0.7071, so 2/‖w‖ = 2.828 and 1/‖w‖ = 1.414.
Notation
Annotate
From Derive the geometric margin = 1/‖w‖ — read this one piece at a time. What is each part doing?
On: \( \text{dist}(x_i) = \frac{\lvert w^\top x_i + b\rvert}{\lVert w\rVert} \)
Concept
We want the widest street, so maximize 2/‖w‖. A fraction with fixed numerator is largest when its denominator is smallest — so maximizing the margin is the same as minimizing ‖w‖.
Minimize ½‖w‖² instead of ‖w‖
Why: ‖w‖ has an ugly square root; ½‖w‖² = ½wᵀw is smooth and convex with the same minimizer. The ½ is cosmetic — it cancels the 2 when we differentiate. This is exactly the OLS-style trick from Lesson 7.
\[ \max_{w,b}\ \frac{2}{\lVert w\rVert} \quad\Longleftrightarrow\quad \min_{w,b}\ \tfrac{1}{2}\lVert w\rVert^2 \]
Concept
Add the constraint that every point is correctly classified with functional margin at least 1. That gives the hard-margin primal — a convex quadratic program:
\[ \min_{w,b}\ \tfrac{1}{2}\lVert w\rVert^2 \quad \text{s.t. } y_i(w^\top x_i + b) \ge 1 \ \ \forall i \]
Quadratic objective, linear inequality constraints — a textbook QP with a unique optimum when the data is separable. Real data usually isn't perfectly separable, so next we let the constraints bend.
Concept
Introduce a slack ξᵢ ≥ 0 per point: how far point i is allowed to intrude past its margin. Charge C per unit of total slack and add it to the objective.
\[ \min_{w,b,\xi}\ \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \quad \text{s.t. } y_i(w^\top x_i + b) \ge 1 - \xi_i,\ \ \xi_i \ge 0 \]
C = ∞ forbids all slack (back to hard margin). Finite C trades a wider margin against a few violations — the knob we'll tune in Part 5. For the hand-solve we keep it separable, so slack stays 0.
Estimation
Predict first
The dual needs every pairwise dot product Kᵢⱼ = xᵢᵀxⱼ. All four points lie on the line x₁ = x₂, so xᵢᵀxⱼ = (aᵢ)(aⱼ) + (aᵢ)(aⱼ) = 2 aᵢ aⱼ where a = (0, 1, 3, 4) is the shared coordinate.
Commit before you compute: what does The dot products, by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Row p₂: K₂₀=0, K₂₁=6, K₂₂=18, K₂₃=24
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 2·a₂·aⱼ = 2·3·(0,1,3,4) = (0, 6, 18, 24).
Worked example
The dual needs every pairwise dot product Kᵢⱼ = xᵢᵀxⱼ. All four points lie on the line x₁ = x₂, so xᵢᵀxⱼ = (aᵢ)(aⱼ) + (aᵢ)(aⱼ) = 2 aᵢ aⱼ where a = (0, 1, 3, 4) is the shared coordinate.
Row p₁: K₁₀=0, K₁₁=2, K₁₂=6, K₁₃=8
Why: 2·a₁·aⱼ = 2·1·(0,1,3,4) = (0, 2, 6, 8). Anything times a₀ = 0 vanishes.
Row p₂: K₂₀=0, K₂₁=6, K₂₂=18, K₂₃=24
Why: 2·a₂·aⱼ = 2·3·(0,1,3,4) = (0, 6, 18, 24).
| K | p₀ | p₁ | p₂ | p₃ |
|---|---|---|---|---|
| p₀ | 0 | 0 | 0 | 0 |
| p₁ | 0 | 2 | 6 | 8 |
| p₂ | 0 | 6 | 18 | 24 |
| p₃ | 0 | 8 | 24 | 32 |
Trade off
Comparison matrix
From The dot products, by hand: every row here is a choice with a cost. Fill the p₂ column, then say which row you would actually pick and what you give up for it.
| K | p₀ | p₁ | p₂ | p₃ |
|---|---|---|---|---|
| p₀ | 0 | 0 | 0 | 0 |
| p₁ | 0 | 2 | 6 | 8 |
| p₂ | 0 | 6 | 18 | 24 |
| p₃ | 0 | 8 | 24 | 32 |
Missing information
Discussion prompt
Now weight each entry by the labels: the labelled Gram is Gᵢⱼ = yᵢyⱼKᵢⱼ. Confirm the hand table in code — this block re-defines the data and runs on its own:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Any dot product with the origin is 0, so p₀ contributes nothing to K or G. A quiet hint that p₀ will be inert in the dual.
Worked example
Now weight each entry by the labels: the labelled Gram is Gᵢⱼ = yᵢyⱼKᵢⱼ. Confirm the hand table in code — this block re-defines the data and runs on its own:
import numpy as np
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
K = X @ X.T # linear kernel = dot products
G = (y[:,None] * y[None,:]) * K # labelled Gram: G_ij = y_i y_j (x_i . x_j)
print(K)
print(G)Row/col 0 is all zeros because p₀ = (0,0)
Why: Any dot product with the origin is 0, so p₀ contributes nothing to K or G. A quiet hint that p₀ will be inert in the dual.
| p₀ | p₁ | p₂ | p₃ | |
|---|---|---|---|---|
| G (row p₀) | 0 | 0 | 0 | 0 |
| G (row p₁) | 0 | 2 | −6 | −8 |
| G (row p₂) | 0 | −6 | 18 | 24 |
| G (row p₃) | 0 | −8 | 24 | 32 |
Invariant
Step through it
Step through Build the labelled Gram matrix one row at a time. One of these columns never changes — find it, and say why it cannot.
Section
Part 2 of 6 — every step
Intuition
The primal is solvable, so why change variables? Two payoffs the exam cares about.
First, the dual exposes sparsity: its solution assigns a multiplier αᵢ to each point, and almost all come out exactly zero. The nonzero few are the support vectors.
Second, in the dual the data appears only through dot products xᵢᵀxⱼ. Replace each dot product with a kernel and you get nonlinear boundaries for free — the kernel trick.
Concept
Each inequality yᵢ(wᵀxᵢ + b) ≥ 1 becomes 1 − yᵢ(wᵀxᵢ + b) ≤ 0. Attach a multiplier αᵢ ≥ 0 to each and subtract it from the objective — the standard Lagrangian recipe from Lesson 24.
\[ \mathcal{L}(w,b,\alpha) = \tfrac{1}{2}\lVert w\rVert^2 - \sum_i \alpha_i\big[\,y_i(w^\top x_i + b) - 1\,\big], \qquad \alpha_i \ge 0 \]
The dual comes from minimizing L over (w, b) for fixed α, then maximizing over α. Minimizing over w and b means setting their gradients to zero — the next two slides.
Step zero
Discussion prompt
Stationarity in w — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Differentiate L with respect to w and set to 0
Answer:
Worked example
Differentiate L with respect to w and set to 0
Why: ∂(½‖w‖²)/∂w = w, and ∂(−Σ αᵢ yᵢ wᵀxᵢ)/∂w = −Σ αᵢ yᵢ xᵢ. The b-term and the +1 have no w.
\[ \nabla_w \mathcal{L} = w - \sum_i \alpha_i y_i x_i = 0 \]
Solve for w
Why: This is the key structural result: the optimal weight vector is a labelled, α-weighted sum of the training points. w lives in the span of the data.
\[ \boxed{\,w = \sum_i \alpha_i y_i x_i\,} \]
Read off the consequence
Why: Any point with αᵢ = 0 drops out of w entirely. Only points with αᵢ > 0 — the support vectors — build the weight vector. Sparsity is baked in right here.
Blank canvas
Draw it
Draw what Stationarity in w just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From Stationarity in b — finish the line. Write what belongs on the right of the equals sign before you look.
\boxed0\,}}
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Only the term −Σ αᵢ yᵢ b contains b, and ∂/∂b of that is −Σ αᵢ yᵢ.
Worked example
Differentiate L with respect to b and set to 0
Why: Only the term −Σ αᵢ yᵢ b contains b, and ∂/∂b of that is −Σ αᵢ yᵢ.
\[ \frac{\partial \mathcal{L}}{\partial b} = -\sum_i \alpha_i y_i = 0 \]
This is the dual equality constraint
Why: The α's are not free — they must satisfy Σ αᵢ yᵢ = 0. It ties the positive-class multipliers to the negative-class ones.
\[ \boxed{\,\sum_i \alpha_i y_i = 0\,} \]
Worked example
Put w = Σ αᵢyᵢxᵢ back into L and use Σ αᵢyᵢ = 0 to kill the b term. Expand ½‖w‖² first:
½‖w‖² = ½ wᵀw with w = Σ αᵢyᵢxᵢ
Why: A dot of two sums becomes a double sum; each xᵢᵀxⱼ is a dot product.
\[ \tfrac{1}{2}\lVert w\rVert^2 = \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, x_i^\top x_j \]
The cross term −Σ αᵢyᵢ wᵀxᵢ equals −‖w‖² = −Σ αᵢαⱼyᵢyⱼxᵢᵀxⱼ
Why: Because wᵀxᵢ summed against αᵢyᵢ rebuilds wᵀw exactly. And the +Σαᵢ term survives untouched.
\[ \mathcal{L} = \tfrac{1}{2}\!\sum_{i,j}\!\alpha_i\alpha_j y_i y_j x_i^\top x_j - \!\sum_{i,j}\!\alpha_i\alpha_j y_i y_j x_i^\top x_j + \sum_i \alpha_i \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
The cross term −Σ αᵢyᵢ wᵀxᵢ equals −‖w‖² = −Σ αᵢαⱼyᵢyⱼxᵢᵀxⱼ
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Put w = Σ αᵢyᵢxᵢ back into L and use Σ αᵢyᵢ = 0 to kill the b term. Expand ½‖w‖² first:
Step zero
Discussion prompt
Collect terms → the dual QP — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Combine the two double sums: ½ − 1 = −½
Answer:
Worked example
Combine the two double sums: ½ − 1 = −½
Why: The identical double sums differ only by their coefficients ½ and −1, which add to −½. That flips the quadratic term negative — the dual is a MAXIMIZATION.
\[ \max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, x_i^\top x_j \]
Carry the constraints α ≥ 0 and Σ αᵢyᵢ = 0
Why: α ≥ 0 came with the multipliers; Σ αᵢyᵢ = 0 came from ∂L/∂b. Together they bound the QP.
\[ \text{s.t. } \alpha_i \ge 0,\qquad \sum_i \alpha_i y_i = 0 \]
Replace xᵢᵀxⱼ with a kernel k(xᵢ,xⱼ)
Why: The data touches the dual ONLY through these dot products. Swapping in any valid kernel gives a nonlinear SVM without ever changing the optimizer — this is the kernel trick.
\[ \boxed{\,\max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j)\,} \]
Translation
\( \boxed{\,\max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j)\,} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
Redo the same Lagrangian with the slack terms C Σ ξᵢ and the ξᵢ ≥ 0 multipliers. The algebra is identical, and the slack multipliers collapse into a single change: each αᵢ gains an upper bound C.
\[ \max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j) \quad \text{s.t. } 0 \le \alpha_i \le C,\ \sum_i \alpha_i y_i = 0 \]
So the whole soft-margin story lives in one inequality: 0 ≤ αᵢ ≤ C. Hard margin is the C → ∞ limit, where the cap disappears.
Concept
Complementary slackness is the KKT condition that decides which points matter: for each i, αᵢ[yᵢ(wᵀxᵢ+b) − 1 + ξᵢ] = 0. Either the multiplier is zero, or the margin constraint is tight.
| value of αᵢ | what it means geometrically |
|---|---|
| αᵢ = 0 | not a support vector — strictly outside the margin, correct |
| 0 < αᵢ < C | support vector exactly ON the margin (ξᵢ = 0) |
| αᵢ = C | support vector inside the margin or misclassified (ξᵢ > 0) |
Read the trained α and you instantly know each point's status. We measure all three cases in code in Part 5.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The primal is a min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all.
Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.
Why: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.
Trap
The primal is a min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....
min over α of Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk
Why: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.
Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.
max over α of Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk
Why: Correct. It is a concave maximization (equivalently, minimize its negative — which is what solvers do). On our data the max is W = 0.25 at α = [0, ¼, ¼, 0]. Duality: this equals the primal ½‖w‖² = 0.25.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.
Section
Part 3 of 6 — the 4 points
Worked example
Four αᵢ looks hard, but the constraints and geometry collapse it to one variable. First, guess (from the widest-street picture) that only the closest pair p₁, p₂ are support vectors, so α₀ = α₃ = 0. We'll verify this holds.
Apply Σ αᵢyᵢ = 0 to the survivors
Why: With α₀ = α₃ = 0, the equality constraint reads −α₁ + α₂ = 0 (labels y₁ = −1, y₂ = +1).
\[ -\alpha_1 + \alpha_2 = 0 \;\Longrightarrow\; \alpha_1 = \alpha_2 = t \]
One unknown t left
Why: The four-variable QP is now a single-variable concave parabola in t. Maximize it with one derivative.
Ranking
Put in order
Put the moves of Plug the Gram entries into the objective into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p₁ = (1,1), so x₁·x₁ = 1²+1² = 2, and y₁² = 1.
Worked example
Write the dual with only α₁ = α₂ = t. We need three Gram entries: G₁₁, G₂₂, G₁₂.
G₁₁ = y₁²(x₁·x₁) = 1·(1+1) = 2
Why: p₁ = (1,1), so x₁·x₁ = 1²+1² = 2, and y₁² = 1.
G₂₂ = y₂²(x₂·x₂) = 1·(9+9) = 18
Why: p₂ = (3,3), so x₂·x₂ = 3²+3² = 18.
G₁₂ = y₁y₂(x₁·x₂) = (−1)(+1)(3+3) = −6
Why: x₁·x₂ = 1·3 + 1·3 = 6, times the opposite labels gives −6. These match the Gram matrix printed in Part 1.
Step zero
Discussion prompt
Reduce to a one-variable parabola — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Objective with α₁ = α₂ = t
Answer:
Worked example
Objective with α₁ = α₂ = t
Why: Σαᵢ = t + t = 2t. The double sum expands to α₁²G₁₁ + α₂²G₂₂ + 2α₁α₂G₁₂ = t²(2) + t²(18) + 2t²(−6).
\[ W(t) = 2t - \tfrac{1}{2}\big[\,2t^2 + 18t^2 + 2t^2(-6)\,\big] \]
Collect the bracket: 2 + 18 − 12 = 8
Why: The three quadratic coefficients sum to 8, so the bracket is 8t².
\[ W(t) = 2t - \tfrac{1}{2}(8t^2) = 2t - 4t^2 \]
A downward parabola in t
Why: Coefficient of t² is −4 < 0, so W(t) is concave with a single interior maximum — exactly what a well-posed dual should be.
Sorting
Sort into buckets
These are the pieces of Lesson 30: Support Vector Machines, out of order. Put each one back under the part of the lesson it belongs to.
Step zero
Discussion prompt
Maximize: one derivative — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Set dW/dt = 0
Answer:
Worked example
Set dW/dt = 0
Why: d/dt(2t − 4t²) = 2 − 8t. The maximum of the concave parabola is where the slope vanishes.
\[ \frac{dW}{dt} = 2 - 8t = 0 \]
Solve for t
Why: 8t = 2, so t = ¼. Both support-vector multipliers equal 0.25.
\[ t = \tfrac{1}{4} \;\Longrightarrow\; \alpha = \big[\,0,\ \tfrac14,\ \tfrac14,\ 0\,\big] \]
Check α ≥ 0 and the value
Why: t = ¼ > 0 satisfies the bounds. Dual value W(¼) = 2(¼) − 4(¼)² = 0.5 − 0.25 = 0.25. We'll confirm every digit against a QP solver next.
Pattern
Predict first
The table runs: α₀, α₃ | 0, 0 | 0.0, 0.0 · α₁, α₂ | ¼, ¼ | 0.25, 0.25
In Verify the hand dual against a QP solver, given the rows so far: what is the next one — the row where quantity is dual value W?
Correct: dual value W | 0.25 | 0.25
| quantity | by hand | solver (verified) |
|---|---|---|
| α₀, α₃ | 0, 0 | 0.0, 0.0 |
| α₁, α₂ | ¼, ¼ | 0.25, 0.25 |
| dual value W | 0.25 | 0.25 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Exactly the hand answer. α₀ = α₃ = 0 confirms p₀ and p₃ are NOT support vectors — our guess was right, not assumed.
Worked example
Solve the same dual numerically with scipy.optimize.minimize (minimizing −W). This block re-defines the data and runs on its own:
import numpy as np
from scipy.optimize import minimize
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
G = (y[:,None]*y[None,:]) * (X @ X.T)
neg_W = lambda a: -(a.sum() - 0.5 * a @ G @ a) # minimize -W
cons = [{'type':'eq', 'fun': lambda a: a @ y}] # sum a_i y_i = 0
res = minimize(neg_W, np.full(4, 0.1), bounds=[(0,None)]*4, constraints=cons)
alpha = np.where(res.x < 1e-7, 0.0, res.x)
print("alpha =", np.round(alpha, 4))
print("dual value W =", round(-res.fun, 4))Solver returns α = [0, 0.25, 0.25, 0], W = 0.25
Why: Exactly the hand answer. α₀ = α₃ = 0 confirms p₀ and p₃ are NOT support vectors — our guess was right, not assumed.
| quantity | by hand | solver (verified) |
|---|---|---|
| α₀, α₃ | 0, 0 | 0.0, 0.0 |
| α₁, α₂ | ¼, ¼ | 0.25, 0.25 |
| dual value W | 0.25 | 0.25 |
Pattern
Step through it
Step through Verify the hand dual against a QP solver one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
Put the moves of Recover w from the multipliers into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Only p₁ and p₂ have nonzero α. So w = α₁y₁x₁ + α₂y₂x₂ = ¼(−1)(1,1) + ¼(+1)(3,3).
Worked example
Use w = Σ αᵢyᵢxᵢ
Why: Only p₁ and p₂ have nonzero α. So w = α₁y₁x₁ + α₂y₂x₂ = ¼(−1)(1,1) + ¼(+1)(3,3).
\[ w = \tfrac14(-1)\!\begin{bmatrix}1\\1\end{bmatrix} + \tfrac14(+1)\!\begin{bmatrix}3\\3\end{bmatrix} = \begin{bmatrix}-\tfrac14 + \tfrac34\\[2pt] -\tfrac14 + \tfrac34\end{bmatrix} \]
Simplify each component
Why: −¼ + ¾ = ½ in both coordinates.
\[ w = \begin{bmatrix} \tfrac12 \\[2pt] \tfrac12 \end{bmatrix} \]
Sanity: ‖w‖ and the margin
Why: ‖w‖ = √(¼+¼) = √½ ≈ 0.7071, so the geometric margin 1/‖w‖ ≈ 1.414 and the full street 2/‖w‖ ≈ 2.828 — matching the Part-1 formula.
Blank canvas
Draw it
Draw what Recover w from the multipliers just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Worked example
A support vector sits exactly on its margin: yₛ(wᵀxₛ + b) = 1. Solve for b using p₁:
Plug in p₁ = (1,1), y₁ = −1
Why: wᵀx₁ = ½·1 + ½·1 = 1. So −1·(1 + b) = 1, giving 1 + b = −1.
\[ y_1(w^\top x_1 + b) = 1 \;\Longrightarrow\; -(1 + b) = 1 \;\Longrightarrow\; b = -2 \]
Cross-check with p₂ = (3,3), y₂ = +1
Why: wᵀx₂ = ½·3 + ½·3 = 3. Then +1·(3 + b) = 1 gives b = −2 again. Both support vectors agree, so b = −2 is consistent.
| source of b | equation | b |
|---|---|---|
| from p₁ | −(1 + b) = 1 | −2 |
| from p₂ | +(3 + b) = 1 | −2 |
Notation
Annotate
From Recover b from a support vector — read this one piece at a time. What is each part doing?
On: \( y_1(w^\top x_1 + b) = 1 \;\Longrightarrow\; -(1 + b) = 1 \;\Longrightarrow\; b = -2 \)
Worked example
With w = [½, ½], b = −2, compute the functional margin yᵢ(wᵀxᵢ + b) for every point. Support vectors land at exactly 1; non-SVs land strictly above.
| point | wᵀxᵢ | wᵀxᵢ+b | yᵢ(wᵀxᵢ+b) | role |
|---|---|---|---|---|
| p₀ (0,0) | 0 | −2 | 2 | non-SV (α=0) |
| p₁ (1,1) | 1 | −1 | 1 | support vector |
| p₂ (3,3) | 3 | 1 | 1 | support vector |
| p₃ (4,4) | 4 | 2 | 2 | non-SV (α=0) |
SVs at margin 1, non-SVs at margin 2
Why: p₀ and p₃ satisfy the constraint with room to spare (margin 2 > 1), so their multipliers are 0 — they don't touch the curb. This is the geometric meaning of a support vector, computed exactly.
Discrimination
Sort into buckets
Sort these by yᵢ(wᵀxᵢ+b), from memory, without looking back at Full margin trace: who sits where. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Fill the middle
Fill in the blanks
From Verify w, b against sklearn — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("w =", clf.coef_[0], " b =", round(clf.intercept_[0], 4))
print("support idx =", clf.support_, " n_support =", clf.n_support_)
print("dual_coef (a_i*y_i) =", clf.dual_coef_[0])
Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. Identical to the hand answer. Note dual_coef_ stores the SIGNED products αᵢyᵢ = [−0.25, +0.25], i.e.
Worked example
Fit SVC with a huge C (≈ hard margin) and confirm it reproduces our by-hand w, b, and support set. Standalone:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("w =", clf.coef_[0], " b =", round(clf.intercept_[0], 4))
print("support idx =", clf.support_, " n_support =", clf.n_support_)
print("dual_coef (a_i*y_i) =", clf.dual_coef_[0])sklearn: w = [0.5, 0.5], b = −2, SVs = points 1 and 2
Why: Identical to the hand answer. Note dual_coef_ stores the SIGNED products αᵢyᵢ = [−0.25, +0.25], i.e. α₁ = α₂ = 0.25 with the class signs attached.
| quantity | by hand | sklearn (verified) |
|---|---|---|
| w | [½, ½] | [0.5, 0.5] |
| b | −2 | −2.0 |
| support indices | 1, 2 | [1 2] |
| dual_coef_ (αᵢyᵢ) | −¼, +¼ | [−0.25, 0.25] |
Concept
The primal minimized ½‖w‖²; the dual maximized W(α). For a convex QP with these constraints, strong duality holds: the two optima are equal, with no gap.
\[ \underbrace{\tfrac{1}{2}\lVert w^\star\rVert^2}_{\text{primal min}} \;=\; \underbrace{\sum_i \alpha_i - \tfrac12\sum_{i,j}\alpha_i\alpha_j y_i y_j x_i^\top x_j}_{\text{dual max } W(\alpha^\star)} \]
On our data both equal 0.25. That equality is not a coincidence — it is what licenses solving the dual instead of the primal and reading w, b back off the α.
Estimation
Predict first
Compute the primal ½‖w‖² from sklearn's w, and the dual W(α★) from the QP solve, and check they match. Standalone:
Commit before you compute: what does Confirm the duality gap is zero come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: primal = dual = 0.25, gap = 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Strong duality confirmed numerically.
Worked example
Compute the primal ½‖w‖² from sklearn's w, and the dual W(α★) from the QP solve, and check they match. Standalone:
import numpy as np
from scipy.optimize import minimize
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
primal = 0.5 * clf.coef_[0] @ clf.coef_[0] # 1/2 ||w||^2
G = (y[:,None]*y[None,:]) * (X @ X.T)
res = minimize(lambda a: -(a.sum()-0.5*a@G@a), np.full(4,0.1),
bounds=[(0,None)]*4, constraints=[{'type':'eq','fun':lambda a:a@y}])
print("primal 1/2||w||^2 =", round(primal,4))
print("dual W* =", round(-res.fun,4))
print("equal?", np.isclose(primal, -res.fun))primal = dual = 0.25, gap = 0
Why: Strong duality confirmed numerically. The solver's dual optimum lands exactly on the primal's ½‖w‖², so nothing is lost by working in α.
| side | quantity | value (verified) |
|---|---|---|
| primal | ½‖w‖² | 0.25 |
| dual | W(α★) | 0.25 |
| gap | primal − dual | 0.0 |
Section
Part 4 of 6
Concept
Substitute w = Σ αᵢyᵢxᵢ into f(x) = wᵀx + b. The dot product xᵢᵀx becomes a kernel, and the sum runs over support vectors only (the rest have αᵢ = 0):
\[ f(x) = \sum_i \alpha_i y_i\, k(x_i, x) + b = \sum_{i \in SV} \alpha_i y_i\, k(x_i, x) + b \]
The prediction is a weighted vote of the support vectors: each SV casts a vote of sign yᵢ, strength αᵢ, scaled by how similar x is to it via the kernel. Sparse and kernelizable at once.
Faded example
Fill in the blanks
Rebuild f(x) from the duals alone, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5]) # a test point
manual = **np.sum(dc * (sv @ xt)) + b** # sum a_i y_i (x_i . xt) + b
print("manual f(xt) =", round(manual, 4))
print("sklearn f(xt) =", round(clf.decision_function([xt])[0], 4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Σ αᵢyᵢ(xᵢ·xt) + b, using only the 2 support vectors, reproduces sklearn's decision_function exactly — proof that f is fully determined by the duals.
Worked example
Take sklearn's dual_coef_ (the αᵢyᵢ) and support_vectors_, and compute f at a test point (2.5, 2.5) by hand. It must match decision_function. Standalone:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5]) # a test point
manual = np.sum(dc * (sv @ xt)) + b # sum a_i y_i (x_i . xt) + b
print("manual f(xt) =", round(manual, 4))
print("sklearn f(xt) =", round(clf.decision_function([xt])[0], 4))manual = sklearn = 0.5
Why: Σ αᵢyᵢ(xᵢ·xt) + b, using only the 2 support vectors, reproduces sklearn's decision_function exactly — proof that f is fully determined by the duals.
| test point | manual f (from duals) | sklearn decision_function |
|---|---|---|
| (2.5, 2.5) | 0.5 | 0.5 |
| (0, 0) = p₀ | −2.0 | −2.0 |
| (2, 2) on boundary | 0.0 | 0.0 |
Fill the middle
Fill in the blanks
From Deleting a non-support-vector changes nothing — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
full = SVC(kernel='linear', C=1e6).fit(X, y)
keep = [1, 2, 3] # drop point 0 (a NON-SV)
part = SVC(kernel='linear', C=1e6).fit(X[keep], y[keep])
print("full w,b =", full.coef_[0], round(full.intercept_[0],4))
print("drop w,b =", part.coef_[0], round(part.intercept_[0],4))
Why: keep is what everything below it consumes, so the wrong expression here fails later and somewhere else. Identical boundary. The non-support-vector p₀ contributed nothing (α₀ = 0), so its removal is invisible to the model — the SVM's signature sparsity.
Worked example
The claim: removing a point with αᵢ = 0 leaves the boundary identical. Drop p₀ (a non-SV) and refit. Standalone:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
full = SVC(kernel='linear', C=1e6).fit(X, y)
keep = [1, 2, 3] # drop point 0 (a NON-SV)
part = SVC(kernel='linear', C=1e6).fit(X[keep], y[keep])
print("full w,b =", full.coef_[0], round(full.intercept_[0],4))
print("drop w,b =", part.coef_[0], round(part.intercept_[0],4))Both fits give w = [0.5, 0.5], b = −2
Why: Identical boundary. The non-support-vector p₀ contributed nothing (α₀ = 0), so its removal is invisible to the model — the SVM's signature sparsity.
| dataset | w | b |
|---|---|---|
| all 4 points | [0.5, 0.5] | −2.0 |
| 3 points (p₀ dropped) | [0.5, 0.5] | −2.0 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
All four points went into the fit, so deleting any of them would move the boundary.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x).
Only the support vectors (αᵢ > 0) shape the boundary; the rest are irrelevant to it.
Why: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x). We just deleted p₀ and got the identical [0.5, 0.5], −2. Non-SVs are dead weight to the model.
Trap
All four points went into the fit, so deleting any of them would move the boundary.
Assume dropping p₀ or p₃ shifts w, b
Why: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x). We just deleted p₀ and got the identical [0.5, 0.5], −2. Non-SVs are dead weight to the model.
Only the support vectors (αᵢ > 0) shape the boundary; the rest are irrelevant to it.
Deleting non-support-vectors leaves w, b unchanged
Why: f(x) = Σ over support vectors only. Here just p₁, p₂ matter. (Deleting a real support vector, or adding a point inside the margin, WOULD move it — that's the meaningful edit.)
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
−1 sits low-left, class +1 high-right. We want the boundary line that separates them with the widest gap.; Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.; A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....; All four points went into the fit, so deleting any of them would move the boundary.Section
Part 5 of 6 — the trade-off
Concept
C is the price of a margin violation. Large C makes violations expensive, so the SVM fits the training data hard with a narrow margin. Small C tolerates violations for a wide, smooth margin.
| C | margin | support set | bias / variance |
|---|---|---|---|
| large | narrow | small (few points) | low bias, high variance |
| small | wide | large (many points) | high bias, low variance |
Because a wide margin engulfs more points, small C produces more support vectors. Let's measure that on a genuinely overlapping dataset.
Faded example
Fill in the blanks
The soft margin is hinge loss + regularization, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
w, b = clf.coef_[0], clf.intercept_[0]
margins = **y * (X @ w + b)**
hinge = np.maximum(0, 1 - margins).sum()
print("sum hinge losses =", round(hinge, 4))
print("1/2||w||^2 + Chinge =", round(0.5w@w + C*hinge, 4))
print("points with hinge > 0 =", int((margins < 1).sum()))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The 6 points with functional margin below 1 are exactly the ones paying hinge loss — and they are the α = C support vectors from the KKT slide.
Worked example
At the optimum the slack equals the hinge loss ξᵢ = max(0, 1 − yᵢ(wᵀxᵢ+b)), so the primal objective is ½‖w‖² + C Σ hinge — an L2-regularized hinge classifier. Verify the identity in code (C=1). Standalone:
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
w, b = clf.coef_[0], clf.intercept_[0]
margins = y * (X @ w + b)
hinge = np.maximum(0, 1 - margins).sum()
print("sum hinge losses =", round(hinge, 4))
print("1/2||w||^2 + C*hinge =", round(0.5*w@w + C*hinge, 4))
print("points with hinge > 0 =", int((margins < 1).sum()))hinge sum 5.857, full objective 6.397, 6 points inside the margin
Why: The 6 points with functional margin below 1 are exactly the ones paying hinge loss — and they are the α = C support vectors from the KKT slide. Hinge loss and the slack tell the same story.
| quantity | value (verified) |
|---|---|
| Σ hinge losses | 5.857 |
| ½‖w‖² + C·Σhinge | 6.397 |
| points with hinge > 0 | 6 |
Comparison
Comparison matrix
From The soft margin is hinge loss + regularization: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| quantity | value (verified) |
|---|---|
| Σ hinge losses | 5.857 |
| ½‖w‖² + C·Σhinge | 6.397 |
| points with hinge > 0 | 6 |
Fill the middle
Fill in the blanks
From Sweep C: margin and support count — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 0.1, 1.0, 100.0]:
clf = SVC(kernel='linear', C=C).fit(X, y)
margin = 2.0 / np.linalg.norm(clf.coef_[0])
print(f"C=___: SVs=___ margin=___")
Why: margin is what everything below it consumes, so the wrong expression here fails later and somewhere else. C=0.01 gives a wide 5.635 margin with 30 SVs; C=100 gives a narrow 0.991 margin with only 7.
Worked example
Sixty overlapping blob points (not separable). Sweep C and watch the margin 2/‖w‖ and the support-vector count move together. Standalone:
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 0.1, 1.0, 100.0]:
clf = SVC(kernel='linear', C=C).fit(X, y)
margin = 2.0 / np.linalg.norm(clf.coef_[0])
print(f"C={C:>6}: SVs={clf.n_support_.sum():>2} margin={margin:.3f}")As C rises, margin shrinks and SVs drop
Why: C=0.01 gives a wide 5.635 margin with 30 SVs; C=100 gives a narrow 0.991 margin with only 7. Small C = wide street = many points on/inside it = many SVs.
| C | support vectors | margin 2/‖w‖ |
|---|---|---|
| 0.01 | 30 | 5.635 |
| 0.1 | 14 | 2.410 |
| 1.0 | 9 | 1.925 |
| 100.0 | 7 | 0.991 |
Pattern
Step through it
Step through Sweep C: margin and support count one row at a time. What is driving the change, and what would the row after the last one be?
Pattern
Predict first
The table runs: on the margin | 0 < αᵢ < C | 3 · inside / misclassified | αᵢ = C | 6
In Read the KKT cases off the α values, given the rows so far: what is the next one — the row where KKT group is total support vectors?
Correct: total support vectors | αᵢ > 0 | 9
| KKT group | condition | count (C=1) |
|---|---|---|
| on the margin | 0 < αᵢ < C | 3 |
| inside / misclassified | αᵢ = C | 6 |
| total support vectors | αᵢ > 0 | 9 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Because the blobs overlap, most support vectors are pinned at α = C (they violate the margin).
Worked example
At finite C, the support vectors split into two KKT groups: 0 < αᵢ < C (exactly on the margin) and αᵢ = C (inside the margin or misclassified). Count them at C = 1. Standalone:
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
a = np.abs(clf.dual_coef_[0]) # recover a_i = |a_i y_i|
on_margin = np.sum((a > 1e-6) & (a < C - 1e-6))
at_C = np.sum(a > C - 1e-6)
print("total SVs =", len(a))
print("0 < a < C (on margin) =", on_margin)
print("a = C (inside/misclassified) =", at_C)9 SVs = 3 on-margin + 6 capped at C
Why: Because the blobs overlap, most support vectors are pinned at α = C (they violate the margin). Only 3 sit cleanly on it. The α values encode the geometry exactly as the KKT table predicted.
| KKT group | condition | count (C=1) |
|---|---|---|
| on the margin | 0 < αᵢ < C | 3 |
| inside / misclassified | αᵢ = C | 6 |
| total support vectors | αᵢ > 0 | 9 |
Section
Part 6 of 6 — nonlinear for free
Intuition
Look back at the dual and at f(x): the data appears only as dot products xᵢᵀxⱼ and xᵢᵀx. Nowhere do we need the raw coordinates on their own.
So replace every dot product with a kernel k(xᵢ, xⱼ) = φ(xᵢ)ᵀφ(xⱼ) — the dot product in some richer feature space φ. The optimizer never changes; only the entries of the Gram matrix do.
You get a nonlinear boundary in the original space without ever computing φ. The RBF kernel k(u,v) = exp(−γ‖u−v‖²) corresponds to an infinite-dimensional φ.
Explain it
Discussion prompt
Explain Why the dual kernelizes for free to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Look back at the dual and at f(x): the data appears only as dot products xᵢᵀxⱼ and xᵢᵀx. Nowhere do we need the raw coordinates on their own.
Missing information
Discussion prompt
The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
No straight line separates XOR, so the linear SVM is stuck at 50%. Swapping the linear kernel for RBF — same dual, same code, one keyword change — bends the boundary and separates all four corners.
Worked example
The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[0.,1.],[1.,0.]])
y = np.array([-1., -1., 1., 1.]) # XOR: not linearly separable
lin = SVC(kernel='linear', C=1e6).fit(X, y)
rbf = SVC(kernel='rbf', C=1e6, gamma=1.0).fit(X, y)
print("linear train accuracy =", lin.score(X, y))
print("rbf train accuracy =", rbf.score(X, y))
print("rbf predictions =", rbf.predict(X))linear 0.5 (chance), RBF 1.0 (perfect)
Why: No straight line separates XOR, so the linear SVM is stuck at 50%. Swapping the linear kernel for RBF — same dual, same code, one keyword change — bends the boundary and separates all four corners.
| kernel | train accuracy | separates XOR? |
|---|---|---|
| linear | 0.5 | no |
| rbf (γ=1) | 1.0 | yes |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
linear 0.5 (chance), RBF 1.0 (perfect)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:
Concept
The RBF kernel k(u,v) = exp(−γ‖u−v‖²) has its own hyperparameter γ — the reach of each support vector's influence. Small γ = broad, smooth bumps; large γ = narrow, local spikes.
| γ | influence radius | boundary | risk |
|---|---|---|---|
| small | wide | smooth, almost linear | underfit |
| large | narrow | wiggly, hugs points | overfit |
γ and C are tuned together (usually by grid search with cross-validation). Next slide watches γ bend a ring-shaped boundary.
Analogy
Discussion prompt
Explain The RBF γ is a second knob by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The RBF kernel k(u,v) = exp(−γ‖u−v‖²) has its own hyperparameter γ — the reach of each support vector's influence. Small γ = broad, smooth bumps; large γ = narrow, local spikes.
Estimation
Predict first
Concentric circles — one class inside a ring of the other, impossible for a line. Sweep γ on an RBF SVM and watch training accuracy climb as the boundary tightens. Standalone:
Commit before you compute: what does γ bends a nonlinear boundary come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: γ=0.1 acc 0.990, γ=1 acc 1.000, γ=10 acc 1.000
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Even a modest γ separates the rings.
Worked example
Concentric circles — one class inside a ring of the other, impossible for a line. Sweep γ on an RBF SVM and watch training accuracy climb as the boundary tightens. Standalone:
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_circles
X, y = make_circles(n_samples=100, noise=0.08, factor=0.4, random_state=1)
y = np.where(y == 0, -1, 1)
for g in [0.1, 1.0, 10.0]:
clf = SVC(kernel='rbf', C=1.0, gamma=g).fit(X, y)
print(f"gamma={g:>5}: train acc={clf.score(X,y):.3f} SVs={clf.n_support_.sum()}")γ=0.1 acc 0.990, γ=1 acc 1.000, γ=10 acc 1.000
Why: Even a modest γ separates the rings. Push γ too high and the boundary would hug individual points (overfit) — here training accuracy already saturates at 1.0, so cross-validation, not train accuracy, must pick γ.
| γ | train accuracy | support vectors |
|---|---|---|
| 0.1 | 0.990 | 100 |
| 1.0 | 1.000 | 26 |
| 10.0 | 1.000 | 47 |
Pattern
Step through it
Step through γ bends a nonlinear boundary one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
These are the steps of The SVM toolkit, scrambled. Put them back in order before the next slide shows you.
2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢΣαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors onlyC → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundariesWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢΣαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors onlyC → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundariesEdge cases
Discussion prompt
The SVM toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢΣαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors onlyC → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundariesElimination
Eliminate the wrong options
After training, an SVM's decision function f(x) depends on:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and αᵢ = 0 for every non-support-vector. On our data only p₁, p₂ (α = ¼) contribute; deleting p₀ left w, b unchanged. The model is sparse in the data.
Check
Think back to the margin trace: p₀ and p₃ sat at functional margin 2.
Check your understanding
After training, an SVM's decision function f(x) depends on:
Answer: A
Why: f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and αᵢ = 0 for every non-support-vector. On our data only p₁, p₂ (α = ¼) contribute; deleting p₀ left w, b unchanged. The model is sparse in the data.
Prediction
Predict first
Decreasing C (toward a wider margin) tends to:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: increase bias and decrease variance — a more regularized model
Why: Small C tolerates more margin violations to buy a wider, smoother margin — a higher-bias, lower-variance model. Large C fits the training data hard (low bias, high variance). C is the SVM's regularization strength.
Check
Recall the sweep: C=0.01 gave 30 SVs and a 5.6-wide margin; C=100 gave 7 SVs and a 0.99 margin.
Check your understanding
Decreasing C (toward a wider margin) tends to:
Answer: A
Why: Small C tolerates more margin violations to buy a wider, smoother margin — a higher-bias, lower-variance model. Large C fits the training data hard (low bias, high variance). C is the SVM's regularization strength.
Elimination
Eliminate the wrong options
Why can an SVM swap in an RBF kernel without changing the optimizer?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Both the dual objective Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) and f(x) = Σ αᵢyᵢk(xᵢ,x)+b use the data ONLY via k. Change k and everything else — the constraints, the solver, the sparsity — stays the same. That's the kernel trick, and it flipped XOR from 0.5 to 1.0.
Check
Look at where the data enters the dual objective.
Check your understanding
Why can an SVM swap in an RBF kernel without changing the optimizer?
Answer: A
Why: Both the dual objective Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) and f(x) = Σ αᵢyᵢk(xᵢ,x)+b use the data ONLY via k. Change k and everything else — the constraints, the solver, the sparsity — stays the same. That's the kernel trick, and it flipped XOR from 0.5 to 1.0.
Section
The project
Concept
Fit a linear SVM on our four points, confirm the support set, rebuild the decision function from the duals, then watch C reshape the support set on the noisy blobs. You derived every piece — now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | fit linear SVM, find the support vectors | SVC, support_, n_support_ |
| 2 | rebuild f(x) = Σαᵢyᵢ(xᵢ·x)+b from the duals | dual_coef_, support_vectors_ |
| 3 | vary C, watch the SV count | loop over C on blobs |
Build rules: type every line yourself, run after each, and note dual_coef_ already holds the signed products αᵢyᵢ, not the bare αᵢ.
Counterexample
Discussion prompt
Build rules: type every line yourself, run after each, and note dual_coef_ already holds the signed products αᵢyᵢ, not the bare αᵢ.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: fit a linear SVM on the four points and read the support indices. Predict aloud: which two points, and why?
Hint: SVC(kernel='linear', C=1e6).fit(X, y); read clf.support_ (indices) and clf.n_support_ (per-class counts).
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("support idx =", clf.support_)
print("n_support =", clf.n_support_, "-> total", clf.n_support_.sum())| check | value |
|---|---|
| support idx | [1 2] (p₁, p₂) |
| n_support (per class) | [1 1] |
| total SVs | 2 of 4 |
Worked example
Your turn: reconstruct the decision value at (2.5, 2.5) from the duals and confirm it matches decision_function. Predict the sign before you run.
Hint: dc = clf.dual_coef_[0] (= αᵢyᵢ), sv = clf.support_vectors_; then np.sum(dc * (sv @ xt)) + clf.intercept_[0].
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5])
manual = float(np.sum(dc * (sv @ xt)) + b)
print("manual =", round(manual, 4))
print("sklearn =", round(clf.decision_function([xt])[0], 4))| source | f(2.5, 2.5) |
|---|---|
| manual (from duals) | 0.5 |
| sklearn | 0.5 |
Worked example
Your turn: on the overlapping blobs, loop over C and print the support-vector count. Predict the direction as C grows.
Hint: small C → wide margin → MORE support vectors; large C → tight margin → fewer.
import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 1.0, 100.0]:
print("C =", C, "->", SVC(kernel='linear', C=C).fit(X, y).n_support_.sum(), "SVs")| C | support vectors |
|---|---|
| 0.01 | 30 |
| 1.0 | 9 |
| 100.0 | 7 |
Trade off
Comparison matrix
From Milestone 3 — the C trade-off: every row here is a choice with a cost. Fill the support vectors column, then say which row you would actually pick and what you give up for it.
| C | support vectors |
|---|---|
| 0.01 | 30 |
| 1.0 | 9 |
| 100.0 | 7 |
Concept
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5])
manual = float(np.sum(dc * (sv @ xt)) + b)
print("SVs:", clf.n_support_.sum(), "of", len(X), " idx", clf.support_)
print("w =", clf.coef_[0], " b =", round(b, 4))
print("f(2.5,2.5): manual", round(manual,4),
"== sklearn", np.isclose(manual, clf.decision_function([xt])[0]))| printed line | value (verified) |
|---|---|
| SVs | 2 of 4 idx [1 2] |
| w, b | [0.5 0.5], −2.0 |
| f(2.5,2.5): manual == sklearn | 0.5, True |
If your support set is {p₁, p₂}, w = [0.5, 0.5], b = −2, and your hand-built f matches sklearn — you understand the SVM from the dual up.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| SVs | 2 of 4 idx [1 2] |
| w, b | [0.5 0.5], −2.0 |
| f(2.5,2.5): manual == sklearn | 0.5, True |
Concept
Slides closed, out loud: explain (1) why maximizing the margin is minimizing ½‖w‖², (2) how ∂L/∂w and ∂L/∂b produce w = Σαᵢyᵢxᵢ and Σαᵢyᵢ = 0, and (3) what each αᵢ value — 0, between, or C — says about a point.
Stretch (homework): re-derive the soft-margin dual from the full Lagrangian to see the 0 ≤ αᵢ ≤ C cap appear, then plot the RBF decision boundary for a few γ and C. SVMs sit on Lagrangian duality (Lesson 24) and the kernel trick (Lesson 28).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The margin · The Lagrangian & the dual · Solve the dual by hand · Support vectors & f(x) · Soft margin & C · Kernels. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
2/‖w‖ and the primal min ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢw = Σαᵢyᵢxᵢ and Σαᵢyᵢ = 0, and derive the dual QP with 0 ≤ αᵢ ≤ Cα = [0, ¼, ¼, 0], w = [½, ½], b = −2, margin √2f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and rebuild f from the dualsC as a bias–variance knob and kernelize (RBF) — all matched digit-for-digit against sklearn| move | the one thing to remember |
|---|---|
| margin | widest street = min ½‖w‖²; width = 2/‖w‖ |
| dual | max Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk; data as kernels only |
| support vectors | only αᵢ > 0 points; the rest are inert |
| decision | f(x) = Σ αᵢyᵢk(xᵢ,x) + b, over SVs |
| C | large = tight margin (variance); small = wide (bias) |
| kernel | swap k for nonlinear boundaries — XOR: 0.5 → 1.0 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.