Lesson 30: Support Vector Machines

USAAIO Lesson 30, from Week 10, fully worked. It derives the geometric margin as 2/‖w‖, sets up the primal max-margin QP, and writes out the Lagrangian and the KKT conditions. It then derives the dual QP in α term by term and solves that dual BY HAND on a four-point set, giving α=[0,¼,¼,0], explains why only the support vectors survive, and rebuilds the decision function from the duals. It measures the soft-margin C trade-off and closes with the kernel trick on XOR. One running four-point dataset carries the whole deck; every snippet runs standalone, and every number came from real execution. The lesson runs to 63 slides.

Subject: Machine Learning · 118 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Support Vector Machines

Title

USAAIO · Lesson 30 · Week 10

The maximum-margin classifier, where only a handful of points matter. We derive the margin 2/‖w‖, build the primal QP, pass to the dual through the Lagrangian, then solve that dual by hand on four points and rebuild f(x) from the multipliers alone.

2. By the end of this lesson you can

Objectives

  1. Derive the geometric margin as 2/‖w‖ and state the primal max-margin QP with slack ξᵢ
  2. Build the Lagrangian, apply the KKT stationarity conditions, and derive the dual QP in α step by step
  3. Solve that dual by hand on a 4-point set — get α = [0, ¼, ¼, 0], w = [½, ½], b = −2
  4. Explain why only support vectors (αᵢ > 0) enter f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and rebuild f from the duals
  5. Tune the soft-margin C as a bias–variance knob and kernelize (RBF) to bend the boundary — all matched against sklearn

3. What survived from Mutual Information & Contrastive Learning?

Warm-up

Discussion prompt

Before we open Lesson 30: Support Vector Machines: without looking back, what was the main idea of Mutual Information & Contrastive Learning, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

mutual information I(X;Y) as H(X)+H(Y)−H(X,Y) and as KL(joint‖product), the independence test I=0, feature selection by mutual information, the InfoNCE contrastive loss (CLIP), and the link to compression. Compute MI three equivalent ways and rank features with sklearn.

4. The margin

Section

Part 1 of 6 — the primal

5. The data: four labelled points

Concept

Four points in the plane, two per class. Class −1 sits low-left, class +1 high-right. We want the boundary line that separates them with the widest gap.

pointx₁x₂label y
p₀00−1
p₁11−1
p₂33+1
p₃44+1

Keep this set in your head — every derivation, hand-solve, and code block in the lesson uses these exact four points.

6. Fill in: x₁ for The data: four labelled points

Comparison

Comparison matrix

From The data: four labelled points: refill the x₁ column from what you know. The rest of the table is as it appeared.

pointx₁x₂label y
p₀00−1
p₁11−1
p₂33+1
p₃44+1

7. The widest street

Intuition

Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.

The closest opposing pair here is p₁ = (1,1) and p₂ = (3,3). The street's centerline is their perpendicular bisector; the curbs just touch those two points.

The far points p₀ and p₃ sit well back from the curb. Intuition to test later: they should have no say in where the street goes.

8. Break it if you can: The widest street

Counterexample

Discussion prompt

Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The closest opposing pair here is p₁ = (1,1) and p₂ = (3,3). The street's centerline is their perpendicular bisector; the curbs just touch those two points.

9. A linear classifier and its boundary

Concept

A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.

\[ \hat y = \operatorname{sign}(w^\top x + b), \qquad \text{boundary: } w^\top x + b = 0 \]

w is perpendicular to the boundary (it is the boundary's normal). b shifts the boundary off the origin. Our job is to choose w and b so the gap around the boundary is as wide as possible.

10. By analogy: A linear classifier and its boundary

Analogy

Discussion prompt

Explain A linear classifier and its boundary by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.

11. Why w is perpendicular to the boundary

Intuition

Take any two points xₐ, x_b that both sit on the boundary: wᵀxₐ + b = 0 and wᵀx_b + b = 0. Subtract to get wᵀ(xₐ − x_b) = 0.

xₐ − x_b is a direction along the boundary, and its dot with w is zero — so w is perpendicular to every in-boundary direction. w points straight across the street.

That is why the perpendicular distance formula uses w/‖w‖: moving along w is the fastest way off the boundary, and the margin is measured in exactly that direction.

12. Teach it back: Why w is perpendicular to the boundary

Explain it

Discussion prompt

Explain Why w is perpendicular to the boundary to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Take any two points xₐ, x_b that both sit on the boundary: wᵀxₐ + b = 0 and wᵀx_b + b = 0. Subtract to get wᵀ(xₐ − x_b) = 0.

13. The functional margin and its scaling ambiguity

Concept

Call yᵢ(wᵀxᵢ + b) the functional margin of point i. It is positive when the point is on the correct side, and larger when the point is more confidently classified.

But it has a flaw: scaling (w, b) → (2w, 2b) doubles every functional margin without moving the boundary an inch. Size alone is meaningless until we pin the scale.

Fix the scale: force the closest points to have functional margin exactly 1

Why: This is the canonical normalization. It removes the ambiguity by definition, so 'nearest points' means yᵢ(wᵀxᵢ+b) = 1. Everything downstream depends on this convention.

14. What has to happen first: Derive the geometric margin = 1/‖w‖

Ranking

Put in order

Put the moves of Derive the geometric margin = 1/‖w‖ into the order they have to happen.

  1. Distance from a point to the hyperplane
  2. For a nearest point, the numerator is exactly 1
  3. The full street width (curb to curb) is 2/‖w‖

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Projecting the offset onto the unit normal w/‖w‖ gives the signed perpendicular distance.

15. Derive the geometric margin = 1/‖w‖

Worked example

The geometric margin is the true perpendicular distance from a point to the boundary. Distance from x to the hyperplane wᵀx + b = 0 is the standard point-to-plane formula:

Distance from a point to the hyperplane

Why: Projecting the offset onto the unit normal w/‖w‖ gives the signed perpendicular distance.

\[ \text{dist}(x_i) = \frac{\lvert w^\top x_i + b\rvert}{\lVert w\rVert} \]

For a nearest point, the numerator is exactly 1

Why: Our canonical scaling forces yᵢ(wᵀxᵢ+b) = 1 at the closest points, so |wᵀxᵢ+b| = 1 there.

\[ \gamma = \frac{1}{\lVert w\rVert} \]

The full street width (curb to curb) is 2/‖w‖

Why: One margin γ = 1/‖w‖ on each side of the centerline, so the total separation between the +1 and −1 margin lines is 2/‖w‖. Verified on our fit: ‖w‖ = 0.7071, so 2/‖w‖ = 2.828 and 1/‖w‖ = 1.414.

16. Decode the notation: Derive the geometric margin = 1/‖w‖

Notation

Annotate

From Derive the geometric margin = 1/‖w‖ — read this one piece at a time. What is each part doing?

On: \( \text{dist}(x_i) = \frac{\lvert w^\top x_i + b\rvert}{\lVert w\rVert} \)

  • Projecting the offset onto the unit normal w/‖w‖ gives the signed perpendicular distance.
  • Our canonical scaling forces yᵢ(wᵀxᵢ+b) = 1 at the closest points, so |wᵀxᵢ+b| = 1 there.
  • One margin γ = 1/‖w‖ on each side of the centerline, so the total separation between the +1 and −1 margin lines is 2/‖w‖. Verified on our fit: ‖w‖ = 0.7071, so 2/‖w‖ = 2.828 and 1/‖w‖ = 1.414.

17. Maximize the margin ⇔ minimize ‖w‖

Concept

We want the widest street, so maximize 2/‖w‖. A fraction with fixed numerator is largest when its denominator is smallest — so maximizing the margin is the same as minimizing ‖w‖.

Minimize ½‖w‖² instead of ‖w‖

Why: ‖w‖ has an ugly square root; ½‖w‖² = ½wᵀw is smooth and convex with the same minimizer. The ½ is cosmetic — it cancels the 2 when we differentiate. This is exactly the OLS-style trick from Lesson 7.

\[ \max_{w,b}\ \frac{2}{\lVert w\rVert} \quad\Longleftrightarrow\quad \min_{w,b}\ \tfrac{1}{2}\lVert w\rVert^2 \]

18. The hard-margin primal

Concept

Add the constraint that every point is correctly classified with functional margin at least 1. That gives the hard-margin primal — a convex quadratic program:

\[ \min_{w,b}\ \tfrac{1}{2}\lVert w\rVert^2 \quad \text{s.t. } y_i(w^\top x_i + b) \ge 1 \ \ \forall i \]

Quadratic objective, linear inequality constraints — a textbook QP with a unique optimum when the data is separable. Real data usually isn't perfectly separable, so next we let the constraints bend.

19. Soft margin: let constraints bend with slack

Concept

Introduce a slack ξᵢ ≥ 0 per point: how far point i is allowed to intrude past its margin. Charge C per unit of total slack and add it to the objective.

\[ \min_{w,b,\xi}\ \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \quad \text{s.t. } y_i(w^\top x_i + b) \ge 1 - \xi_i,\ \ \xi_i \ge 0 \]

C = ∞ forbids all slack (back to hard margin). Finite C trades a wider margin against a few violations — the knob we'll tune in Part 5. For the hand-solve we keep it separable, so slack stays 0.

20. Guess the shape of the answer: The dot products, by hand

Estimation

Predict first

The dual needs every pairwise dot product Kᵢⱼ = xᵢᵀxⱼ. All four points lie on the line x₁ = x₂, so xᵢᵀxⱼ = (aᵢ)(aⱼ) + (aᵢ)(aⱼ) = 2 aᵢ aⱼ where a = (0, 1, 3, 4) is the shared coordinate.

Commit before you compute: what does The dot products, by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Row p₂: K₂₀=0, K₂₁=6, K₂₂=18, K₂₃=24

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. 2·a₂·aⱼ = 2·3·(0,1,3,4) = (0, 6, 18, 24).

21. The dot products, by hand

Worked example

The dual needs every pairwise dot product Kᵢⱼ = xᵢᵀxⱼ. All four points lie on the line x₁ = x₂, so xᵢᵀxⱼ = (aᵢ)(aⱼ) + (aᵢ)(aⱼ) = 2 aᵢ aⱼ where a = (0, 1, 3, 4) is the shared coordinate.

Row p₁: K₁₀=0, K₁₁=2, K₁₂=6, K₁₃=8

Why: 2·a₁·aⱼ = 2·1·(0,1,3,4) = (0, 2, 6, 8). Anything times a₀ = 0 vanishes.

Row p₂: K₂₀=0, K₂₁=6, K₂₂=18, K₂₃=24

Why: 2·a₂·aⱼ = 2·3·(0,1,3,4) = (0, 6, 18, 24).

Kp₀p₁p₂p₃
p₀0000
p₁0268
p₂061824
p₃082432

22. What each one costs: The dot products, by hand

Trade off

Comparison matrix

From The dot products, by hand: every row here is a choice with a cost. Fill the p₂ column, then say which row you would actually pick and what you give up for it.

Kp₀p₁p₂p₃
p₀0000
p₁0268
p₂061824
p₃082432

23. What has to be given first: Build the labelled Gram matrix

Missing information

Discussion prompt

Now weight each entry by the labels: the labelled Gram is Gᵢⱼ = yᵢyⱼKᵢⱼ. Confirm the hand table in code — this block re-defines the data and runs on its own:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Any dot product with the origin is 0, so p₀ contributes nothing to K or G. A quiet hint that p₀ will be inert in the dual.

24. Build the labelled Gram matrix

Worked example

Now weight each entry by the labels: the labelled Gram is Gᵢⱼ = yᵢyⱼKᵢⱼ. Confirm the hand table in code — this block re-defines the data and runs on its own:

import numpy as np
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
K = X @ X.T                       # linear kernel = dot products
G = (y[:,None] * y[None,:]) * K   # labelled Gram: G_ij = y_i y_j (x_i . x_j)
print(K)
print(G)

Row/col 0 is all zeros because p₀ = (0,0)

Why: Any dot product with the origin is 0, so p₀ contributes nothing to K or G. A quiet hint that p₀ will be inert in the dual.

p₀p₁p₂p₃
G (row p₀)0000
G (row p₁)02−6−8
G (row p₂)0−61824
G (row p₃)0−82432

25. What stays fixed: Build the labelled Gram matrix

Invariant

Step through it

Step through Build the labelled Gram matrix one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: is G (row p₀)
  2. Step 2: is G (row p₁)
  3. Step 3: is G (row p₂)
  4. Step 4: is G (row p₃)

26. The Lagrangian & the dual

Section

Part 2 of 6 — every step

27. Why leave the primal at all?

Intuition

The primal is solvable, so why change variables? Two payoffs the exam cares about.

First, the dual exposes sparsity: its solution assigns a multiplier αᵢ to each point, and almost all come out exactly zero. The nonzero few are the support vectors.

Second, in the dual the data appears only through dot products xᵢᵀxⱼ. Replace each dot product with a kernel and you get nonlinear boundaries for free — the kernel trick.

28. One Lagrange multiplier per constraint

Concept

Each inequality yᵢ(wᵀxᵢ + b) ≥ 1 becomes 1 − yᵢ(wᵀxᵢ + b) ≤ 0. Attach a multiplier αᵢ ≥ 0 to each and subtract it from the objective — the standard Lagrangian recipe from Lesson 24.

\[ \mathcal{L}(w,b,\alpha) = \tfrac{1}{2}\lVert w\rVert^2 - \sum_i \alpha_i\big[\,y_i(w^\top x_i + b) - 1\,\big], \qquad \alpha_i \ge 0 \]

The dual comes from minimizing L over (w, b) for fixed α, then maximizing over α. Minimizing over w and b means setting their gradients to zero — the next two slides.

29. Plan first: Stationarity in w

Step zero

Discussion prompt

Stationarity in w — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Differentiate L with respect to w and set to 0

Answer:

  1. Differentiate L with respect to w and set to 0
  2. Solve for w
  3. Read off the consequence

30. Stationarity in w

Worked example

Differentiate L with respect to w and set to 0

Why: ∂(½‖w‖²)/∂w = w, and ∂(−Σ αᵢ yᵢ wᵀxᵢ)/∂w = −Σ αᵢ yᵢ xᵢ. The b-term and the +1 have no w.

\[ \nabla_w \mathcal{L} = w - \sum_i \alpha_i y_i x_i = 0 \]

Solve for w

Why: This is the key structural result: the optimal weight vector is a labelled, α-weighted sum of the training points. w lives in the span of the data.

\[ \boxed{\,w = \sum_i \alpha_i y_i x_i\,} \]

Read off the consequence

Why: Any point with αᵢ = 0 drops out of w entirely. Only points with αᵢ > 0 — the support vectors — build the weight vector. Sparsity is baked in right here.

31. Draw the shape of it: Stationarity in w

Blank canvas

Draw it

Draw what Stationarity in w just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

32. Complete the line: Stationarity in b

Fill the middle

Fill in the blanks

From Stationarity in b — finish the line. Write what belongs on the right of the equals sign before you look.

\boxed0\,}}

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Only the term −Σ αᵢ yᵢ b contains b, and ∂/∂b of that is −Σ αᵢ yᵢ.

33. Stationarity in b

Worked example

Differentiate L with respect to b and set to 0

Why: Only the term −Σ αᵢ yᵢ b contains b, and ∂/∂b of that is −Σ αᵢ yᵢ.

\[ \frac{\partial \mathcal{L}}{\partial b} = -\sum_i \alpha_i y_i = 0 \]

This is the dual equality constraint

Why: The α's are not free — they must satisfy Σ αᵢ yᵢ = 0. It ties the positive-class multipliers to the negative-class ones.

\[ \boxed{\,\sum_i \alpha_i y_i = 0\,} \]

34. Substitute back: expand the objective

Worked example

Put w = Σ αᵢyᵢxᵢ back into L and use Σ αᵢyᵢ = 0 to kill the b term. Expand ½‖w‖² first:

½‖w‖² = ½ wᵀw with w = Σ αᵢyᵢxᵢ

Why: A dot of two sums becomes a double sum; each xᵢᵀxⱼ is a dot product.

\[ \tfrac{1}{2}\lVert w\rVert^2 = \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, x_i^\top x_j \]

The cross term −Σ αᵢyᵢ wᵀxᵢ equals −‖w‖² = −Σ αᵢαⱼyᵢyⱼxᵢᵀxⱼ

Why: Because wᵀxᵢ summed against αᵢyᵢ rebuilds wᵀw exactly. And the +Σαᵢ term survives untouched.

\[ \mathcal{L} = \tfrac{1}{2}\!\sum_{i,j}\!\alpha_i\alpha_j y_i y_j x_i^\top x_j - \!\sum_{i,j}\!\alpha_i\alpha_j y_i y_j x_i^\top x_j + \sum_i \alpha_i \]

35. Work backwards from the answer: Substitute back: expand the objective

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

The cross term −Σ αᵢyᵢ wᵀxᵢ equals −‖w‖² = −Σ αᵢαⱼyᵢyⱼxᵢᵀxⱼ

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Put w = Σ αᵢyᵢxᵢ back into L and use Σ αᵢyᵢ = 0 to kill the b term. Expand ½‖w‖² first:

36. Plan first: Collect terms → the dual QP

Step zero

Discussion prompt

Collect terms → the dual QP — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Combine the two double sums: ½ − 1 = −½

Answer:

  1. Combine the two double sums: ½ − 1 = −½
  2. Carry the constraints α ≥ 0 and Σ αᵢyᵢ = 0
  3. Replace xᵢᵀxⱼ with a kernel k(xᵢ,xⱼ)

37. Collect terms → the dual QP

Worked example

Combine the two double sums: ½ − 1 = −½

Why: The identical double sums differ only by their coefficients ½ and −1, which add to −½. That flips the quadratic term negative — the dual is a MAXIMIZATION.

\[ \max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, x_i^\top x_j \]

Carry the constraints α ≥ 0 and Σ αᵢyᵢ = 0

Why: α ≥ 0 came with the multipliers; Σ αᵢyᵢ = 0 came from ∂L/∂b. Together they bound the QP.

\[ \text{s.t. } \alpha_i \ge 0,\qquad \sum_i \alpha_i y_i = 0 \]

Replace xᵢᵀxⱼ with a kernel k(xᵢ,xⱼ)

Why: The data touches the dual ONLY through these dot products. Swapping in any valid kernel gives a nonlinear SVM without ever changing the optimizer — this is the kernel trick.

\[ \boxed{\,\max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j)\,} \]

38. Say it in words: Collect terms → the dual QP

Translation

\( \boxed{\,\max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j)\,} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

39. Soft margin only adds an upper cap C

Concept

Redo the same Lagrangian with the slack terms C Σ ξᵢ and the ξᵢ ≥ 0 multipliers. The algebra is identical, and the slack multipliers collapse into a single change: each αᵢ gains an upper bound C.

\[ \max_{\alpha}\ \sum_i \alpha_i - \tfrac{1}{2}\sum_{i,j}\alpha_i\alpha_j y_i y_j\, k(x_i,x_j) \quad \text{s.t. } 0 \le \alpha_i \le C,\ \sum_i \alpha_i y_i = 0 \]

So the whole soft-margin story lives in one inequality: 0 ≤ αᵢ ≤ C. Hard margin is the C → ∞ limit, where the cap disappears.

40. The KKT conditions read off the solution

Concept

Complementary slackness is the KKT condition that decides which points matter: for each i, αᵢ[yᵢ(wᵀxᵢ+b) − 1 + ξᵢ] = 0. Either the multiplier is zero, or the margin constraint is tight.

value of αᵢwhat it means geometrically
αᵢ = 0not a support vector — strictly outside the margin, correct
0 < αᵢ < Csupport vector exactly ON the margin (ξᵢ = 0)
αᵢ = Csupport vector inside the margin or misclassified (ξᵢ > 0)

Read the trained α and you instantly know each point's status. We measure all three cases in code in Part 5.

41. Something is wrong here: which way does the dual optimize?

Anomaly

Predict first

A student writes this, and it looks reasonable:

The primal is a min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all.

Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.

Why: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.

42. Trap: which way does the dual optimize?

Trap

The trap

The primal is a min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....

min over α of Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk

Why: Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.

The fix

Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.

max over α of Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk

Why: Correct. It is a concave maximization (equivalently, minimize its negative — which is what solvers do). On our data the max is W = 0.25 at α = [0, ¼, ¼, 0]. Duality: this equals the primal ½‖w‖² = 0.25.

43. Break it on purpose: which way does the dual optimize?

Break the constraint

Discussion prompt

The rule this trap just fixed:

Combining the ½ and −1 coefficients flipped the quadratic term negative, so the dual is a maximization.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Wrong direction. Driving this DOWN sends the positive Σαᵢ term toward 0 and the α's to the trivial all-zero point — no classifier at all. The sign flip in the substitution matters.

44. Solve the dual by hand

Section

Part 3 of 6 — the 4 points

45. Use the constraints to shrink the problem

Worked example

Four αᵢ looks hard, but the constraints and geometry collapse it to one variable. First, guess (from the widest-street picture) that only the closest pair p₁, p₂ are support vectors, so α₀ = α₃ = 0. We'll verify this holds.

Apply Σ αᵢyᵢ = 0 to the survivors

Why: With α₀ = α₃ = 0, the equality constraint reads −α₁ + α₂ = 0 (labels y₁ = −1, y₂ = +1).

\[ -\alpha_1 + \alpha_2 = 0 \;\Longrightarrow\; \alpha_1 = \alpha_2 = t \]

One unknown t left

Why: The four-variable QP is now a single-variable concave parabola in t. Maximize it with one derivative.

46. What has to happen first: Plug the Gram entries into the objective

Ranking

Put in order

Put the moves of Plug the Gram entries into the objective into the order they have to happen.

  1. G₁₁ = y₁²(x₁·x₁) = 1·(1+1) = 2
  2. G₂₂ = y₂²(x₂·x₂) = 1·(9+9) = 18
  3. G₁₂ = y₁y₂(x₁·x₂) = (−1)(+1)(3+3) = −6

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p₁ = (1,1), so x₁·x₁ = 1²+1² = 2, and y₁² = 1.

47. Plug the Gram entries into the objective

Worked example

Write the dual with only α₁ = α₂ = t. We need three Gram entries: G₁₁, G₂₂, G₁₂.

G₁₁ = y₁²(x₁·x₁) = 1·(1+1) = 2

Why: p₁ = (1,1), so x₁·x₁ = 1²+1² = 2, and y₁² = 1.

G₂₂ = y₂²(x₂·x₂) = 1·(9+9) = 18

Why: p₂ = (3,3), so x₂·x₂ = 3²+3² = 18.

G₁₂ = y₁y₂(x₁·x₂) = (−1)(+1)(3+3) = −6

Why: x₁·x₂ = 1·3 + 1·3 = 6, times the opposite labels gives −6. These match the Gram matrix printed in Part 1.

48. Plan first: Reduce to a one-variable parabola

Step zero

Discussion prompt

Reduce to a one-variable parabola — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Objective with α₁ = α₂ = t

Answer:

  1. Objective with α₁ = α₂ = t
  2. Collect the bracket: 2 + 18 − 12 = 8
  3. A downward parabola in t

49. Reduce to a one-variable parabola

Worked example

Objective with α₁ = α₂ = t

Why: Σαᵢ = t + t = 2t. The double sum expands to α₁²G₁₁ + α₂²G₂₂ + 2α₁α₂G₁₂ = t²(2) + t²(18) + 2t²(−6).

\[ W(t) = 2t - \tfrac{1}{2}\big[\,2t^2 + 18t^2 + 2t^2(-6)\,\big] \]

Collect the bracket: 2 + 18 − 12 = 8

Why: The three quadratic coefficients sum to 8, so the bracket is 8t².

\[ W(t) = 2t - \tfrac{1}{2}(8t^2) = 2t - 4t^2 \]

A downward parabola in t

Why: Coefficient of t² is −4 < 0, so W(t) is concave with a single interior maximum — exactly what a well-posed dual should be.

50. Where does each piece belong: Lesson 30: Support Vector Machines

Sorting

Sort into buckets

These are the pieces of Lesson 30: Support Vector Machines, out of order. Put each one back under the part of the lesson it belongs to.

The margin
The data: four labelled points; The widest street; A linear classifier and its boundary
The Lagrangian & the dual
Why leave the primal at all?; One Lagrange multiplier per constraint; Stationarity in w
Solve the dual by hand
Use the constraints to shrink the problem; Plug the Gram entries into the objective; Reduce to a one-variable parabola
s1
The margin is where Lesson 30: Support Vector Machines puts The data: four labelled points, The widest street, A linear classifier and its boundary. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The Lagrangian & the dual is where Lesson 30: Support Vector Machines puts Why leave the primal at all?, One Lagrange multiplier per constraint, Stationarity in w. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Solve the dual by hand is where Lesson 30: Support Vector Machines puts Use the constraints to shrink the problem, Plug the Gram entries into the objective, Reduce to a one-variable parabola. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

51. Plan first: Maximize: one derivative

Step zero

Discussion prompt

Maximize: one derivative — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Set dW/dt = 0

Answer:

  1. Set dW/dt = 0
  2. Solve for t
  3. Check α ≥ 0 and the value

52. Maximize: one derivative

Worked example

Set dW/dt = 0

Why: d/dt(2t − 4t²) = 2 − 8t. The maximum of the concave parabola is where the slope vanishes.

\[ \frac{dW}{dt} = 2 - 8t = 0 \]

Solve for t

Why: 8t = 2, so t = ¼. Both support-vector multipliers equal 0.25.

\[ t = \tfrac{1}{4} \;\Longrightarrow\; \alpha = \big[\,0,\ \tfrac14,\ \tfrac14,\ 0\,\big] \]

Check α ≥ 0 and the value

Why: t = ¼ > 0 satisfies the bounds. Dual value W(¼) = 2(¼) − 4(¼)² = 0.5 − 0.25 = 0.25. We'll confirm every digit against a QP solver next.

53. Predict the next row: Verify the hand dual against a QP solver

Pattern

Predict first

The table runs: α₀, α₃ | 0, 0 | 0.0, 0.0 · α₁, α₂ | ¼, ¼ | 0.25, 0.25

In Verify the hand dual against a QP solver, given the rows so far: what is the next one — the row where quantity is dual value W?

Correct: dual value W | 0.25 | 0.25

quantityby handsolver (verified)
α₀, α₃0, 00.0, 0.0
α₁, α₂¼, ¼0.25, 0.25
dual value W0.250.25

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Exactly the hand answer. α₀ = α₃ = 0 confirms p₀ and p₃ are NOT support vectors — our guess was right, not assumed.

54. Verify the hand dual against a QP solver

Worked example

Solve the same dual numerically with scipy.optimize.minimize (minimizing −W). This block re-defines the data and runs on its own:

import numpy as np
from scipy.optimize import minimize
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
G = (y[:,None]*y[None,:]) * (X @ X.T)
neg_W = lambda a: -(a.sum() - 0.5 * a @ G @ a)   # minimize -W
cons  = [{'type':'eq', 'fun': lambda a: a @ y}]  # sum a_i y_i = 0
res   = minimize(neg_W, np.full(4, 0.1), bounds=[(0,None)]*4, constraints=cons)
alpha = np.where(res.x < 1e-7, 0.0, res.x)
print("alpha =", np.round(alpha, 4))
print("dual value W =", round(-res.fun, 4))

Solver returns α = [0, 0.25, 0.25, 0], W = 0.25

Why: Exactly the hand answer. α₀ = α₃ = 0 confirms p₀ and p₃ are NOT support vectors — our guess was right, not assumed.

quantityby handsolver (verified)
α₀, α₃0, 00.0, 0.0
α₁, α₂¼, ¼0.25, 0.25
dual value W0.250.25

55. Watch it run: Verify the hand dual against a QP solver

Pattern

Step through it

Step through Verify the hand dual against a QP solver one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is α₀, α₃
  2. Step 2: quantity is α₁, α₂
  3. Step 3: quantity is dual value W

56. What has to happen first: Recover w from the multipliers

Ranking

Put in order

Put the moves of Recover w from the multipliers into the order they have to happen.

  1. Use w = Σ αᵢyᵢxᵢ
  2. Simplify each component
  3. Sanity: ‖w‖ and the margin

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Only p₁ and p₂ have nonzero α. So w = α₁y₁x₁ + α₂y₂x₂ = ¼(−1)(1,1) + ¼(+1)(3,3).

57. Recover w from the multipliers

Worked example

Use w = Σ αᵢyᵢxᵢ

Why: Only p₁ and p₂ have nonzero α. So w = α₁y₁x₁ + α₂y₂x₂ = ¼(−1)(1,1) + ¼(+1)(3,3).

\[ w = \tfrac14(-1)\!\begin{bmatrix}1\\1\end{bmatrix} + \tfrac14(+1)\!\begin{bmatrix}3\\3\end{bmatrix} = \begin{bmatrix}-\tfrac14 + \tfrac34\\[2pt] -\tfrac14 + \tfrac34\end{bmatrix} \]

Simplify each component

Why: −¼ + ¾ = ½ in both coordinates.

\[ w = \begin{bmatrix} \tfrac12 \\[2pt] \tfrac12 \end{bmatrix} \]

Sanity: ‖w‖ and the margin

Why: ‖w‖ = √(¼+¼) = √½ ≈ 0.7071, so the geometric margin 1/‖w‖ ≈ 1.414 and the full street 2/‖w‖ ≈ 2.828 — matching the Part-1 formula.

58. Draw the shape of it: Recover w from the multipliers

Blank canvas

Draw it

Draw what Recover w from the multipliers just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

59. Recover b from a support vector

Worked example

A support vector sits exactly on its margin: yₛ(wᵀxₛ + b) = 1. Solve for b using p₁:

Plug in p₁ = (1,1), y₁ = −1

Why: wᵀx₁ = ½·1 + ½·1 = 1. So −1·(1 + b) = 1, giving 1 + b = −1.

\[ y_1(w^\top x_1 + b) = 1 \;\Longrightarrow\; -(1 + b) = 1 \;\Longrightarrow\; b = -2 \]

Cross-check with p₂ = (3,3), y₂ = +1

Why: wᵀx₂ = ½·3 + ½·3 = 3. Then +1·(3 + b) = 1 gives b = −2 again. Both support vectors agree, so b = −2 is consistent.

source of bequationb
from p₁−(1 + b) = 1−2
from p₂+(3 + b) = 1−2

60. Decode the notation: Recover b from a support vector

Notation

Annotate

From Recover b from a support vector — read this one piece at a time. What is each part doing?

On: \( y_1(w^\top x_1 + b) = 1 \;\Longrightarrow\; -(1 + b) = 1 \;\Longrightarrow\; b = -2 \)

  • wᵀx₁ = ½·1 + ½·1 = 1. So −1·(1 + b) = 1, giving 1 + b = −1.
  • wᵀx₂ = ½·3 + ½·3 = 3. Then +1·(3 + b) = 1 gives b = −2 again. Both support vectors agree, so b = −2 is consistent.

61. Full margin trace: who sits where

Worked example

With w = [½, ½], b = −2, compute the functional margin yᵢ(wᵀxᵢ + b) for every point. Support vectors land at exactly 1; non-SVs land strictly above.

pointwᵀxᵢwᵀxᵢ+byᵢ(wᵀxᵢ+b)role
p₀ (0,0)0−22non-SV (α=0)
p₁ (1,1)1−11support vector
p₂ (3,3)311support vector
p₃ (4,4)422non-SV (α=0)

SVs at margin 1, non-SVs at margin 2

Why: p₀ and p₃ satisfy the constraint with room to spare (margin 2 > 1), so their multipliers are 0 — they don't touch the curb. This is the geometric meaning of a support vector, computed exactly.

62. Which is which, by yᵢ(wᵀxᵢ+b)

Discrimination

Sort into buckets

Sort these by yᵢ(wᵀxᵢ+b), from memory, without looking back at Full margin trace: who sits where. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

2
p₀ (0,0); p₃ (4,4)
1
p₁ (1,1); p₂ (3,3)
g1
yᵢ(wᵀxᵢ+b) is "2" for p₀ (0,0), p₃ (4,4) — that is what the table on "Full margin trace: who sits where" records, and it is the single property separating this group from the rest.
g2
yᵢ(wᵀxᵢ+b) is "1" for p₁ (1,1), p₂ (3,3) — that is what the table on "Full margin trace: who sits where" records, and it is the single property separating this group from the rest.

63. Restore the missing line: Verify w, b against sklearn

Fill the middle

Fill in the blanks

From Verify w, b against sklearn — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("w =", clf.coef_[0], " b =", round(clf.intercept_[0], 4))
print("support idx =", clf.support_, " n_support =", clf.n_support_)
print("dual_coef (a_i*y_i) =", clf.dual_coef_[0])

Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. Identical to the hand answer. Note dual_coef_ stores the SIGNED products αᵢyᵢ = [−0.25, +0.25], i.e.

64. Verify w, b against sklearn

Worked example

Fit SVC with a huge C (≈ hard margin) and confirm it reproduces our by-hand w, b, and support set. Standalone:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("w =", clf.coef_[0], " b =", round(clf.intercept_[0], 4))
print("support idx =", clf.support_, " n_support =", clf.n_support_)
print("dual_coef (a_i*y_i) =", clf.dual_coef_[0])

sklearn: w = [0.5, 0.5], b = −2, SVs = points 1 and 2

Why: Identical to the hand answer. Note dual_coef_ stores the SIGNED products αᵢyᵢ = [−0.25, +0.25], i.e. α₁ = α₂ = 0.25 with the class signs attached.

quantityby handsklearn (verified)
w[½, ½][0.5, 0.5]
b−2−2.0
support indices1, 2[1 2]
dual_coef_ (αᵢyᵢ)−¼, +¼[−0.25, 0.25]

65. Strong duality: the two values meet

Concept

The primal minimized ½‖w‖²; the dual maximized W(α). For a convex QP with these constraints, strong duality holds: the two optima are equal, with no gap.

\[ \underbrace{\tfrac{1}{2}\lVert w^\star\rVert^2}_{\text{primal min}} \;=\; \underbrace{\sum_i \alpha_i - \tfrac12\sum_{i,j}\alpha_i\alpha_j y_i y_j x_i^\top x_j}_{\text{dual max } W(\alpha^\star)} \]

On our data both equal 0.25. That equality is not a coincidence — it is what licenses solving the dual instead of the primal and reading w, b back off the α.

66. Guess the shape of the answer: Confirm the duality gap is zero

Estimation

Predict first

Compute the primal ½‖w‖² from sklearn's w, and the dual W(α★) from the QP solve, and check they match. Standalone:

Commit before you compute: what does Confirm the duality gap is zero come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: primal = dual = 0.25, gap = 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Strong duality confirmed numerically.

67. Confirm the duality gap is zero

Worked example

Compute the primal ½‖w‖² from sklearn's w, and the dual W(α★) from the QP solve, and check they match. Standalone:

import numpy as np
from scipy.optimize import minimize
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
primal = 0.5 * clf.coef_[0] @ clf.coef_[0]        # 1/2 ||w||^2
G = (y[:,None]*y[None,:]) * (X @ X.T)
res = minimize(lambda a: -(a.sum()-0.5*a@G@a), np.full(4,0.1),
               bounds=[(0,None)]*4, constraints=[{'type':'eq','fun':lambda a:a@y}])
print("primal 1/2||w||^2 =", round(primal,4))
print("dual   W*         =", round(-res.fun,4))
print("equal?", np.isclose(primal, -res.fun))

primal = dual = 0.25, gap = 0

Why: Strong duality confirmed numerically. The solver's dual optimum lands exactly on the primal's ½‖w‖², so nothing is lost by working in α.

sidequantityvalue (verified)
primal½‖w‖²0.25
dualW(α★)0.25
gapprimal − dual0.0

68. Support vectors & f(x)

Section

Part 4 of 6

69. The decision function in dual form

Concept

Substitute w = Σ αᵢyᵢxᵢ into f(x) = wᵀx + b. The dot product xᵢᵀx becomes a kernel, and the sum runs over support vectors only (the rest have αᵢ = 0):

\[ f(x) = \sum_i \alpha_i y_i\, k(x_i, x) + b = \sum_{i \in SV} \alpha_i y_i\, k(x_i, x) + b \]

The prediction is a weighted vote of the support vectors: each SV casts a vote of sign yᵢ, strength αᵢ, scaled by how similar x is to it via the kernel. Sparse and kernelizable at once.

70. Finish it with less help: Rebuild f(x) from the duals alone

Faded example

Fill in the blanks

Rebuild f(x) from the duals alone, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5]) # a test point
manual = **np.sum(dc * (sv @ xt)) + b** # sum a_i y_i (x_i . xt) + b
print("manual f(xt) =", round(manual, 4))
print("sklearn f(xt) =", round(clf.decision_function([xt])[0], 4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Σ αᵢyᵢ(xᵢ·xt) + b, using only the 2 support vectors, reproduces sklearn's decision_function exactly — proof that f is fully determined by the duals.

71. Rebuild f(x) from the duals alone

Worked example

Take sklearn's dual_coef_ (the αᵢyᵢ) and support_vectors_, and compute f at a test point (2.5, 2.5) by hand. It must match decision_function. Standalone:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5])                        # a test point
manual = np.sum(dc * (sv @ xt)) + b              # sum a_i y_i (x_i . xt) + b
print("manual  f(xt) =", round(manual, 4))
print("sklearn f(xt) =", round(clf.decision_function([xt])[0], 4))

manual = sklearn = 0.5

Why: Σ αᵢyᵢ(xᵢ·xt) + b, using only the 2 support vectors, reproduces sklearn's decision_function exactly — proof that f is fully determined by the duals.

test pointmanual f (from duals)sklearn decision_function
(2.5, 2.5)0.50.5
(0, 0) = p₀−2.0−2.0
(2, 2) on boundary0.00.0

72. Restore the missing line: Deleting a non-support-vector changes nothing

Fill the middle

Fill in the blanks

From Deleting a non-support-vector changes nothing — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
full = SVC(kernel='linear', C=1e6).fit(X, y)
keep = [1, 2, 3] # drop point 0 (a NON-SV)
part = SVC(kernel='linear', C=1e6).fit(X[keep], y[keep])
print("full w,b =", full.coef_[0], round(full.intercept_[0],4))
print("drop w,b =", part.coef_[0], round(part.intercept_[0],4))

Why: keep is what everything below it consumes, so the wrong expression here fails later and somewhere else. Identical boundary. The non-support-vector p₀ contributed nothing (α₀ = 0), so its removal is invisible to the model — the SVM's signature sparsity.

73. Deleting a non-support-vector changes nothing

Worked example

The claim: removing a point with αᵢ = 0 leaves the boundary identical. Drop p₀ (a non-SV) and refit. Standalone:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
full = SVC(kernel='linear', C=1e6).fit(X, y)
keep = [1, 2, 3]                                  # drop point 0 (a NON-SV)
part = SVC(kernel='linear', C=1e6).fit(X[keep], y[keep])
print("full w,b =", full.coef_[0], round(full.intercept_[0],4))
print("drop w,b =", part.coef_[0], round(part.intercept_[0],4))

Both fits give w = [0.5, 0.5], b = −2

Why: Identical boundary. The non-support-vector p₀ contributed nothing (α₀ = 0), so its removal is invisible to the model — the SVM's signature sparsity.

datasetwb
all 4 points[0.5, 0.5]−2.0
3 points (p₀ dropped)[0.5, 0.5]−2.0

74. Something is wrong here: every training point shapes the boundary

Anomaly

Predict first

A student writes this, and it looks reasonable:

All four points went into the fit, so deleting any of them would move the boundary.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x).

Only the support vectors (αᵢ > 0) shape the boundary; the rest are irrelevant to it.

Why: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x). We just deleted p₀ and got the identical [0.5, 0.5], −2. Non-SVs are dead weight to the model.

75. Trap: every training point shapes the boundary

Trap

The trap

All four points went into the fit, so deleting any of them would move the boundary.

Assume dropping p₀ or p₃ shifts w, b

Why: False. p₀ and p₃ have αᵢ = 0, so they appear in neither w = Σαᵢyᵢxᵢ nor f(x). We just deleted p₀ and got the identical [0.5, 0.5], −2. Non-SVs are dead weight to the model.

The fix

Only the support vectors (αᵢ > 0) shape the boundary; the rest are irrelevant to it.

Deleting non-support-vectors leaves w, b unchanged

Why: f(x) = Σ over support vectors only. Here just p₁, p₂ matter. (Deleting a real support vector, or adding a point inside the margin, WOULD move it — that's the meaningful edit.)

76. Which of these survive contact with Lesson 30: Support Vector Machines?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Four points in the plane, two per class. Class −1 sits low-left, class +1 high-right. We want the boundary line that separates them with the widest gap.; Many lines separate these classes. The SVM picks the one that carves the widest possible street down the middle — maximum clearance to the nearest point on each side.; A linear classifier scores a point with f(x) = wᵀx + b and predicts the sign. The decision boundary is where the score is zero.
Breaks
The primal is a min of ½‖w‖², so the dual must also be a minimization of Σαᵢ − ½ΣΣ....; All four points went into the fit, so deleting any of them would move the boundary.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 30: Support Vector Machines puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

77. Soft margin & C

Section

Part 5 of 6 — the trade-off

78. C is a bias–variance knob

Concept

C is the price of a margin violation. Large C makes violations expensive, so the SVM fits the training data hard with a narrow margin. Small C tolerates violations for a wide, smooth margin.

Cmarginsupport setbias / variance
largenarrowsmall (few points)low bias, high variance
smallwidelarge (many points)high bias, low variance

Because a wide margin engulfs more points, small C produces more support vectors. Let's measure that on a genuinely overlapping dataset.

79. Finish it with less help: The soft margin is hinge loss +…

Faded example

Fill in the blanks

The soft margin is hinge loss + regularization, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
w, b = clf.coef_[0], clf.intercept_[0]
margins = **y * (X @ w + b)**
hinge = np.maximum(0, 1 - margins).sum()
print("sum hinge losses =", round(hinge, 4))
print("1/2||w||^2 + Chinge =", round(0.5w@w + C*hinge, 4))
print("points with hinge > 0 =", int((margins < 1).sum()))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The 6 points with functional margin below 1 are exactly the ones paying hinge loss — and they are the α = C support vectors from the KKT slide.

80. The soft margin is hinge loss + regularization

Worked example

At the optimum the slack equals the hinge loss ξᵢ = max(0, 1 − yᵢ(wᵀxᵢ+b)), so the primal objective is ½‖w‖² + C Σ hinge — an L2-regularized hinge classifier. Verify the identity in code (C=1). Standalone:

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
w, b = clf.coef_[0], clf.intercept_[0]
margins = y * (X @ w + b)
hinge = np.maximum(0, 1 - margins).sum()
print("sum hinge losses     =", round(hinge, 4))
print("1/2||w||^2 + C*hinge  =", round(0.5*w@w + C*hinge, 4))
print("points with hinge > 0 =", int((margins < 1).sum()))

hinge sum 5.857, full objective 6.397, 6 points inside the margin

Why: The 6 points with functional margin below 1 are exactly the ones paying hinge loss — and they are the α = C support vectors from the KKT slide. Hinge loss and the slack tell the same story.

quantityvalue (verified)
Σ hinge losses5.857
½‖w‖² + C·Σhinge6.397
points with hinge > 06

81. Fill in: value (verified) for The soft margin is hinge loss +…

Comparison

Comparison matrix

From The soft margin is hinge loss + regularization: refill the value (verified) column from what you know. The rest of the table is as it appeared.

quantityvalue (verified)
Σ hinge losses5.857
½‖w‖² + C·Σhinge6.397
points with hinge > 06

82. Restore the missing line: Sweep C: margin and support count

Fill the middle

Fill in the blanks

From Sweep C: margin and support count — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 0.1, 1.0, 100.0]:
clf = SVC(kernel='linear', C=C).fit(X, y)
margin = 2.0 / np.linalg.norm(clf.coef_[0])
print(f"C=___: SVs=___ margin=___")

Why: margin is what everything below it consumes, so the wrong expression here fails later and somewhere else. C=0.01 gives a wide 5.635 margin with 30 SVs; C=100 gives a narrow 0.991 margin with only 7.

83. Sweep C: margin and support count

Worked example

Sixty overlapping blob points (not separable). Sweep C and watch the margin 2/‖w‖ and the support-vector count move together. Standalone:

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 0.1, 1.0, 100.0]:
    clf = SVC(kernel='linear', C=C).fit(X, y)
    margin = 2.0 / np.linalg.norm(clf.coef_[0])
    print(f"C={C:>6}: SVs={clf.n_support_.sum():>2}  margin={margin:.3f}")

As C rises, margin shrinks and SVs drop

Why: C=0.01 gives a wide 5.635 margin with 30 SVs; C=100 gives a narrow 0.991 margin with only 7. Small C = wide street = many points on/inside it = many SVs.

Csupport vectorsmargin 2/‖w‖
0.01305.635
0.1142.410
1.091.925
100.070.991

84. Watch it run: Sweep C: margin and support count

Pattern

Step through it

Step through Sweep C: margin and support count one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: C is 0.01
  2. Step 2: C is 0.1
  3. Step 3: C is 1.0
  4. Step 4: C is 100.0

85. Predict the next row: Read the KKT cases off the α values

Pattern

Predict first

The table runs: on the margin | 0 < αᵢ < C | 3 · inside / misclassified | αᵢ = C | 6

In Read the KKT cases off the α values, given the rows so far: what is the next one — the row where KKT group is total support vectors?

Correct: total support vectors | αᵢ > 0 | 9

KKT groupconditioncount (C=1)
on the margin0 < αᵢ < C3
inside / misclassifiedαᵢ = C6
total support vectorsαᵢ > 09

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Because the blobs overlap, most support vectors are pinned at α = C (they violate the margin).

86. Read the KKT cases off the α values

Worked example

At finite C, the support vectors split into two KKT groups: 0 < αᵢ < C (exactly on the margin) and αᵢ = C (inside the margin or misclassified). Count them at C = 1. Standalone:

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
C = 1.0
clf = SVC(kernel='linear', C=C).fit(X, y)
a = np.abs(clf.dual_coef_[0])            # recover a_i = |a_i y_i|
on_margin = np.sum((a > 1e-6) & (a < C - 1e-6))
at_C      = np.sum(a > C - 1e-6)
print("total SVs =", len(a))
print("0 < a < C (on margin) =", on_margin)
print("a = C (inside/misclassified) =", at_C)

9 SVs = 3 on-margin + 6 capped at C

Why: Because the blobs overlap, most support vectors are pinned at α = C (they violate the margin). Only 3 sit cleanly on it. The α values encode the geometry exactly as the KKT table predicted.

KKT groupconditioncount (C=1)
on the margin0 < αᵢ < C3
inside / misclassifiedαᵢ = C6
total support vectorsαᵢ > 09

87. Kernels

Section

Part 6 of 6 — nonlinear for free

88. Why the dual kernelizes for free

Intuition

Look back at the dual and at f(x): the data appears only as dot products xᵢᵀxⱼ and xᵢᵀx. Nowhere do we need the raw coordinates on their own.

So replace every dot product with a kernel k(xᵢ, xⱼ) = φ(xᵢ)ᵀφ(xⱼ) — the dot product in some richer feature space φ. The optimizer never changes; only the entries of the Gram matrix do.

You get a nonlinear boundary in the original space without ever computing φ. The RBF kernel k(u,v) = exp(−γ‖u−v‖²) corresponds to an infinite-dimensional φ.

89. Teach it back: Why the dual kernelizes for free

Explain it

Discussion prompt

Explain Why the dual kernelizes for free to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Look back at the dual and at f(x): the data appears only as dot products xᵢᵀxⱼ and xᵢᵀx. Nowhere do we need the raw coordinates on their own.

90. What has to be given first: XOR: where linear fails and RBF wins

Missing information

Discussion prompt

The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

No straight line separates XOR, so the linear SVM is stuck at 50%. Swapping the linear kernel for RBF — same dual, same code, one keyword change — bends the boundary and separates all four corners.

91. XOR: where linear fails and RBF wins

Worked example

The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[0.,1.],[1.,0.]])
y = np.array([-1., -1., 1., 1.])                  # XOR: not linearly separable
lin = SVC(kernel='linear', C=1e6).fit(X, y)
rbf = SVC(kernel='rbf', C=1e6, gamma=1.0).fit(X, y)
print("linear train accuracy =", lin.score(X, y))
print("rbf    train accuracy =", rbf.score(X, y))
print("rbf predictions =", rbf.predict(X))

linear 0.5 (chance), RBF 1.0 (perfect)

Why: No straight line separates XOR, so the linear SVM is stuck at 50%. Swapping the linear kernel for RBF — same dual, same code, one keyword change — bends the boundary and separates all four corners.

kerneltrain accuracyseparates XOR?
linear0.5no
rbf (γ=1)1.0yes

92. Work backwards from the answer: XOR: where linear fails and RBF wins

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

linear 0.5 (chance), RBF 1.0 (perfect)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The XOR set is the canonical non-linearly-separable problem: opposite corners share a class. A linear SVM cannot beat 50%; an RBF SVM nails it. Standalone:

93. The RBF γ is a second knob

Concept

The RBF kernel k(u,v) = exp(−γ‖u−v‖²) has its own hyperparameter γ — the reach of each support vector's influence. Small γ = broad, smooth bumps; large γ = narrow, local spikes.

γinfluence radiusboundaryrisk
smallwidesmooth, almost linearunderfit
largenarrowwiggly, hugs pointsoverfit

γ and C are tuned together (usually by grid search with cross-validation). Next slide watches γ bend a ring-shaped boundary.

94. By analogy: The RBF γ is a second knob

Analogy

Discussion prompt

Explain The RBF γ is a second knob by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The RBF kernel k(u,v) = exp(−γ‖u−v‖²) has its own hyperparameter γ — the reach of each support vector's influence. Small γ = broad, smooth bumps; large γ = narrow, local spikes.

95. Guess the shape of the answer: γ bends a nonlinear boundary

Estimation

Predict first

Concentric circles — one class inside a ring of the other, impossible for a line. Sweep γ on an RBF SVM and watch training accuracy climb as the boundary tightens. Standalone:

Commit before you compute: what does γ bends a nonlinear boundary come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: γ=0.1 acc 0.990, γ=1 acc 1.000, γ=10 acc 1.000

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Even a modest γ separates the rings.

96. γ bends a nonlinear boundary

Worked example

Concentric circles — one class inside a ring of the other, impossible for a line. Sweep γ on an RBF SVM and watch training accuracy climb as the boundary tightens. Standalone:

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_circles
X, y = make_circles(n_samples=100, noise=0.08, factor=0.4, random_state=1)
y = np.where(y == 0, -1, 1)
for g in [0.1, 1.0, 10.0]:
    clf = SVC(kernel='rbf', C=1.0, gamma=g).fit(X, y)
    print(f"gamma={g:>5}: train acc={clf.score(X,y):.3f}  SVs={clf.n_support_.sum()}")

γ=0.1 acc 0.990, γ=1 acc 1.000, γ=10 acc 1.000

Why: Even a modest γ separates the rings. Push γ too high and the boundary would hug individual points (overfit) — here training accuracy already saturates at 1.0, so cross-validation, not train accuracy, must pick γ.

γtrain accuracysupport vectors
0.10.990100
1.01.00026
10.01.00047

97. Watch it run: γ bends a nonlinear boundary

Pattern

Step through it

Step through γ bends a nonlinear boundary one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: γ is 0.1
  2. Step 2: γ is 1.0
  3. Step 3: γ is 10.0

98. Rebuild the recipe: The SVM toolkit

Ranking

Put in order

These are the steps of The SVM toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Primal: maximize the margin 2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢ
  2. Dual: maximize Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0
  3. Recover: w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)
  4. Predict: f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors only
  5. Tune: large C → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundaries

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

99. The SVM toolkit

Pattern

  1. Primal: maximize the margin 2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢ
  2. Dual: maximize Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0
  3. Recover: w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)
  4. Predict: f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors only
  5. Tune: large C → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundaries

100. Where does it stop working: The SVM toolkit

Edge cases

Discussion prompt

The SVM toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Primal: maximize the margin 2/‖w‖ = minimize ½‖w‖² + C Σξᵢ subject to yᵢ(wᵀxᵢ+b) ≥ 1 − ξᵢ
  2. Dual: maximize Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) subject to 0 ≤ αᵢ ≤ C and Σαᵢyᵢ = 0
  3. Recover: w = Σαᵢyᵢxᵢ, and b from any on-margin SV (0 < αᵢ < C)
  4. Predict: f(x) = Σ_{i∈SV} αᵢyᵢk(xᵢ,x) + b — support vectors only
  5. Tune: large C → narrow margin (variance); small C → wide margin (bias). Swap the kernel for nonlinear boundaries

101. Rule out three: Check yourself — support vectors

Elimination

Eliminate the wrong options

After training, an SVM's decision function f(x) depends on:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. only the support vectors — the points with αᵢ > 0
  • B. all training points, weighted equally
  • C. only the two class means (centroids)
  • D. the points farthest from the boundary

Survives elimination: A

Why: f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and αᵢ = 0 for every non-support-vector. On our data only p₁, p₂ (α = ¼) contribute; deleting p₀ left w, b unchanged. The model is sparse in the data.

102. Check yourself — support vectors

Check

Think back to the margin trace: p₀ and p₃ sat at functional margin 2.

Check your understanding

After training, an SVM's decision function f(x) depends on:

  • A. only the support vectors — the points with αᵢ > 0 (correct)
  • B. all training points, weighted equally
  • C. only the two class means (centroids)
  • D. the points farthest from the boundary

Answer: A

Why: f(x) = Σ αᵢyᵢk(xᵢ,x) + b, and αᵢ = 0 for every non-support-vector. On our data only p₁, p₂ (α = ¼) contribute; deleting p₀ left w, b unchanged. The model is sparse in the data.

Why B tempts people
Most points have αᵢ = 0 and contribute nothing. Here 2 of 4 points carry the entire boundary; the other two are dead weight.
Why C tempts people
SVMs are not centroid/nearest-mean classifiers. The boundary is fixed by the margin-defining points, and moving a far point (or the mean) need not move it.
Why D tempts people
Exactly backwards — the support vectors are the CLOSEST points (on the margin, functional margin 1). The far points p₀, p₃ have αᵢ = 0.

103. Answer it before you see the options: Check yourself — the C knob

Prediction

Predict first

Decreasing C (toward a wider margin) tends to:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: increase bias and decrease variance — a more regularized model

Why: Small C tolerates more margin violations to buy a wider, smoother margin — a higher-bias, lower-variance model. Large C fits the training data hard (low bias, high variance). C is the SVM's regularization strength.

104. Check yourself — the C knob

Check

Recall the sweep: C=0.01 gave 30 SVs and a 5.6-wide margin; C=100 gave 7 SVs and a 0.99 margin.

Check your understanding

Decreasing C (toward a wider margin) tends to:

  • A. increase bias and decrease variance — a more regularized model (correct)
  • B. decrease bias and increase variance
  • C. have no effect on the margin or the support set
  • D. always raise test accuracy

Answer: A

Why: Small C tolerates more margin violations to buy a wider, smoother margin — a higher-bias, lower-variance model. Large C fits the training data hard (low bias, high variance). C is the SVM's regularization strength.

Why B tempts people
That's LARGE C — a tight margin that fits training points hard, lowering bias and raising variance. Small C does the opposite.
Why C tempts people
C strongly reshapes both: in the sweep the margin ran 5.635 → 0.991 and the support count 30 → 7 as C rose from 0.01 to 100.
Why D tempts people
Neither extreme is universally best. Too-small C underfits, too-large C overfits; C is tuned on validation data, not maximized.

105. Rule out three: Check yourself — the dual & kernels

Elimination

Eliminate the wrong options

Why can an SVM swap in an RBF kernel without changing the optimizer?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The dual and f(x) touch the data only through inner products k(xᵢ,xⱼ), so replacing k is the only change
  • B. It first computes the explicit feature map φ(x), then dots the vectors
  • C. The labels yᵢ are dropped once a kernel is used
  • D. Every point becomes a support vector under a kernel

Survives elimination: A

Why: Both the dual objective Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) and f(x) = Σ αᵢyᵢk(xᵢ,x)+b use the data ONLY via k. Change k and everything else — the constraints, the solver, the sparsity — stays the same. That's the kernel trick, and it flipped XOR from 0.5 to 1.0.

106. Check yourself — the dual & kernels

Check

Look at where the data enters the dual objective.

Check your understanding

Why can an SVM swap in an RBF kernel without changing the optimizer?

  • A. The dual and f(x) touch the data only through inner products k(xᵢ,xⱼ), so replacing k is the only change (correct)
  • B. It first computes the explicit feature map φ(x), then dots the vectors
  • C. The labels yᵢ are dropped once a kernel is used
  • D. Every point becomes a support vector under a kernel

Answer: A

Why: Both the dual objective Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk(xᵢ,xⱼ) and f(x) = Σ αᵢyᵢk(xᵢ,x)+b use the data ONLY via k. Change k and everything else — the constraints, the solver, the sparsity — stays the same. That's the kernel trick, and it flipped XOR from 0.5 to 1.0.

Why B tempts people
The whole point is to AVOID forming φ — for the RBF kernel φ is infinite-dimensional. k computes the needed inner product directly.
Why C tempts people
The labels stay in the sum αᵢyᵢk(...); they set the sign of each support vector's vote and appear in the constraint Σαᵢyᵢ = 0.
Why D tempts people
Kernels don't force every point to be an SV. Sparsity persists — most αᵢ are still 0; only margin-defining points survive.

107. Your turn: fit an SVM

Section

The project

108. Project: support vectors & the dual

Concept

Fit a linear SVM on our four points, confirm the support set, rebuild the decision function from the duals, then watch C reshape the support set on the noisy blobs. You derived every piece — now assemble it.

#requirementtool
1fit linear SVM, find the support vectorsSVC, support_, n_support_
2rebuild f(x) = Σαᵢyᵢ(xᵢ·x)+b from the dualsdual_coef_, support_vectors_
3vary C, watch the SV countloop over C on blobs

Build rules: type every line yourself, run after each, and note dual_coef_ already holds the signed products αᵢyᵢ, not the bare αᵢ.

109. Break it if you can: Project: support vectors & the dual

Counterexample

Discussion prompt

Build rules: type every line yourself, run after each, and note dual_coef_ already holds the signed products αᵢyᵢ, not the bare αᵢ.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

110. Milestone 1 — fit & find the SVs

Worked example

Your turn: fit a linear SVM on the four points and read the support indices. Predict aloud: which two points, and why?

Hint: SVC(kernel='linear', C=1e6).fit(X, y); read clf.support_ (indices) and clf.n_support_ (per-class counts).

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print("support idx =", clf.support_)
print("n_support   =", clf.n_support_, "-> total", clf.n_support_.sum())
checkvalue
support idx[1 2] (p₁, p₂)
n_support (per class)[1 1]
total SVs2 of 4

111. Milestone 2 — rebuild f(x)

Worked example

Your turn: reconstruct the decision value at (2.5, 2.5) from the duals and confirm it matches decision_function. Predict the sign before you run.

Hint: dc = clf.dual_coef_[0] (= αᵢyᵢ), sv = clf.support_vectors_; then np.sum(dc * (sv @ xt)) + clf.intercept_[0].

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5])
manual = float(np.sum(dc * (sv @ xt)) + b)
print("manual  =", round(manual, 4))
print("sklearn =", round(clf.decision_function([xt])[0], 4))
sourcef(2.5, 2.5)
manual (from duals)0.5
sklearn0.5

112. Milestone 3 — the C trade-off

Worked example

Your turn: on the overlapping blobs, loop over C and print the support-vector count. Predict the direction as C grows.

Hint: small C → wide margin → MORE support vectors; large C → tight margin → fewer.

import numpy as np
from sklearn.svm import SVC
from sklearn.datasets import make_blobs
X, y = make_blobs(n_samples=60, centers=2, cluster_std=2.2, random_state=3)
y = np.where(y == 0, -1, 1)
for C in [0.01, 1.0, 100.0]:
    print("C =", C, "->", SVC(kernel='linear', C=C).fit(X, y).n_support_.sum(), "SVs")
Csupport vectors
0.0130
1.09
100.07

113. What each one costs: Milestone 3 — the C trade-off

Trade off

Comparison matrix

From Milestone 3 — the C trade-off: every row here is a choice with a cost. Fill the support vectors column, then say which row you would actually pick and what you give up for it.

Csupport vectors
0.0130
1.09
100.07

114. The full program

Concept

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[1.,1.],[3.,3.],[4.,4.]])
y = np.array([-1.,-1.,1.,1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
sv, dc, b = clf.support_vectors_, clf.dual_coef_[0], clf.intercept_[0]
xt = np.array([2.5, 2.5])
manual = float(np.sum(dc * (sv @ xt)) + b)
print("SVs:", clf.n_support_.sum(), "of", len(X), "  idx", clf.support_)
print("w =", clf.coef_[0], " b =", round(b, 4))
print("f(2.5,2.5): manual", round(manual,4),
      "== sklearn", np.isclose(manual, clf.decision_function([xt])[0]))
printed linevalue (verified)
SVs2 of 4 idx [1 2]
w, b[0.5 0.5], −2.0
f(2.5,2.5): manual == sklearn0.5, True

If your support set is {p₁, p₂}, w = [0.5, 0.5], b = −2, and your hand-built f matches sklearn — you understand the SVM from the dual up.

115. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
SVs2 of 4 idx [1 2]
w, b[0.5 0.5], −2.0
f(2.5,2.5): manual == sklearn0.5, True

116. Show it off

Concept

Slides closed, out loud: explain (1) why maximizing the margin is minimizing ½‖w‖², (2) how ∂L/∂w and ∂L/∂b produce w = Σαᵢyᵢxᵢ and Σαᵢyᵢ = 0, and (3) what each αᵢ value — 0, between, or C — says about a point.

Stretch (homework): re-derive the soft-margin dual from the full Lagrangian to see the 0 ≤ αᵢ ≤ C cap appear, then plot the RBF decision boundary for a few γ and C. SVMs sit on Lagrangian duality (Lesson 24) and the kernel trick (Lesson 28).

117. Connect it up: Lesson 30: Support Vector Machines

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The margin · The Lagrangian & the dual · Solve the dual by hand · Support vectors & f(x) · Soft margin & C · Kernels. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

118. What you can do now

Recap

movethe one thing to remember
marginwidest street = min ½‖w‖²; width = 2/‖w‖
dualmax Σαᵢ − ½ΣΣ αᵢαⱼyᵢyⱼk; data as kernels only
support vectorsonly αᵢ > 0 points; the rest are inert
decisionf(x) = Σ αᵢyᵢk(xᵢ,x) + b, over SVs
Clarge = tight margin (variance); small = wide (bias)
kernelswap k for nonlinear boundaries — XOR: 0.5 → 1.0

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 30 (Week 10 — Support Vector Machines) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn SVC (dual_coef_, support_vectors_, decision_function)
  3. Bishop, Pattern Recognition and Machine Learning, Ch. 7 (Sparse Kernel Machines) — Springer, 2006
  4. Every α, w, b, support-vector count, margin, and decision value produced by real execution — numpy 2.2 + scipy 1.16 + scikit-learn 1.9, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108