USAAIO Lesson 43, from Week 15 of Phase 2, fully worked. It derives the maximum-margin hyperplane from the signed-distance formula with no steps skipped, explains the ±1 canonical scaling, and proves that the margin is 2/‖w‖ on a hand-solvable four-point toy. It then sets up the hard-margin QP with its Lagrangian and KKT conditions and verifies w = Σαᵢyᵢxᵢ against sklearn's dual coefficients. From there it covers soft-margin slack and the exact role of C, the kernel trick with an RBF value checked by hand, grid search with cross-validation, one-versus-one against one-versus-rest for multi-class, and SVM against logistic regression, ending with a from-scratch Iris pipeline. Every snippet runs standalone, and every number came from real execution. The lesson runs to 64 slides.
Subject: Machine Learning · 104 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 43 · Week 15 (Phase 2)
We derive the maximum-margin hyperplane from the signed-distance formula — every step — prove margin = 2/‖w‖ on a 4-point toy you can solve by hand, turn it into the hard-margin QP, read off support vectors from the dual, then soften it with slack, kernelize it, and tune it. Every number comes from real execution.
Objectives
margin = 2/‖w‖ with no skipped stepsmin ½‖w‖² and read off the support vectorsw = Σᵢ αᵢyᵢxᵢ against sklearn's dual coefficientsξᵢ and explain exactly how C trades margin width against violationsC/gamma with grid search + CV, and choose OvO/OvR and SVM vs logistic regressionWarm-up
Discussion prompt
Before we open Lesson 43: Support Vector Machines: without looking back, what was the main idea of PyTorch Data Pipeline & Model Serialization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
custom Dataset (__len__, __getitem__), DataLoader (batching, shuffling, num_workers), manual feature transforms, nn.Module parameter inspection (.parameters(), .named_parameters()), and model serialization (state_dict vs full-model save, checkpoint save/resume). Build a complete pipeline: 120-sample tabular dataset, TwoLayerNet(4→8→1), train with Adam, save checkpoint at epoch 5, resume and reach loss 0.1666 at epoch 10.
Section
Part 1 of 6 — the geometry
Concept
A linear classifier draws a flat boundary — a hyperplane — and predicts by which side a point falls on. In d dimensions the boundary is the set of points where a linear score is zero:
\[ f(x) = w^\top x + b = 0 \]
w is the weight (normal) vector — it points perpendicular to the boundary. b is the bias that shifts the boundary off the origin. We predict class +1 when f(x) > 0 and class −1 when f(x) < 0.
Counterexample
Discussion prompt
A linear classifier draws a flat boundary — a hyperplane — and predicts by which side a point falls on. In d dimensions the boundary is the set of points where a linear score is zero:
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Picture it
Figure (svg): Two clusters of points, one class lower-left and one upper-right, with three different straight lines all separating them; one line runs down the middle of the gap, the other two hug one cluster.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
If two classes are linearly separable, infinitely many hyperplanes split them perfectly. A perceptron stops at the first one it finds. The SVM asks a sharper question: which separator is safest?
Concept
If two classes are linearly separable, infinitely many hyperplanes split them perfectly. A perceptron stops at the first one it finds. The SVM asks a sharper question: which separator is safest?
Figure (svg): Two clusters of points, one class lower-left and one upper-right, with three different straight lines all separating them; one line runs down the middle of the gap, the other two hug one cluster.
Intuition
Picture the boundary as a street you widen until it touches the nearest point on each side. A separator squeezed right up against one cloud will misjudge a slightly-shifted test point; a separator down the middle of the widest street has the most slack before a new point crosses.
The half-width of that street is the margin. The SVM is the unique separator that makes the margin as large as possible — so 'safest' becomes a precise, solvable objective.
Analogy
Discussion prompt
Explain Widest street wins by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The half-width of that street is the margin. The SVM is the unique separator that makes the margin as large as possible — so 'safest' becomes a precise, solvable objective.
Concept
We carry one dataset through the whole geometry: four points, two per class, chosen so every number is clean and checkable by hand and against sklearn.
| point | x₁ | x₂ | class y |
|---|---|---|---|
| a | 0 | 0 | −1 |
| b | 2 | 0 | −1 |
| c | 0 | 4 | +1 |
| d | 2 | 4 | +1 |
The two classes are stacked: the negatives sit at height x₂ = 0, the positives at x₂ = 4. By symmetry the best boundary is the horizontal line x₂ = 2 — halfway between. The whole point is to derive that, then read w and b off it.
Comparison
Comparison matrix
From Our hand-solvable toy: refill the x₁ column from what you know. The rest of the table is as it appeared.
| point | x₁ | x₂ | class y |
|---|---|---|---|
| a | 0 | 0 | −1 |
| b | 2 | 0 | −1 |
| c | 0 | 4 | +1 |
| d | 2 | 4 | +1 |
Concept
The perpendicular distance from a point x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:
\[ \operatorname{dist}(x_0) = \frac{w^\top x_0 + b}{\lVert w \rVert} \]
The sign tells you the side; the magnitude tells you how far. Dividing by ‖w‖ removes the arbitrary scale of w — double w and b and the boundary is unchanged, so the true distance can't depend on that scale.
Explain it
Discussion prompt
Explain Signed distance from a point to the plane to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The perpendicular distance from a point x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:
Estimation
Predict first
Take the boundary x₂ = 2, written as wᵀx + b = 0 with w = [0, ½], b = −1 (so ½·x₂ − 1 = 0 ⇒ x₂ = 2). Then ‖w‖ = ½. Verify the distances match reality — this snippet runs on its own:
Commit before you compute: what does Signed distance, computed on the toy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Negatives land at distance −2, positives at +2
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For a=(0,0): (0·0 + ½·0 − 1)/½ = −1/0.5 = −2.
Worked example
Take the boundary x₂ = 2, written as wᵀx + b = 0 with w = [0, ½], b = −1 (so ½·x₂ − 1 = 0 ⇒ x₂ = 2). Then ‖w‖ = ½. Verify the distances match reality — this snippet runs on its own:
import numpy as np
w = np.array([0.0, 0.5]); b = -1.0
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]]) # a, b, c, d
for x in X:
signed = (w @ x + b) / np.linalg.norm(w)
print(x, 'signed dist =', round(signed, 3))Negatives land at distance −2, positives at +2
Why: For a=(0,0): (0·0 + ½·0 − 1)/½ = −1/0.5 = −2. For c=(0,4): (½·4 − 1)/½ = 1/0.5 = +2. Both classes are exactly 2 units off the boundary — the line really is centered.
| point | wᵀx + b | ÷ ‖w‖ = signed dist |
|---|---|---|
| a (0,0) | −1.0 | −2.0 |
| b (2,0) | −1.0 | −2.0 |
| c (0,4) | +1.0 | +2.0 |
| d (2,4) | +1.0 | +2.0 |
Discrimination
Sort into buckets
Sort these by wᵀx + b, from memory, without looking back at Signed distance, computed on the toy. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Because scaling (w, b) doesn't move the boundary, we are free to fix the scale. The SVM convention pins it so the closest points score exactly ±1:
\[ \min_i \; \lvert w^\top x_i + b \rvert = 1 \]
With that choice every point obeys yᵢ(wᵀxᵢ + b) ≥ 1, and the two closest points sit on the margin planes wᵀx + b = +1 and wᵀx + b = −1. Those points are the support vectors.
Ranking
Put in order
Put the moves of Derive w and b from the ±1 condition into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Point a=(0,0) sits on the wᵀx+b = −1 margin plane.
Worked example
Force the closest negative a=(0,0) to score −1 and the closest positive c=(0,4) to score +1. By the toy's symmetry w = [0, w₂] (the boundary is horizontal), so only w₂ and b are unknown:
Negative support: w₂·0 + b = −1 ⇒ b = −1
Why: Point a=(0,0) sits on the wᵀx+b = −1 margin plane. Its x₂ is 0, so the whole score is just b.
Positive support: w₂·4 + b = +1 ⇒ 4w₂ − 1 = 1
Why: Point c=(0,4) sits on the +1 plane. Substitute b = −1 from the previous step.
Solve: 4w₂ = 2 ⇒ w₂ = 0.5
Why: So w = [0, 0.5], b = −1. The boundary wᵀx+b = 0 is ½x₂ − 1 = 0, i.e. x₂ = 2 — exactly the midline we predicted.
\[ w = \begin{bmatrix} 0 \\ 0.5 \end{bmatrix}, \qquad b = -1 \]
Notation
Annotate
From Derive w and b from the ±1 condition — read this one piece at a time. What is each part doing?
On: \( w = \begin{bmatrix} 0 \\ 0.5 \end{bmatrix}, \qquad b = -1 \)
Concept
A positive support x₊ scores +1 and a negative support x₋ scores −1. Their signed distances to the boundary are therefore:
\[ \frac{w^\top x_+ + b}{\lVert w\rVert} = \frac{+1}{\lVert w\rVert}, \qquad \frac{w^\top x_- + b}{\lVert w\rVert} = \frac{-1}{\lVert w\rVert} \]
The full street width is the gap between the two margin planes — the positive half-width plus the negative half-width:
\[ \text{margin} = \frac{1}{\lVert w\rVert} + \frac{1}{\lVert w\rVert} = \frac{2}{\lVert w\rVert} \]
Missing information
Discussion prompt
We found w = [0, 0.5], so ‖w‖ = 0.5. The formula predicts margin = 2/0.5 = 4. Independently, the supports a=(0,0) and c=(0,4) are 4 apart in x₂ — those must agree. Verify both, standalone:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
2/‖w‖ = 2/0.5 = 4.0, and the two support points are √((0−0)²+(4−0)²) = 4 apart. The geometric picture and the algebra agree.
Worked example
We found w = [0, 0.5], so ‖w‖ = 0.5. The formula predicts margin = 2/0.5 = 4. Independently, the supports a=(0,0) and c=(0,4) are 4 apart in x₂ — those must agree. Verify both, standalone:
import numpy as np
w = np.array([0.0, 0.5])
x_neg = np.array([0.,0.]); x_pos = np.array([0.,4.])
margin_formula = 2 / np.linalg.norm(w)
margin_direct = np.linalg.norm(x_pos - x_neg) # raw gap between supports
print('2 / ||w|| =', margin_formula)
print('||x_pos - x_neg|| =', margin_direct)Both give 4.0 — the formula is exact
Why: 2/‖w‖ = 2/0.5 = 4.0, and the two support points are √((0−0)²+(4−0)²) = 4 apart. The geometric picture and the algebra agree.
| quantity | value |
|---|---|
| ‖w‖ | 0.5 |
| margin = 2/‖w‖ | 4.0 |
| ‖x₊ − x₋‖ (direct gap) | 4.0 |
Trade off
Comparison matrix
From Check margin = 2/‖w‖ on the toy: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| quantity | value |
|---|---|
| ‖w‖ | 0.5 |
| margin = 2/‖w‖ | 4.0 |
| ‖x₊ − x₋‖ (direct gap) | 4.0 |
Concept
Maximizing 2/‖w‖ is the same as minimizing ‖w‖, which is the same as minimizing ½‖w‖² (the square and the ½ are monotone conveniences that make the derivative clean):
\[ \max \frac{2}{\lVert w\rVert} \;\Longleftrightarrow\; \min \lVert w\rVert \;\Longleftrightarrow\; \min \tfrac{1}{2}\lVert w\rVert^2 \]
The ½‖w‖² form is a convex quadratic — a smooth bowl — so the optimizer has a single global minimum and no local traps. That is what makes the SVM's answer unique.
Worked example
Stack the objective with one inequality constraint per training point (every point must sit on the correct side of its margin plane). This is the hard-margin support vector machine:
\[ \begin{aligned} &\min_{w,\,b} \;\; \tfrac{1}{2}\lVert w\rVert^2 \\ &\text{s.t.}\;\; y_i\,(w^\top x_i + b) \ge 1 \quad \text{for every } i \end{aligned} \]
It is a convex Quadratic Program (QP)
Why: Quadratic objective, linear constraints — a QP. Convexity guarantees a global optimum, and duality (next part) reveals the support vectors. 'Hard' means we demand ≥ 1 for every point: no violations allowed, so it needs separable data.
Our toy satisfies it with w=[0,0.5], b=−1
Why: Every point scores yᵢ(wᵀxᵢ+b) = +1 exactly at the supports and stays ≥ 1 elsewhere — check c,d give +1 and a,b give +1 after the yᵢ sign flip. The margin 2/‖w‖ = 4 is the largest achievable.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Our toy satisfies it with w=[0,0.5], b=−1
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Stack the objective with one inequality constraint per training point (every point must sit on the correct side of its margin plane). This is the hard-margin support vector machine:
Anomaly
Predict first
A student writes this, and it looks reasonable:
The support points score ±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The RAW score ±1 is not a distance until you divide by ‖w‖.
Divide the ±1 scores by ‖w‖ to turn them into distances, then add the two half-widths.
Why: The RAW score ±1 is not a distance until you divide by ‖w‖. And bigger ‖w‖ makes the SAME ±1 scores correspond to a NARROWER street: margin = 2/‖w‖ shrinks as ‖w‖ grows. That is exactly why we MINIMIZE ‖w‖.
Trap
The support points score ±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.
margin = 1, or margin grows with ‖w‖ ✗
Why: The RAW score ±1 is not a distance until you divide by ‖w‖. And bigger ‖w‖ makes the SAME ±1 scores correspond to a NARROWER street: margin = 2/‖w‖ shrinks as ‖w‖ grows. That is exactly why we MINIMIZE ‖w‖.
Divide the ±1 scores by ‖w‖ to turn them into distances, then add the two half-widths.
margin = 2/‖w‖, shrinks as ‖w‖ grows ✓
Why: Half-width 1/‖w‖ on each side ⇒ full width 2/‖w‖. On the toy ‖w‖=0.5 ⇒ margin 4. Maximizing the margin ⇒ minimizing ‖w‖ ⇒ minimizing ½‖w‖². Never read a raw score as a distance.
Pattern
Predict first
The table runs: w | [0. , 0.5] | [0, 0.5] · b | −1.0 | −1 · margin | 4.0 | 4
In sklearn confirms the hand solution, given the rows so far: what is the next one — the row where quantity is n_support?
Correct: n_support | 2 | 2 (a and c side)
| quantity | sklearn | by hand |
|---|---|---|
| w | [0. , 0.5] | [0, 0.5] |
| b | −1.0 | −1 |
| margin | 4.0 | 4 |
| n_support | 2 | 2 (a and c side) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Identical to the hand derivation.
Worked example
Fit a linear SVC with a huge C (which forces hard margin) on the toy and read off w, b, the margin, and the number of support vectors. It must match our by-hand w=[0,0.5], b=−1, margin 4:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1]) # class 0 vs 1
svm = SVC(kernel='linear', C=1e6).fit(X, y) # C huge => hard margin
w = svm.coef_[0]; b = svm.intercept_[0]
print('w =', w.round(4))
print('b =', round(b, 4))
print('margin =', round(2/np.linalg.norm(w), 4))
print('n_SV =', len(svm.support_vectors_))w=[0, 0.5], b=−1.0, margin=4.0, n_SV=2
Why: Identical to the hand derivation. Only 2 support vectors — one closest point per class — pin the whole boundary; the other two points are irrelevant to it.
| quantity | sklearn | by hand |
|---|---|---|
| w | [0. , 0.5] | [0, 0.5] |
| b | −1.0 | −1 |
| margin | 4.0 | 4 |
| n_support | 2 | 2 (a and c side) |
Section
Part 2 of 6 — where SVs come from
Concept
To solve a constrained minimum we use Lagrange multipliers αᵢ ≥ 0, one per constraint yᵢ(wᵀxᵢ+b) ≥ 1. The Lagrangian folds the constraints into the objective:
\[ \mathcal{L}(w,b,\alpha) = \tfrac{1}{2}\lVert w\rVert^2 - \sum_i \alpha_i\big[\,y_i(w^\top x_i + b) - 1\,\big] \]
Each αᵢ is the 'price' of its constraint. Minimize over w, b; maximize over α ≥ 0. Setting the w- and b-gradients to zero is where the support vectors appear.
Fill the middle
Fill in the blanks
From Stationarity gives w = Σ αᵢyᵢxᵢ — finish the line. Write what belongs on the right of the equals sign before you look.
\boxed\sum_i \alpha_i y_i x_i \;}}
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Differentiate ½‖w‖² → w, and −Σαᵢyᵢ(wᵀxᵢ) → −Σαᵢyᵢxᵢ.
Worked example
∂L/∂w = 0
Why: Differentiate ½‖w‖² → w, and −Σαᵢyᵢ(wᵀxᵢ) → −Σαᵢyᵢxᵢ. Set the sum to zero.
\[ \nabla_w \mathcal{L} = w - \sum_i \alpha_i y_i x_i = 0 \]
Solve for w
Why: The optimal weight vector is a weighted sum of the training points — weighted by αᵢyᵢ. Points with αᵢ = 0 contribute nothing.
\[ \boxed{\; w = \sum_i \alpha_i y_i x_i \;} \]
∂L/∂b = 0 gives the balance condition
Why: The b-derivative of −Σαᵢyᵢb is −Σαᵢyᵢ. Setting it to zero forces the signed multipliers to cancel.
\[ \sum_i \alpha_i y_i = 0 \]
Notation
Annotate
From Stationarity gives w = Σ αᵢyᵢxᵢ — read this one piece at a time. What is each part doing?
On: \( \sum_i \alpha_i y_i = 0 \)
Concept
The complementary slackness KKT condition says, for each point, αᵢ · [yᵢ(wᵀxᵢ+b) − 1] = 0. So for every i, one of the two factors is zero:
yᵢ(wᵀxᵢ+b) > 1, so the bracket is nonzero ⇒ its αᵢ = 0 (it does nothing).= 0, so its αᵢ is allowed to be positive — it is a support vector.That is the deep reason w depends on only a handful of points: all the αᵢ = 0 terms drop out of w = Σαᵢyᵢxᵢ.
Fill the middle
Fill in the blanks
From Reconstruct w from the dual coefficients — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1])
svm = SVC(kernel='linear', C=1e6).fit(X, y)
dc = svm.dual_coef_[0] # signed alphas (alpha_i * y_i)
sv = svm.support_vectors_
w_dual = (dc[:, None] * sv).sum(axis=0)
print('support idx :', svm.support_)
print('dual_coef :', dc.round(4))
print('w (from sum):', w_dual.round(4))
print('sum alpha*y :', round(dc.sum(), 6))
Why: sv is what everything below it consumes, so the wrong expression here fails later and somewhere else. dual_coef_ = [−0.125, +0.125] on supports (2,0)?
Worked example
sklearn exposes the signed multipliers αᵢyᵢ as dual_coef_ and the support vectors as support_vectors_. Rebuild w = Σ(αᵢyᵢ)xᵢ by hand and confirm it equals coef_, and that Σαᵢyᵢ = 0. Standalone:
import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1])
svm = SVC(kernel='linear', C=1e6).fit(X, y)
dc = svm.dual_coef_[0] # signed alphas (alpha_i * y_i)
sv = svm.support_vectors_
w_dual = (dc[:, None] * sv).sum(axis=0)
print('support idx :', svm.support_)
print('dual_coef :', dc.round(4))
print('w (from sum):', w_dual.round(4))
print('sum alpha*y :', round(dc.sum(), 6))Σ(αᵢyᵢ)xᵢ = [0, 0.5] = coef_, and Σαᵢyᵢ = 0
Why: dual_coef_ = [−0.125, +0.125] on supports (2,0)? no — on the two closest points; (−0.125)·x_neg + (0.125)·x_pos rebuilds [0, 0.5] exactly. The signed alphas sum to 0, satisfying the KKT balance condition.
| quantity | value (verified) |
|---|---|
| support_ (indices) | [1, 3] |
| dual_coef_ (αᵢyᵢ) | [−0.125, +0.125] |
| w from Σ(αᵢyᵢ)xᵢ | [0. , 0.5] |
| Σ αᵢyᵢ | 0.0 |
Concept
Substituting w = Σαᵢyᵢxᵢ into the score wᵀx + b turns the decision into a sum over support vectors — the training data enters only through dot products xᵢᵀx:
\[ f(x) = \sum_{i \in \text{SV}} \alpha_i y_i \, (x_i^\top x) + b \]
This dot-product-only form is the door to the kernel trick in Part 4: replace xᵢᵀx with K(xᵢ, x) and you get a non-linear boundary for free.
Section
Part 3 of 6 — real, noisy data
Concept
The hard-margin QP demands yᵢ(wᵀxᵢ+b) ≥ 1 for every point. One noisy overlap and no (w, b) satisfies all constraints — the QP is infeasible. Real datasets almost always overlap.
The fix: let points break the margin, but charge for it. We introduce a per-point slack ξᵢ ≥ 0 that measures how far inside (or across) the margin a point sits.
Concept
Relax each constraint by its slack, and add the total slack — priced by C — to the objective:
\[ \begin{aligned} &\min_{w,\,b,\,\xi} \;\; \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \\ &\text{s.t.}\;\; y_i(w^\top x_i + b) \ge 1 - \xi_i,\;\; \xi_i \ge 0 \end{aligned} \]
ξᵢ = max(0, 1 − yᵢ(wᵀxᵢ+b)) — the hinge loss
Why: ξᵢ = 0 for a point safely beyond its margin; 0 < ξᵢ < 1 for a point inside the margin but still correct; ξᵢ > 1 for a misclassified point. The sum Σξᵢ is exactly the total hinge loss (Lesson 39).
Translation
\( \begin{aligned} &\min_{w,\,b,\,\xi} \;\; \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \\ &\text{s.t.}\;\; y_i(w^\top x_i + b) \ge 1 - \xi_i,\;\; \xi_i \ge 0 \end{aligned} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Concept
C is the price of a margin violation — the exchange rate between a wide margin (small ‖w‖) and few violations (small Σξᵢ).
| C | priority | margin | violations | bias/variance |
|---|---|---|---|---|
| small (0.01) | wide margin | wide | many allowed | high bias (regularized) |
| medium (1) | balanced | medium | some | typical default |
| large (100) | few errors | narrow | few | high variance (overfits) |
C → ∞ recovers the hard margin (no violation tolerated). C → 0 ignores the data and just shrinks w. The sweet spot is found by cross-validation.
Sorting
Sort into buckets
These are the pieces of Lesson 43: Support Vector Machines, out of order. Put each one back under the part of the lesson it belongs to.
Fill the middle
Fill in the blanks
From C sweep: accuracy and support-vector count — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
acc = accuracy_score(yte, m.predict(Xte))
print(f'C=___ acc=___ n_sv=___')
Why: n_redundant is what everything below it consumes, so the wrong expression here fails later and somewhere else. At C=0.01 the wide margin catches 120 SVs and scores 0.867; by C=1 it is 36 SVs at 0.900; C≥10 stays at 32 SVs and 0.900.
Worked example
Sweep C from 0.01 to 100 on a 200-sample 2-D task (class_sep=0.8, 30% test, standardized). Predict: does the support-vector count go up or down as C grows? Standalone:
import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
acc = accuracy_score(yte, m.predict(Xte))
print(f'C={C:6.2f} acc={acc:.3f} n_sv={len(m.support_vectors_)}')Larger C → fewer support vectors and (up to a point) higher accuracy, then it plateaus
Why: At C=0.01 the wide margin catches 120 SVs and scores 0.867; by C=1 it is 36 SVs at 0.900; C≥10 stays at 32 SVs and 0.900. Beyond C=10, nothing improves — a classic diminishing-returns plateau.
| C | test acc | n_support_vectors |
|---|---|---|
| 0.01 | 0.867 | 120 |
| 0.10 | 0.883 | 64 |
| 1.00 | 0.900 | 36 |
| 10.00 | 0.900 | 32 |
| 100.00 | 0.900 | 32 |
Pattern
Step through it
Step through C sweep: accuracy and support-vector count one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
Slack is not hidden — compute ξᵢ = max(0, 1 − yᵢf(xᵢ)) straight from the decision function at C=1, and count how many points violate the margin versus how many are actually misclassified. Standalone:
Commit before you compute: what does Read the slack directly come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: 33 points have ξ>0, but only 12 are actually misclassified
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A point with 0 < ξ < 1 is inside the margin yet on the correct side — it costs slack but is not an error.
Worked example
Slack is not hidden — compute ξᵢ = max(0, 1 − yᵢf(xᵢ)) straight from the decision function at C=1, and count how many points violate the margin versus how many are actually misclassified. Standalone:
import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
m = SVC(kernel='linear', C=1.0).fit(Xtr, ytr)
ys = np.where(ytr == 1, 1.0, -1.0)
margin = ys * m.decision_function(Xtr) # y_i (w.x_i + b)
xi = np.maximum(0.0, 1.0 - margin) # slack
print('n with xi>0 :', int((xi > 1e-9).sum()))
print('n misclassified:', int((margin < 0).sum()))
print('max slack :', round(xi.max(), 4))33 points have ξ>0, but only 12 are actually misclassified
Why: A point with 0 < ξ < 1 is inside the margin yet on the correct side — it costs slack but is not an error. So violations (33) always outnumber misclassifications (12). The worst point has slack 2.97 (well across the boundary).
| quantity | value (verified) |
|---|---|
| train points | 140 |
| points with ξᵢ > 0 (margin violators) | 33 |
| points misclassified (ξᵢ > 1) | 12 |
| max slack ξᵢ | 2.9721 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls.
Treat C as a regularization knob and tune it with cross-validation on held-out folds.
Why: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.
Trap
Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.
C = 10000 to eliminate violations ✗
Why: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.
Treat C as a regularization knob and tune it with cross-validation on held-out folds.
Grid-search C over [0.01, 0.1, 1, 10, 100] with 5-fold CV ✓
Why: CV rewards generalization, not memorization. The chosen C balances margin width against training fit — the identical bias–variance trade-off as L2 weight decay (Lesson 40). Never fix C by staring at training accuracy.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Treat C as a regularization knob and tune it with cross-validation on held-out folds.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.
Section
Part 4 of 6 — non-linear boundaries
Concept
Recall the dual decision function f(x) = Σ αᵢyᵢ(xᵢᵀx) + b. The data enters only through dot products xᵢᵀx — never through the raw coordinates on their own.
So if we had a feature map φ lifting x into a richer space, we would only ever need φ(xᵢ)ᵀφ(x) — a single number — not φ(x) itself. That is the opening the kernel trick walks through.
Concept
A kernel K(xᵢ, xⱼ) computes φ(xᵢ)ᵀφ(xⱼ) for some (possibly infinite-dimensional) φ without ever forming φ. Swap every dot product for a kernel and the SVM learns a non-linear boundary at linear cost:
\[ K_{\text{RBF}}(x_i, x_j) = \exp\!\big(-\gamma\,\lVert x_i - x_j\rVert^2\big), \qquad K_{\text{poly}}(x_i, x_j) = (x_i^\top x_j + r)^d \]
RBF (Gaussian): similarity decays with distance; gamma sets how fast (large gamma ⇒ sharp, wiggly boundary). Polynomial: degree d sets the curve order. The RBF φ is literally infinite-dimensional — yet K is one exp.
Fill the middle
Fill in the blanks
From An RBF kernel value, by hand — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.metrics.pairwise import rbf_kernel
xi = np.array([[0.0, 0.0]]); xj = np.array([[1.0, 2.0]])
gamma = 0.5
d2 = ((xi - xj) 2).sum()** # ||xi - xj||^2 = 1 + 4 = 5
by_hand = np.exp(-gamma * d2) # exp(-0.5 * 5) = exp(-2.5)
print('||xi-xj||^2 =', d2)
print('by hand =', round(by_hand, 6))
print('sklearn rbf =', round(rbf_kernel(xi, xj, gamma=gamma)[0, 0], 6))
Why: d2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. (0−1)² + (0−2)² = 1 + 4 = 5; times −γ = −2.5; exp(−2.5) = 0.082085.
Worked example
Compute K_RBF(xᵢ, xⱼ) for xᵢ=(0,0), xⱼ=(1,2), γ=0.5 by hand, then confirm against sklearn's rbf_kernel. Standalone:
import numpy as np
from sklearn.metrics.pairwise import rbf_kernel
xi = np.array([[0.0, 0.0]]); xj = np.array([[1.0, 2.0]])
gamma = 0.5
d2 = ((xi - xj) ** 2).sum() # ||xi - xj||^2 = 1 + 4 = 5
by_hand = np.exp(-gamma * d2) # exp(-0.5 * 5) = exp(-2.5)
print('||xi-xj||^2 =', d2)
print('by hand =', round(by_hand, 6))
print('sklearn rbf =', round(rbf_kernel(xi, xj, gamma=gamma)[0, 0], 6))‖xᵢ−xⱼ‖² = 5 ⇒ exp(−2.5) = 0.082085
Why: (0−1)² + (0−2)² = 1 + 4 = 5; times −γ = −2.5; exp(−2.5) = 0.082085. sklearn returns the identical value — the kernel is just this formula applied to every pair.
| quantity | value (verified) |
|---|---|
| ‖xᵢ − xⱼ‖² | 5.0 |
| −γ‖xᵢ − xⱼ‖² | −2.5 |
| K by hand = exp(−2.5) | 0.082085 |
| sklearn rbf_kernel | 0.082085 |
Discrimination
Sort into buckets
Sort these by value (verified), from memory, without looking back at An RBF kernel value, by hand. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
The moons dataset is two interleaved half-circles — no straight line separates them. Compare a linear SVM to an RBF SVM (both C=1) on 400 samples, noise=0.2, standardized. Standalone:
import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for name, m in [('linear', SVC(kernel='linear', C=1)),
('rbf', SVC(kernel='rbf', C=1, gamma='scale'))]:
m.fit(Xtr, ytr)
print(name, 'acc =', round(accuracy_score(yte, m.predict(Xte)), 4))linear 0.825 → RBF 0.950
Why: The linear boundary can't bend around the interleaving, capping it near 0.825. The RBF kernel lifts the points into a space where they ARE separable, jumping to 0.950 — a 12.5-point gain for one kernel change.
| classifier | moons test acc |
|---|---|
| Linear SVM (C=1) | 0.825 |
| RBF SVM (C=1, gamma='scale') | 0.950 |
Concept
gamma='scale' sets gamma = 1 / (n_features · Var(X)) from the data, so the kernel width adapts to the feature spread. On standardized moons that is 1/(2 · 1.0) = 0.5.
Use it as the starting point, then grid-search around it. Too-large gamma overfits (each point its own island); too-small gamma underfits (one giant blob).
Explain it
Discussion prompt
Explain gamma='scale' — a data-adaptive default to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
gamma='scale' sets gamma = 1 / (n_features · Var(X)) from the data, so the kernel width adapts to the feature spread. On standardized moons that is 1/(2 · 1.0) = 0.5.
Pattern
Predict first
The table runs: 0.1 | 10 | 0.9679 · 10 | 1 | 0.9679 · 1 | 1 | 0.9643 · 10 | 10 | 0.9500
In Grid search: tune C and gamma with CV, given the rows so far: what is the next one — the row where C is 0.1?
Correct: 0.1 | 0.1 | 0.8429
| C | gamma | 5-fold CV acc |
|---|---|---|
| 0.1 | 10 | 0.9679 |
| 10 | 1 | 0.9679 |
| 1 | 1 | 0.9643 |
| 10 | 10 | 0.9500 |
| 0.1 | 0.1 | 0.8429 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. A complex kernel (high gamma) kept in check by a wide, regularized margin (small C) generalizes best here.
Worked example
Search C ∈ {0.1, 1, 10} × gamma ∈ {0.1, 1, 10} on moons with 5-fold CV. GridSearchCV scores every combination on held-out folds and refits the best. Standalone:
import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
{'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5)
gs.fit(Xtr, ytr)
print('best :', gs.best_params_)
print('CV acc :', round(gs.best_score_, 4))
print('test :', round(accuracy_score(yte, gs.predict(Xte)), 4))Best: C=0.1, gamma=10 → CV 0.9679, test 0.9583
Why: A complex kernel (high gamma) kept in check by a wide, regularized margin (small C) generalizes best here. The grid compared all 9 cells on the folds — the student would never guess this pairing by hand.
| C | gamma | 5-fold CV acc |
|---|---|---|
| 0.1 | 10 | 0.9679 |
| 10 | 1 | 0.9679 |
| 1 | 1 | 0.9643 |
| 10 | 10 | 0.9500 |
| 0.1 | 0.1 | 0.8429 |
Pattern
Step through it
Step through Grid search: tune C and gamma with CV one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
To find the best C and gamma, try all 9 combinations and keep the one with the highest test accuracy — that's the real target metric.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise.
Use GridSearchCV with CV on the training split only; score the winner on the test set exactly once.
Why: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise. The test set must be touched exactly ONCE, after every hyperparameter is frozen.
Trap
To find the best C and gamma, try all 9 combinations and keep the one with the highest test accuracy — that's the real target metric.
Evaluate every combo on the test set, pick the winner ✗
Why: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise. The test set must be touched exactly ONCE, after every hyperparameter is frozen.
Use GridSearchCV with CV on the training split only; score the winner on the test set exactly once.
gs.fit(Xtr, ytr); test = gs.score(Xte, yte) ✓
Why: Cross-validation resamples inside the training data — no test information leaks in. The single final test evaluation gives an honest generalization estimate. This is the same discipline as every tuning workflow in the course.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:; Because scaling (w, b) doesn't move the boundary, we are free to fix the scale. The SVM convention pins it so the closest points score exactly ±1:±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.; Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.Section
Part 5 of 6
Concept
An SVM is binary. For K classes we combine binary SVMs two ways: One-vs-Rest (OvR) trains K classifiers (each class vs all others); One-vs-One (OvO) trains K(K−1)/2 classifiers (every class pair).
| strategy | # classifiers (K=3) | # classifiers (K=10) | decision rule |
|---|---|---|---|
| OvR | 3 | 10 | highest score wins |
| OvO | 3 | 45 | majority vote over pairs |
sklearn's SVC always fits OvO internally (each binary problem is smaller, so it's fast for moderate K). decision_function_shape only changes how the scores are reported, not which classifiers are trained.
Comparison
Comparison matrix
From Multi-class SVM: OvO and OvR: refill the decision rule column from what you know. The rest of the table is as it appeared.
| strategy | # classifiers (K=3) | # classifiers (K=10) | decision rule |
|---|---|---|---|
| OvR | 3 | 10 | highest score wins |
| OvO | 3 | 45 | majority vote over pairs |
Missing information
Discussion prompt
Fit an RBF SVM on Iris (150 samples, 4 features, 3 classes) under both decision_function_shape settings. Predict: will accuracy differ? Standalone:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The underlying classifiers are the same OvO set; only the reported decision scores reshape. Predictions are identical. (For K=3 the shapes even coincide: K(K−1)/2 = 3 = K — a numerical accident, but the score VALUES still differ.)
Worked example
Fit an RBF SVM on Iris (150 samples, 4 features, 3 classes) under both decision_function_shape settings. Predict: will accuracy differ? Standalone:
import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for shape in ['ovo', 'ovr']:
m = SVC(kernel='rbf', C=1, gamma='scale',
decision_function_shape=shape).fit(Xtr, ytr)
print(shape, 'acc =', round(accuracy_score(yte, m.predict(Xte)), 4),
' n_sv =', m.n_support_)Both give acc=0.9778, n_sv=[7,18,17] — identical fit
Why: The underlying classifiers are the same OvO set; only the reported decision scores reshape. Predictions are identical. (For K=3 the shapes even coincide: K(K−1)/2 = 3 = K — a numerical accident, but the score VALUES still differ.)
| class | n_support_vectors |
|---|---|
| setosa (0) | 7 |
| versicolor (1) | 18 |
| virginica (2) | 17 |
| total | 42 |
Pattern
Step through it
Step through OvO vs OvR on Iris (K=3) one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Both draw linear boundaries in their base form, but they optimize different losses and offer different extras. Pick by what the task needs:
| criterion | prefer SVM | prefer Logistic Regression |
|---|---|---|
| non-linearity | free via kernel (RBF/poly) | needs hand-built features |
| probabilities | not native (no P(y|x)) | native, calibrated log-odds |
| interpretability | less (dual weights) | more (coeff = log-odds) |
| scale (large n) | slow past ~100k rows | fast; supports online SGD |
| margin / robustness | maximizes the margin | no explicit margin |
Rule of thumb: start with LR (fast, interpretable, gives probabilities). Upgrade to an RBF SVM when training accuracy is high but test accuracy lags — the signature of a needed non-linear boundary.
Analogy
Discussion prompt
Explain SVM vs logistic regression by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Both draw linear boundaries in their base form, but they optimize different losses and offer different extras. Pick by what the task needs:
Estimation
Predict first
On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:
Commit before you compute: what does SVM vs LR on moons — the non-linear gap come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: LogReg 0.8083 vs RBF SVM 0.9500
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Logistic regression is stuck with a straight boundary (0.8083).
Worked example
On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:
import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
lr = LogisticRegression().fit(Xtr, ytr)
sv = SVC(kernel='rbf', C=1, gamma='scale').fit(Xtr, ytr)
print('LogReg acc =', round(accuracy_score(yte, lr.predict(Xte)), 4))
print('RBF SVM acc =', round(accuracy_score(yte, sv.predict(Xte)), 4))LogReg 0.8083 vs RBF SVM 0.9500
Why: Logistic regression is stuck with a straight boundary (0.8083). The RBF SVM curves around the interleaving (0.9500) — a 14-point gap that is entirely the kernel. On linearly separable data the two would tie.
| classifier | moons test acc | boundary |
|---|---|---|
| Logistic Regression | 0.8083 | linear |
| RBF SVM | 0.9500 | non-linear (kernel) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
LogReg 0.8083 vs RBF SVM 0.9500
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:
Ranking
Put in order
These are the steps of The SVM workflow recipe, scrambled. Put them back in order before the next slide shows you.
StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)SVC(kernel='linear', C=1) to test whether the data is linearly separableSVC(kernel='rbf', C=1, gamma='scale') for non-linear boundariesC and gamma with GridSearchCV(cv=5) on training data onlySVC does OvO automatically; LinearSVC does OvR by defaultWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)SVC(kernel='linear', C=1) to test whether the data is linearly separableSVC(kernel='rbf', C=1, gamma='scale') for non-linear boundariesC and gamma with GridSearchCV(cv=5) on training data onlySVC does OvO automatically; LinearSVC does OvR by defaultEdge cases
Discussion prompt
The SVM workflow recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)SVC(kernel='linear', C=1) to test whether the data is linearly separableSVC(kernel='rbf', C=1, gamma='scale') for non-linear boundariesC and gamma with GridSearchCV(cv=5) on training data onlySVC does OvO automatically; LinearSVC does OvR by defaultElimination
Eliminate the wrong options
An SVM is scaled so the closest points score ±1 and the fitted weights have ‖w‖ = 0.25. What is the margin (full street width)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The margin is 2/‖w‖ = 2/0.25 = 8. Each support point sits at signed distance 1/‖w‖ = 4 from the boundary; the two half-widths add to 8.
Check
Reason from the derivation, not from memory.
Check your understanding
An SVM is scaled so the closest points score ±1 and the fitted weights have ‖w‖ = 0.25. What is the margin (full street width)?
Answer: A
Why: The margin is 2/‖w‖ = 2/0.25 = 8. Each support point sits at signed distance 1/‖w‖ = 4 from the boundary; the two half-widths add to 8.
Prediction
Predict first
You remove a training point that is NOT a support vector, then refit a hard-margin SVM. What happens to the boundary?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Nothing — non-support-vectors have αᵢ = 0 and don't affect w or b
Why: By complementary slackness, a non-support-vector satisfies yᵢ(wᵀxᵢ+b) > 1, so its αᵢ = 0. It contributes nothing to w = Σαᵢyᵢxᵢ, so removing it leaves w, b, and the margin unchanged.
Check
Think about the KKT conditions and w = Σαᵢyᵢxᵢ.
Check your understanding
You remove a training point that is NOT a support vector, then refit a hard-margin SVM. What happens to the boundary?
Answer: A
Why: By complementary slackness, a non-support-vector satisfies yᵢ(wᵀxᵢ+b) > 1, so its αᵢ = 0. It contributes nothing to w = Σαᵢyᵢxᵢ, so removing it leaves w, b, and the margin unchanged.
Prediction
Predict first
An RBF SVM overfits: train acc 0.99, test acc 0.75. Which change is most likely to help?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Decrease C and/or decrease gamma
Why: Overfitting means the model is too complex. Lowering C widens the margin (more regularization); lowering gamma broadens each RBF bump (smoother boundary). Both raise bias and cut variance — exactly the correction needed.
Check
Reason about bias and variance.
Check your understanding
An RBF SVM overfits: train acc 0.99, test acc 0.75. Which change is most likely to help?
Answer: A
Why: Overfitting means the model is too complex. Lowering C widens the margin (more regularization); lowering gamma broadens each RBF bump (smoother boundary). Both raise bias and cut variance — exactly the correction needed.
Elimination
Eliminate the wrong options
The kernel trick works because the SVM's dual decision function depends on the data only through:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Substituting w = Σαᵢyᵢxᵢ gives f(x) = Σαᵢyᵢ(xᵢᵀx) + b — the inputs appear only inside dot products. Replacing each xᵢᵀx with K(xᵢ, x) = φ(xᵢ)ᵀφ(x) uses a high-dimensional map φ without ever computing it.
Check
Why is the trick even possible?
Check your understanding
The kernel trick works because the SVM's dual decision function depends on the data only through:
Answer: A
Why: Substituting w = Σαᵢyᵢxᵢ gives f(x) = Σαᵢyᵢ(xᵢᵀx) + b — the inputs appear only inside dot products. Replacing each xᵢᵀx with K(xᵢ, x) = φ(xᵢ)ᵀφ(x) uses a high-dimensional map φ without ever computing it.
Section
Part 6 of 6 — the project
Concept
Build the full workflow on Iris: scale → linear baseline → RBF with grid search → compare to logistic regression → report test accuracy once. You've derived every piece — now wire it together.
| # | milestone | tool |
|---|---|---|
| 1 | scale + linear SVM baseline | StandardScaler + SVC(kernel='linear') |
| 2 | RBF SVM with grid search | GridSearchCV, cv=5, C+gamma |
| 3 | compare to LogisticRegression | sklearn.linear_model.LogisticRegression |
Build rules: scale before fitting; call GridSearchCV.fit on the training split only; touch the test set exactly once at the very end.
Counterexample
Discussion prompt
Build the full workflow on Iris: scale → linear baseline → RBF with grid search → compare to logistic regression → report test accuracy once. You've derived every piece — now wire it together.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: scale before fitting; call GridSearchCV.fit on the training split only; touch the test set exactly once at the very end.
Worked example
Your turn: scale Iris, fit a linear SVM at C=1, and record test accuracy and support-vector count per class. Predict the accuracy before you print.
Hint: split with random_state=0, test_size=0.3; StandardScaler().fit_transform(Xtr) then .transform(Xte) — never fit the scaler on test data.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
svm = SVC(kernel='linear', C=1).fit(Xtr, ytr)
print('acc :', round(accuracy_score(yte, svm.predict(Xte)), 4))
print('n_sv:', svm.n_support_)| metric | value (verified) |
|---|---|
| test accuracy | 0.9778 |
| n_sv setosa (0) | 2 |
| n_sv versicolor (1) | 12 |
| n_sv virginica (2) | 10 |
Worked example
Your turn: grid-search C ∈ {0.1, 1, 10} × gamma ∈ {0.1, 1, 10} with 5-fold CV on the training split. Predict whether RBF will beat the linear baseline on Iris.
Hint: GridSearchCV(SVC(kernel='rbf'), {...}, cv=5).fit(Xtr, ytr), then read gs.best_params_ and gs.score(Xte, yte).
import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
{'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
print('best params:', gs.best_params_)
print('CV acc :', round(gs.best_score_, 4))
print('test acc :', round(gs.score(Xte, yte), 4))| setting | CV acc | test acc |
|---|---|---|
| linear SVM, C=1 (baseline) | — | 0.9778 |
| RBF, best {C:10, gamma:0.1} | 0.9524 | 0.9778 |
Worked example
Your turn: fit logistic regression on the same scaled split and compare to the grid-searched SVM. Predict which wins on Iris and why.
Hint: reuse the exact Xtr, Xte from Milestone 2; LogisticRegression(C=1, max_iter=200, random_state=0). Iris is nearly linearly separable, so think about whether the kernel can even help.
import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
{'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
lr = LogisticRegression(C=1, max_iter=200, random_state=0).fit(Xtr, ytr)
print('SVM test acc:', round(gs.score(Xte, yte), 4))
print('LR test acc:', round(accuracy_score(yte, lr.predict(Xte)), 4))| classifier | test acc | probabilities? |
|---|---|---|
| Logistic Regression | 0.9778 | yes (native) |
| RBF SVM (grid-searched) | 0.9778 | no (by default) |
Trade off
Comparison matrix
From Milestone 3 — compare to logistic regression: every row here is a choice with a cost. Fill the probabilities? column, then say which row you would actually pick and what you give up for it.
| classifier | test acc | probabilities? |
|---|---|---|
| Logistic Regression | 0.9778 | yes (native) |
| RBF SVM (grid-searched) | 0.9778 | no (by default) |
Concept
import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
{'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
lr = LogisticRegression(C=1, max_iter=200, random_state=0).fit(Xtr, ytr)
print('SVM best:', gs.best_params_, 'test acc:', round(gs.score(Xte, yte), 4))
print('LR test acc:', round(accuracy_score(yte, lr.predict(Xte)), 4))| printed line | value (verified) |
|---|---|
| SVM best params | {'C': 10, 'gamma': 0.1} |
| SVM test acc | 0.9778 |
| LR test acc | 0.9778 |
| verdict | tie on Iris — LR preferred for native probabilities |
On Iris both hit 97.8% — the data is nearly linear, so the kernel earns nothing. The SVM's edge shows on moons (SVM 0.950 vs LR 0.808). Reach for LR when you need calibrated probabilities; reach for the RBF SVM when the boundary must bend.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| SVM best params | {'C': 10, 'gamma': 0.1} |
| SVM test acc | 0.9778 |
| LR test acc | 0.9778 |
| verdict | tie on Iris — LR preferred for native probabilities |
Concept
Slides closed, out loud: explain (1) how margin = 2/‖w‖ falls out of the signed-distance formula and the ±1 scaling, (2) why the KKT conditions force αᵢ = 0 for every non-support-vector, and (3) how replacing xᵢᵀx with K(xᵢ, x) gives a non-linear boundary without ever computing φ(x).
Stretch (homework): plot the moons decision boundary for gamma ∈ {0.1, 1, 10} and watch it go from underfit to overfit; derive the full SVM dual objective Σαᵢ − ½Σᵢⱼαᵢαⱼyᵢyⱼ(xᵢᵀxⱼ); and confirm that dropping a support vector does move the boundary. Next up: decision trees and random forests.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The maximum-margin hyperplane · The dual & support vectors · Soft margin: slack and C · The kernel trick & grid search · Multi-class & SVM vs LR · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
margin = 2/‖w‖ from the signed-distance formula and the ±1 canonical scaling — no skipped stepsmin ½‖w‖² and its Lagrangian/KKT, and rebuild w = Σαᵢyᵢxᵢ from sklearn's dual coefficientsξᵢ and explain how C trades margin width against violations (small C = wide/regularized, large C = narrow/overfit)C/gamma with GridSearchCV on training data only| idea | the one thing to remember |
|---|---|
| margin | 2/‖w‖ — divide the ±1 score by ‖w‖; minimize ½‖w‖² |
| support vectors | only αᵢ > 0 points define w; KKT zeroes the rest |
| soft margin C | large C → narrow, overfits; small C → wide, regularized |
| kernel trick | swap xᵢᵀx for K(xᵢ,x) — never compute φ(x) |
| workflow | scale, grid-search on train CV, touch test once |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.