Lesson 43: Support Vector Machines

USAAIO Lesson 43, from Week 15 of Phase 2, fully worked. It derives the maximum-margin hyperplane from the signed-distance formula with no steps skipped, explains the ±1 canonical scaling, and proves that the margin is 2/‖w‖ on a hand-solvable four-point toy. It then sets up the hard-margin QP with its Lagrangian and KKT conditions and verifies w = Σαᵢyᵢxᵢ against sklearn's dual coefficients. From there it covers soft-margin slack and the exact role of C, the kernel trick with an RBF value checked by hand, grid search with cross-validation, one-versus-one against one-versus-rest for multi-class, and SVM against logistic regression, ending with a from-scratch Iris pipeline. Every snippet runs standalone, and every number came from real execution. The lesson runs to 64 slides.

Subject: Machine Learning · 104 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Support Vector Machines

Title

USAAIO · Lesson 43 · Week 15 (Phase 2)

We derive the maximum-margin hyperplane from the signed-distance formula — every step — prove margin = 2/‖w‖ on a 4-point toy you can solve by hand, turn it into the hard-margin QP, read off support vectors from the dual, then soften it with slack, kernelize it, and tune it. Every number comes from real execution.

2. By the end of this lesson you can

Objectives

  1. Write the signed distance from a point to a hyperplane and derive margin = 2/‖w‖ with no skipped steps
  2. Turn 'maximize the margin' into the hard-margin QP min ½‖w‖² and read off the support vectors
  3. State the Lagrangian / KKT conditions and verify w = Σᵢ αᵢyᵢxᵢ against sklearn's dual coefficients
  4. Add slack ξᵢ and explain exactly how C trades margin width against violations
  5. Apply the kernel trick (RBF, polynomial), tune C/gamma with grid search + CV, and choose OvO/OvR and SVM vs logistic regression

3. What survived from PyTorch Data Pipeline & Model Serialization?

Warm-up

Discussion prompt

Before we open Lesson 43: Support Vector Machines: without looking back, what was the main idea of PyTorch Data Pipeline & Model Serialization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

custom Dataset (__len__, __getitem__), DataLoader (batching, shuffling, num_workers), manual feature transforms, nn.Module parameter inspection (.parameters(), .named_parameters()), and model serialization (state_dict vs full-model save, checkpoint save/resume). Build a complete pipeline: 120-sample tabular dataset, TwoLayerNet(4→8→1), train with Adam, save checkpoint at epoch 5, resume and reach loss 0.1666 at epoch 10.

4. The maximum-margin hyperplane

Section

Part 1 of 6 — the geometry

5. A hyperplane splits the space

Concept

A linear classifier draws a flat boundary — a hyperplane — and predicts by which side a point falls on. In d dimensions the boundary is the set of points where a linear score is zero:

\[ f(x) = w^\top x + b = 0 \]

w is the weight (normal) vector — it points perpendicular to the boundary. b is the bias that shifts the boundary off the origin. We predict class +1 when f(x) > 0 and class −1 when f(x) < 0.

6. Break it if you can: A hyperplane splits the space

Counterexample

Discussion prompt

A linear classifier draws a flat boundary — a hyperplane — and predicts by which side a point falls on. In d dimensions the boundary is the set of points where a linear score is zero:

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Picture it first: Many lines separate — which is best?

Picture it

Figure (svg): Two clusters of points, one class lower-left and one upper-right, with three different straight lines all separating them; one line runs down the middle of the gap, the other two hug one cluster.

All three separate the training data perfectly — but only the middle one keeps its distance from both clouds.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

If two classes are linearly separable, infinitely many hyperplanes split them perfectly. A perceptron stops at the first one it finds. The SVM asks a sharper question: which separator is safest?

8. Many lines separate — which is best?

Concept

If two classes are linearly separable, infinitely many hyperplanes split them perfectly. A perceptron stops at the first one it finds. The SVM asks a sharper question: which separator is safest?

Figure (svg): Two clusters of points, one class lower-left and one upper-right, with three different straight lines all separating them; one line runs down the middle of the gap, the other two hug one cluster.

All three separate the training data perfectly — but only the middle one keeps its distance from both clouds.

9. Widest street wins

Intuition

Picture the boundary as a street you widen until it touches the nearest point on each side. A separator squeezed right up against one cloud will misjudge a slightly-shifted test point; a separator down the middle of the widest street has the most slack before a new point crosses.

The half-width of that street is the margin. The SVM is the unique separator that makes the margin as large as possible — so 'safest' becomes a precise, solvable objective.

10. By analogy: Widest street wins

Analogy

Discussion prompt

Explain Widest street wins by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The half-width of that street is the margin. The SVM is the unique separator that makes the margin as large as possible — so 'safest' becomes a precise, solvable objective.

11. Our hand-solvable toy

Concept

We carry one dataset through the whole geometry: four points, two per class, chosen so every number is clean and checkable by hand and against sklearn.

pointx₁x₂class y
a00−1
b20−1
c04+1
d24+1

The two classes are stacked: the negatives sit at height x₂ = 0, the positives at x₂ = 4. By symmetry the best boundary is the horizontal line x₂ = 2 — halfway between. The whole point is to derive that, then read w and b off it.

12. Fill in: x₁ for Our hand-solvable toy

Comparison

Comparison matrix

From Our hand-solvable toy: refill the x₁ column from what you know. The rest of the table is as it appeared.

pointx₁x₂class y
a00−1
b20−1
c04+1
d24+1

13. Signed distance from a point to the plane

Concept

The perpendicular distance from a point x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:

\[ \operatorname{dist}(x_0) = \frac{w^\top x_0 + b}{\lVert w \rVert} \]

The sign tells you the side; the magnitude tells you how far. Dividing by ‖w‖ removes the arbitrary scale of w — double w and b and the boundary is unchanged, so the true distance can't depend on that scale.

14. Teach it back: Signed distance from a point to the plane

Explain it

Discussion prompt

Explain Signed distance from a point to the plane to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The perpendicular distance from a point x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:

15. Guess the shape of the answer: Signed distance, computed on the toy

Estimation

Predict first

Take the boundary x₂ = 2, written as wᵀx + b = 0 with w = [0, ½], b = −1 (so ½·x₂ − 1 = 0 ⇒ x₂ = 2). Then ‖w‖ = ½. Verify the distances match reality — this snippet runs on its own:

Commit before you compute: what does Signed distance, computed on the toy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Negatives land at distance −2, positives at +2

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. For a=(0,0): (0·0 + ½·0 − 1)/½ = −1/0.5 = −2.

16. Signed distance, computed on the toy

Worked example

Take the boundary x₂ = 2, written as wᵀx + b = 0 with w = [0, ½], b = −1 (so ½·x₂ − 1 = 0 ⇒ x₂ = 2). Then ‖w‖ = ½. Verify the distances match reality — this snippet runs on its own:

import numpy as np
w = np.array([0.0, 0.5]); b = -1.0
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])   # a, b, c, d
for x in X:
    signed = (w @ x + b) / np.linalg.norm(w)
    print(x, 'signed dist =', round(signed, 3))

Negatives land at distance −2, positives at +2

Why: For a=(0,0): (0·0 + ½·0 − 1)/½ = −1/0.5 = −2. For c=(0,4): (½·4 − 1)/½ = 1/0.5 = +2. Both classes are exactly 2 units off the boundary — the line really is centered.

pointwᵀx + b÷ ‖w‖ = signed dist
a (0,0)−1.0−2.0
b (2,0)−1.0−2.0
c (0,4)+1.0+2.0
d (2,4)+1.0+2.0

17. Which is which, by wᵀx + b

Discrimination

Sort into buckets

Sort these by wᵀx + b, from memory, without looking back at Signed distance, computed on the toy. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

−1.0
a (0,0); b (2,0)
+1.0
c (0,4); d (2,4)
g1
wᵀx + b is "−1.0" for a (0,0), b (2,0) — that is what the table on "Signed distance, computed on the toy" records, and it is the single property separating this group from the rest.
g2
wᵀx + b is "+1.0" for c (0,4), d (2,4) — that is what the table on "Signed distance, computed on the toy" records, and it is the single property separating this group from the rest.

18. The ±1 canonical scaling

Concept

Because scaling (w, b) doesn't move the boundary, we are free to fix the scale. The SVM convention pins it so the closest points score exactly ±1:

\[ \min_i \; \lvert w^\top x_i + b \rvert = 1 \]

With that choice every point obeys yᵢ(wᵀxᵢ + b) ≥ 1, and the two closest points sit on the margin planes wᵀx + b = +1 and wᵀx + b = −1. Those points are the support vectors.

19. What has to happen first: Derive w and b from the ±1 condition

Ranking

Put in order

Put the moves of Derive w and b from the ±1 condition into the order they have to happen.

  1. Negative support: w₂·0 + b = −1 ⇒ b = −1
  2. Positive support: w₂·4 + b = +1 ⇒ 4w₂ − 1 = 1
  3. Solve: 4w₂ = 2 ⇒ w₂ = 0.5

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Point a=(0,0) sits on the wᵀx+b = −1 margin plane.

20. Derive w and b from the ±1 condition

Worked example

Force the closest negative a=(0,0) to score −1 and the closest positive c=(0,4) to score +1. By the toy's symmetry w = [0, w₂] (the boundary is horizontal), so only w₂ and b are unknown:

Negative support: w₂·0 + b = −1 ⇒ b = −1

Why: Point a=(0,0) sits on the wᵀx+b = −1 margin plane. Its x₂ is 0, so the whole score is just b.

Positive support: w₂·4 + b = +1 ⇒ 4w₂ − 1 = 1

Why: Point c=(0,4) sits on the +1 plane. Substitute b = −1 from the previous step.

Solve: 4w₂ = 2 ⇒ w₂ = 0.5

Why: So w = [0, 0.5], b = −1. The boundary wᵀx+b = 0 is ½x₂ − 1 = 0, i.e. x₂ = 2 — exactly the midline we predicted.

\[ w = \begin{bmatrix} 0 \\ 0.5 \end{bmatrix}, \qquad b = -1 \]

21. Decode the notation: Derive w and b from the ±1 condition

Notation

Annotate

From Derive w and b from the ±1 condition — read this one piece at a time. What is each part doing?

On: \( w = \begin{bmatrix} 0 \\ 0.5 \end{bmatrix}, \qquad b = -1 \)

  • Point a=(0,0) sits on the wᵀx+b = −1 margin plane. Its x₂ is 0, so the whole score is just b.
  • Point c=(0,4) sits on the +1 plane. Substitute b = −1 from the previous step.
  • So w = [0, 0.5], b = −1. The boundary wᵀx+b = 0 is ½x₂ − 1 = 0, i.e. x₂ = 2 — exactly the midline we predicted.

22. Why the margin is 2/‖w‖

Concept

A positive support x₊ scores +1 and a negative support x₋ scores −1. Their signed distances to the boundary are therefore:

\[ \frac{w^\top x_+ + b}{\lVert w\rVert} = \frac{+1}{\lVert w\rVert}, \qquad \frac{w^\top x_- + b}{\lVert w\rVert} = \frac{-1}{\lVert w\rVert} \]

The full street width is the gap between the two margin planes — the positive half-width plus the negative half-width:

\[ \text{margin} = \frac{1}{\lVert w\rVert} + \frac{1}{\lVert w\rVert} = \frac{2}{\lVert w\rVert} \]

23. What has to be given first: Check margin = 2/‖w‖ on the toy

Missing information

Discussion prompt

We found w = [0, 0.5], so ‖w‖ = 0.5. The formula predicts margin = 2/0.5 = 4. Independently, the supports a=(0,0) and c=(0,4) are 4 apart in x₂ — those must agree. Verify both, standalone:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

2/‖w‖ = 2/0.5 = 4.0, and the two support points are √((0−0)²+(4−0)²) = 4 apart. The geometric picture and the algebra agree.

24. Check margin = 2/‖w‖ on the toy

Worked example

We found w = [0, 0.5], so ‖w‖ = 0.5. The formula predicts margin = 2/0.5 = 4. Independently, the supports a=(0,0) and c=(0,4) are 4 apart in x₂ — those must agree. Verify both, standalone:

import numpy as np
w = np.array([0.0, 0.5])
x_neg = np.array([0.,0.]); x_pos = np.array([0.,4.])
margin_formula = 2 / np.linalg.norm(w)
margin_direct  = np.linalg.norm(x_pos - x_neg)   # raw gap between supports
print('2 / ||w||         =', margin_formula)
print('||x_pos - x_neg|| =', margin_direct)

Both give 4.0 — the formula is exact

Why: 2/‖w‖ = 2/0.5 = 4.0, and the two support points are √((0−0)²+(4−0)²) = 4 apart. The geometric picture and the algebra agree.

quantityvalue
‖w‖0.5
margin = 2/‖w‖4.0
‖x₊ − x₋‖ (direct gap)4.0

25. What each one costs: Check margin = 2/‖w‖ on the toy

Trade off

Comparison matrix

From Check margin = 2/‖w‖ on the toy: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

quantityvalue
‖w‖0.5
margin = 2/‖w‖4.0
‖x₊ − x₋‖ (direct gap)4.0

26. Maximize margin ⇔ minimize ½‖w‖²

Concept

Maximizing 2/‖w‖ is the same as minimizing ‖w‖, which is the same as minimizing ½‖w‖² (the square and the ½ are monotone conveniences that make the derivative clean):

\[ \max \frac{2}{\lVert w\rVert} \;\Longleftrightarrow\; \min \lVert w\rVert \;\Longleftrightarrow\; \min \tfrac{1}{2}\lVert w\rVert^2 \]

The ½‖w‖² form is a convex quadratic — a smooth bowl — so the optimizer has a single global minimum and no local traps. That is what makes the SVM's answer unique.

27. The hard-margin QP, written in full

Worked example

Stack the objective with one inequality constraint per training point (every point must sit on the correct side of its margin plane). This is the hard-margin support vector machine:

\[ \begin{aligned} &\min_{w,\,b} \;\; \tfrac{1}{2}\lVert w\rVert^2 \\ &\text{s.t.}\;\; y_i\,(w^\top x_i + b) \ge 1 \quad \text{for every } i \end{aligned} \]

It is a convex Quadratic Program (QP)

Why: Quadratic objective, linear constraints — a QP. Convexity guarantees a global optimum, and duality (next part) reveals the support vectors. 'Hard' means we demand ≥ 1 for every point: no violations allowed, so it needs separable data.

Our toy satisfies it with w=[0,0.5], b=−1

Why: Every point scores yᵢ(wᵀxᵢ+b) = +1 exactly at the supports and stays ≥ 1 elsewhere — check c,d give +1 and a,b give +1 after the yᵢ sign flip. The margin 2/‖w‖ = 4 is the largest achievable.

28. Work backwards from the answer: The hard-margin QP, written in full

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Our toy satisfies it with w=[0,0.5], b=−1

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Stack the objective with one inequality constraint per training point (every point must sit on the correct side of its margin plane). This is the hard-margin support vector machine:

29. Something is wrong here: the margin is not 1/‖w‖ and not ‖w‖

Anomaly

Predict first

A student writes this, and it looks reasonable:

The support points score ±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The RAW score ±1 is not a distance until you divide by ‖w‖.

Divide the ±1 scores by ‖w‖ to turn them into distances, then add the two half-widths.

Why: The RAW score ±1 is not a distance until you divide by ‖w‖. And bigger ‖w‖ makes the SAME ±1 scores correspond to a NARROWER street: margin = 2/‖w‖ shrinks as ‖w‖ grows. That is exactly why we MINIMIZE ‖w‖.

30. Trap: the margin is not 1/‖w‖ and not ‖w‖

Trap

The trap

The support points score ±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.

margin = 1, or margin grows with ‖w‖ ✗

Why: The RAW score ±1 is not a distance until you divide by ‖w‖. And bigger ‖w‖ makes the SAME ±1 scores correspond to a NARROWER street: margin = 2/‖w‖ shrinks as ‖w‖ grows. That is exactly why we MINIMIZE ‖w‖.

The fix

Divide the ±1 scores by ‖w‖ to turn them into distances, then add the two half-widths.

margin = 2/‖w‖, shrinks as ‖w‖ grows ✓

Why: Half-width 1/‖w‖ on each side ⇒ full width 2/‖w‖. On the toy ‖w‖=0.5 ⇒ margin 4. Maximizing the margin ⇒ minimizing ‖w‖ ⇒ minimizing ½‖w‖². Never read a raw score as a distance.

31. Predict the next row: sklearn confirms the hand solution

Pattern

Predict first

The table runs: w | [0. , 0.5] | [0, 0.5] · b | −1.0 | −1 · margin | 4.0 | 4

In sklearn confirms the hand solution, given the rows so far: what is the next one — the row where quantity is n_support?

Correct: n_support | 2 | 2 (a and c side)

quantitysklearnby hand
w[0. , 0.5][0, 0.5]
b−1.0−1
margin4.04
n_support22 (a and c side)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Identical to the hand derivation.

32. sklearn confirms the hand solution

Worked example

Fit a linear SVC with a huge C (which forces hard margin) on the toy and read off w, b, the margin, and the number of support vectors. It must match our by-hand w=[0,0.5], b=−1, margin 4:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1])                    # class 0 vs 1
svm = SVC(kernel='linear', C=1e6).fit(X, y)   # C huge => hard margin
w = svm.coef_[0]; b = svm.intercept_[0]
print('w      =', w.round(4))
print('b      =', round(b, 4))
print('margin =', round(2/np.linalg.norm(w), 4))
print('n_SV   =', len(svm.support_vectors_))

w=[0, 0.5], b=−1.0, margin=4.0, n_SV=2

Why: Identical to the hand derivation. Only 2 support vectors — one closest point per class — pin the whole boundary; the other two points are irrelevant to it.

quantitysklearnby hand
w[0. , 0.5][0, 0.5]
b−1.0−1
margin4.04
n_support22 (a and c side)

33. The dual & support vectors

Section

Part 2 of 6 — where SVs come from

34. Attach a multiplier to each constraint

Concept

To solve a constrained minimum we use Lagrange multipliers αᵢ ≥ 0, one per constraint yᵢ(wᵀxᵢ+b) ≥ 1. The Lagrangian folds the constraints into the objective:

\[ \mathcal{L}(w,b,\alpha) = \tfrac{1}{2}\lVert w\rVert^2 - \sum_i \alpha_i\big[\,y_i(w^\top x_i + b) - 1\,\big] \]

Each αᵢ is the 'price' of its constraint. Minimize over w, b; maximize over α ≥ 0. Setting the w- and b-gradients to zero is where the support vectors appear.

35. Complete the line: Stationarity gives w = Σ αᵢyᵢxᵢ

Fill the middle

Fill in the blanks

From Stationarity gives w = Σ αᵢyᵢxᵢ — finish the line. Write what belongs on the right of the equals sign before you look.

\boxed\sum_i \alpha_i y_i x_i \;}}

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. Differentiate ½‖w‖² → w, and −Σαᵢyᵢ(wᵀxᵢ) → −Σαᵢyᵢxᵢ.

36. Stationarity gives w = Σ αᵢyᵢxᵢ

Worked example

∂L/∂w = 0

Why: Differentiate ½‖w‖² → w, and −Σαᵢyᵢ(wᵀxᵢ) → −Σαᵢyᵢxᵢ. Set the sum to zero.

\[ \nabla_w \mathcal{L} = w - \sum_i \alpha_i y_i x_i = 0 \]

Solve for w

Why: The optimal weight vector is a weighted sum of the training points — weighted by αᵢyᵢ. Points with αᵢ = 0 contribute nothing.

\[ \boxed{\; w = \sum_i \alpha_i y_i x_i \;} \]

∂L/∂b = 0 gives the balance condition

Why: The b-derivative of −Σαᵢyᵢb is −Σαᵢyᵢ. Setting it to zero forces the signed multipliers to cancel.

\[ \sum_i \alpha_i y_i = 0 \]

37. Decode the notation: Stationarity gives w = Σ αᵢyᵢxᵢ

Notation

Annotate

From Stationarity gives w = Σ αᵢyᵢxᵢ — read this one piece at a time. What is each part doing?

On: \( \sum_i \alpha_i y_i = 0 \)

  • Differentiate ½‖w‖² → w, and −Σαᵢyᵢ(wᵀxᵢ) → −Σαᵢyᵢxᵢ. Set the sum to zero.
  • The optimal weight vector is a weighted sum of the training points — weighted by αᵢyᵢ. Points with αᵢ = 0 contribute nothing.
  • The b-derivative of −Σαᵢyᵢb is −Σαᵢyᵢ. Setting it to zero forces the signed multipliers to cancel.

38. KKT: αᵢ > 0 only for support vectors

Concept

The complementary slackness KKT condition says, for each point, αᵢ · [yᵢ(wᵀxᵢ+b) − 1] = 0. So for every i, one of the two factors is zero:

That is the deep reason w depends on only a handful of points: all the αᵢ = 0 terms drop out of w = Σαᵢyᵢxᵢ.

39. Restore the missing line: Reconstruct w from the dual coefficients

Fill the middle

Fill in the blanks

From Reconstruct w from the dual coefficients — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1])
svm = SVC(kernel='linear', C=1e6).fit(X, y)
dc = svm.dual_coef_[0] # signed alphas (alpha_i * y_i)
sv = svm.support_vectors_
w_dual = (dc[:, None] * sv).sum(axis=0)
print('support idx :', svm.support_)
print('dual_coef :', dc.round(4))
print('w (from sum):', w_dual.round(4))
print('sum alpha*y :', round(dc.sum(), 6))

Why: sv is what everything below it consumes, so the wrong expression here fails later and somewhere else. dual_coef_ = [−0.125, +0.125] on supports (2,0)?

40. Reconstruct w from the dual coefficients

Worked example

sklearn exposes the signed multipliers αᵢyᵢ as dual_coef_ and the support vectors as support_vectors_. Rebuild w = Σ(αᵢyᵢ)xᵢ by hand and confirm it equals coef_, and that Σαᵢyᵢ = 0. Standalone:

import numpy as np
from sklearn.svm import SVC
X = np.array([[0.,0.],[2.,0.],[0.,4.],[2.,4.]])
y = np.array([0, 0, 1, 1])
svm = SVC(kernel='linear', C=1e6).fit(X, y)
dc = svm.dual_coef_[0]              # signed alphas (alpha_i * y_i)
sv = svm.support_vectors_
w_dual = (dc[:, None] * sv).sum(axis=0)
print('support idx :', svm.support_)
print('dual_coef   :', dc.round(4))
print('w (from sum):', w_dual.round(4))
print('sum alpha*y :', round(dc.sum(), 6))

Σ(αᵢyᵢ)xᵢ = [0, 0.5] = coef_, and Σαᵢyᵢ = 0

Why: dual_coef_ = [−0.125, +0.125] on supports (2,0)? no — on the two closest points; (−0.125)·x_neg + (0.125)·x_pos rebuilds [0, 0.5] exactly. The signed alphas sum to 0, satisfying the KKT balance condition.

quantityvalue (verified)
support_ (indices)[1, 3]
dual_coef_ (αᵢyᵢ)[−0.125, +0.125]
w from Σ(αᵢyᵢ)xᵢ[0. , 0.5]
Σ αᵢyᵢ0.0

41. The prediction uses only support vectors

Concept

Substituting w = Σαᵢyᵢxᵢ into the score wᵀx + b turns the decision into a sum over support vectors — the training data enters only through dot products xᵢᵀx:

\[ f(x) = \sum_{i \in \text{SV}} \alpha_i y_i \, (x_i^\top x) + b \]

This dot-product-only form is the door to the kernel trick in Part 4: replace xᵢᵀx with K(xᵢ, x) and you get a non-linear boundary for free.

42. Soft margin: slack and C

Section

Part 3 of 6 — real, noisy data

43. Real data is not separable

Concept

The hard-margin QP demands yᵢ(wᵀxᵢ+b) ≥ 1 for every point. One noisy overlap and no (w, b) satisfies all constraints — the QP is infeasible. Real datasets almost always overlap.

The fix: let points break the margin, but charge for it. We introduce a per-point slack ξᵢ ≥ 0 that measures how far inside (or across) the margin a point sits.

44. Slack variables and the soft-margin QP

Concept

Relax each constraint by its slack, and add the total slack — priced by C — to the objective:

\[ \begin{aligned} &\min_{w,\,b,\,\xi} \;\; \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \\ &\text{s.t.}\;\; y_i(w^\top x_i + b) \ge 1 - \xi_i,\;\; \xi_i \ge 0 \end{aligned} \]

ξᵢ = max(0, 1 − yᵢ(wᵀxᵢ+b)) — the hinge loss

Why: ξᵢ = 0 for a point safely beyond its margin; 0 < ξᵢ < 1 for a point inside the margin but still correct; ξᵢ > 1 for a misclassified point. The sum Σξᵢ is exactly the total hinge loss (Lesson 39).

45. Say it in words: Slack variables and the soft-margin QP

Translation

\( \begin{aligned} &\min_{w,\,b,\,\xi} \;\; \tfrac{1}{2}\lVert w\rVert^2 + C\sum_i \xi_i \\ &\text{s.t.}\;\; y_i(w^\top x_i + b) \ge 1 - \xi_i,\;\; \xi_i \ge 0 \end{aligned} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

46. What C actually controls

Concept

C is the price of a margin violation — the exchange rate between a wide margin (small ‖w‖) and few violations (small Σξᵢ).

Cprioritymarginviolationsbias/variance
small (0.01)wide marginwidemany allowedhigh bias (regularized)
medium (1)balancedmediumsometypical default
large (100)few errorsnarrowfewhigh variance (overfits)

C → ∞ recovers the hard margin (no violation tolerated). C → 0 ignores the data and just shrinks w. The sweet spot is found by cross-validation.

47. Where does each piece belong: Lesson 43: Support Vector Machines

Sorting

Sort into buckets

These are the pieces of Lesson 43: Support Vector Machines, out of order. Put each one back under the part of the lesson it belongs to.

The maximum-margin hyperplane
A hyperplane splits the space; Many lines separate — which is best?; Widest street wins
The dual & support vectors
Attach a multiplier to each constraint; Stationarity gives w = Σ αᵢyᵢxᵢ; KKT: αᵢ > 0 only for support vectors
Soft margin: slack and C
Real data is not separable; Slack variables and the soft-margin QP; What C actually controls
s1
The maximum-margin hyperplane is where Lesson 43: Support Vector Machines puts A hyperplane splits the space, Many lines separate — which is best?, Widest street wins. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The dual & support vectors is where Lesson 43: Support Vector Machines puts Attach a multiplier to each constraint, Stationarity gives w = Σ αᵢyᵢxᵢ, KKT: αᵢ > 0 only for support vectors. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Soft margin: slack and C is where Lesson 43: Support Vector Machines puts Real data is not separable, Slack variables and the soft-margin QP, What C actually controls. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

48. Restore the missing line: C sweep: accuracy and support-vector count

Fill the middle

Fill in the blanks

From C sweep: accuracy and support-vector count — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
acc = accuracy_score(yte, m.predict(Xte))
print(f'C=___ acc=___ n_sv=___')

Why: n_redundant is what everything below it consumes, so the wrong expression here fails later and somewhere else. At C=0.01 the wide margin catches 120 SVs and scores 0.867; by C=1 it is 36 SVs at 0.900; C≥10 stays at 32 SVs and 0.900.

49. C sweep: accuracy and support-vector count

Worked example

Sweep C from 0.01 to 100 on a 200-sample 2-D task (class_sep=0.8, 30% test, standardized). Predict: does the support-vector count go up or down as C grows? Standalone:

import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
    n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for C in [0.01, 0.1, 1.0, 10.0, 100.0]:
    m = SVC(kernel='linear', C=C).fit(Xtr, ytr)
    acc = accuracy_score(yte, m.predict(Xte))
    print(f'C={C:6.2f}  acc={acc:.3f}  n_sv={len(m.support_vectors_)}')

Larger C → fewer support vectors and (up to a point) higher accuracy, then it plateaus

Why: At C=0.01 the wide margin catches 120 SVs and scores 0.867; by C=1 it is 36 SVs at 0.900; C≥10 stays at 32 SVs and 0.900. Beyond C=10, nothing improves — a classic diminishing-returns plateau.

Ctest accn_support_vectors
0.010.867120
0.100.88364
1.000.90036
10.000.90032
100.000.90032

50. Watch it run: C sweep: accuracy and support-vector count

Pattern

Step through it

Step through C sweep: accuracy and support-vector count one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: C is 0.01
  2. Step 2: C is 0.10
  3. Step 3: C is 1.00
  4. Step 4: C is 10.00
  5. Step 5: C is 100.00

51. Guess the shape of the answer: Read the slack directly

Estimation

Predict first

Slack is not hidden — compute ξᵢ = max(0, 1 − yᵢf(xᵢ)) straight from the decision function at C=1, and count how many points violate the margin versus how many are actually misclassified. Standalone:

Commit before you compute: what does Read the slack directly come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: 33 points have ξ>0, but only 12 are actually misclassified

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A point with 0 < ξ < 1 is inside the margin yet on the correct side — it costs slack but is not an error.

52. Read the slack directly

Worked example

Slack is not hidden — compute ξᵢ = max(0, 1 − yᵢf(xᵢ)) straight from the decision function at C=1, and count how many points violate the margin versus how many are actually misclassified. Standalone:

import numpy as np
from sklearn.datasets import make_classification
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=200, n_features=2, n_informative=2,
    n_redundant=0, class_sep=0.8, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
m = SVC(kernel='linear', C=1.0).fit(Xtr, ytr)
ys = np.where(ytr == 1, 1.0, -1.0)
margin = ys * m.decision_function(Xtr)   # y_i (w.x_i + b)
xi = np.maximum(0.0, 1.0 - margin)       # slack
print('n with xi>0    :', int((xi > 1e-9).sum()))
print('n misclassified:', int((margin < 0).sum()))
print('max slack      :', round(xi.max(), 4))

33 points have ξ>0, but only 12 are actually misclassified

Why: A point with 0 < ξ < 1 is inside the margin yet on the correct side — it costs slack but is not an error. So violations (33) always outnumber misclassifications (12). The worst point has slack 2.97 (well across the boundary).

quantityvalue (verified)
train points140
points with ξᵢ > 0 (margin violators)33
points misclassified (ξᵢ > 1)12
max slack ξᵢ2.9721

53. Something is wrong here: larger C always means better performance

Anomaly

Predict first

A student writes this, and it looks reasonable:

Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls.

Treat C as a regularization knob and tune it with cross-validation on held-out folds.

Why: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.

54. Trap: larger C always means better performance

Trap

The trap

Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.

C = 10000 to eliminate violations ✗

Why: That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.

The fix

Treat C as a regularization knob and tune it with cross-validation on held-out folds.

Grid-search C over [0.01, 0.1, 1, 10, 100] with 5-fold CV ✓

Why: CV rewards generalization, not memorization. The chosen C balances margin width against training fit — the identical bias–variance trade-off as L2 weight decay (Lesson 40). Never fix C by staring at training accuracy.

55. Break it on purpose: larger C always means better performance

Break the constraint

Discussion prompt

The rule this trap just fixed:

Treat C as a regularization knob and tune it with cross-validation on held-out folds.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

That memorizes the training set: the margin shrinks toward zero, the boundary bends around noise, and held-out accuracy falls. The C-sweep above already plateaus at C=10 (acc 0.900) — larger C buys nothing and risks overfitting on a harder dataset.

56. The kernel trick & grid search

Section

Part 4 of 6 — non-linear boundaries

57. Everything is a dot product

Concept

Recall the dual decision function f(x) = Σ αᵢyᵢ(xᵢᵀx) + b. The data enters only through dot products xᵢᵀx — never through the raw coordinates on their own.

So if we had a feature map φ lifting x into a richer space, we would only ever need φ(xᵢ)ᵀφ(x) — a single number — not φ(x) itself. That is the opening the kernel trick walks through.

58. The kernel trick

Concept

A kernel K(xᵢ, xⱼ) computes φ(xᵢ)ᵀφ(xⱼ) for some (possibly infinite-dimensional) φ without ever forming φ. Swap every dot product for a kernel and the SVM learns a non-linear boundary at linear cost:

\[ K_{\text{RBF}}(x_i, x_j) = \exp\!\big(-\gamma\,\lVert x_i - x_j\rVert^2\big), \qquad K_{\text{poly}}(x_i, x_j) = (x_i^\top x_j + r)^d \]

RBF (Gaussian): similarity decays with distance; gamma sets how fast (large gamma ⇒ sharp, wiggly boundary). Polynomial: degree d sets the curve order. The RBF φ is literally infinite-dimensional — yet K is one exp.

59. Restore the missing line: An RBF kernel value, by hand

Fill the middle

Fill in the blanks

From An RBF kernel value, by hand — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.metrics.pairwise import rbf_kernel
xi = np.array([[0.0, 0.0]]); xj = np.array([[1.0, 2.0]])
gamma = 0.5
d2 = ((xi - xj) 2).sum()** # ||xi - xj||^2 = 1 + 4 = 5
by_hand = np.exp(-gamma * d2) # exp(-0.5 * 5) = exp(-2.5)
print('||xi-xj||^2 =', d2)
print('by hand =', round(by_hand, 6))
print('sklearn rbf =', round(rbf_kernel(xi, xj, gamma=gamma)[0, 0], 6))

Why: d2 is what everything below it consumes, so the wrong expression here fails later and somewhere else. (0−1)² + (0−2)² = 1 + 4 = 5; times −γ = −2.5; exp(−2.5) = 0.082085.

60. An RBF kernel value, by hand

Worked example

Compute K_RBF(xᵢ, xⱼ) for xᵢ=(0,0), xⱼ=(1,2), γ=0.5 by hand, then confirm against sklearn's rbf_kernel. Standalone:

import numpy as np
from sklearn.metrics.pairwise import rbf_kernel
xi = np.array([[0.0, 0.0]]); xj = np.array([[1.0, 2.0]])
gamma = 0.5
d2 = ((xi - xj) ** 2).sum()      # ||xi - xj||^2 = 1 + 4 = 5
by_hand = np.exp(-gamma * d2)     # exp(-0.5 * 5) = exp(-2.5)
print('||xi-xj||^2 =', d2)
print('by hand     =', round(by_hand, 6))
print('sklearn rbf =', round(rbf_kernel(xi, xj, gamma=gamma)[0, 0], 6))

‖xᵢ−xⱼ‖² = 5 ⇒ exp(−2.5) = 0.082085

Why: (0−1)² + (0−2)² = 1 + 4 = 5; times −γ = −2.5; exp(−2.5) = 0.082085. sklearn returns the identical value — the kernel is just this formula applied to every pair.

quantityvalue (verified)
‖xᵢ − xⱼ‖²5.0
−γ‖xᵢ − xⱼ‖²−2.5
K by hand = exp(−2.5)0.082085
sklearn rbf_kernel0.082085

61. Which is which, by value (verified)

Discrimination

Sort into buckets

Sort these by value (verified), from memory, without looking back at An RBF kernel value, by hand. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

5.0
‖xᵢ − xⱼ‖²
−2.5
−γ‖xᵢ − xⱼ‖²
0.082085
K by hand = exp(−2.5); sklearn rbf_kernel
g1
value (verified) is "5.0" for ‖xᵢ − xⱼ‖² — that is what the table on "An RBF kernel value, by hand" records, and it is the single property separating this group from the rest.
g2
value (verified) is "−2.5" for −γ‖xᵢ − xⱼ‖² — that is what the table on "An RBF kernel value, by hand" records, and it is the single property separating this group from the rest.
g3
value (verified) is "0.082085" for K by hand = exp(−2.5), sklearn rbf_kernel — that is what the table on "An RBF kernel value, by hand" records, and it is the single property separating this group from the rest.

62. RBF beats linear on interleaved moons

Worked example

The moons dataset is two interleaved half-circles — no straight line separates them. Compare a linear SVM to an RBF SVM (both C=1) on 400 samples, noise=0.2, standardized. Standalone:

import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for name, m in [('linear', SVC(kernel='linear', C=1)),
                ('rbf', SVC(kernel='rbf', C=1, gamma='scale'))]:
    m.fit(Xtr, ytr)
    print(name, 'acc =', round(accuracy_score(yte, m.predict(Xte)), 4))

linear 0.825 → RBF 0.950

Why: The linear boundary can't bend around the interleaving, capping it near 0.825. The RBF kernel lifts the points into a space where they ARE separable, jumping to 0.950 — a 12.5-point gain for one kernel change.

classifiermoons test acc
Linear SVM (C=1)0.825
RBF SVM (C=1, gamma='scale')0.950

63. gamma='scale' — a data-adaptive default

Concept

gamma='scale' sets gamma = 1 / (n_features · Var(X)) from the data, so the kernel width adapts to the feature spread. On standardized moons that is 1/(2 · 1.0) = 0.5.

Use it as the starting point, then grid-search around it. Too-large gamma overfits (each point its own island); too-small gamma underfits (one giant blob).

64. Teach it back: gamma='scale' — a data-adaptive default

Explain it

Discussion prompt

Explain gamma='scale' — a data-adaptive default to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

gamma='scale' sets gamma = 1 / (n_features · Var(X)) from the data, so the kernel width adapts to the feature spread. On standardized moons that is 1/(2 · 1.0) = 0.5.

65. Predict the next row: Grid search: tune C and gamma with CV

Pattern

Predict first

The table runs: 0.1 | 10 | 0.9679 · 10 | 1 | 0.9679 · 1 | 1 | 0.9643 · 10 | 10 | 0.9500

In Grid search: tune C and gamma with CV, given the rows so far: what is the next one — the row where C is 0.1?

Correct: 0.1 | 0.1 | 0.8429

Cgamma5-fold CV acc
0.1100.9679
1010.9679
110.9643
10100.9500
0.10.10.8429

Why: The relationship between the columns, not the individual numbers, is what generates the next row. A complex kernel (high gamma) kept in check by a wide, regularized margin (small C) generalizes best here.

66. Grid search: tune C and gamma with CV

Worked example

Search C ∈ {0.1, 1, 10} × gamma ∈ {0.1, 1, 10} on moons with 5-fold CV. GridSearchCV scores every combination on held-out folds and refits the best. Standalone:

import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
    {'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5)
gs.fit(Xtr, ytr)
print('best   :', gs.best_params_)
print('CV acc :', round(gs.best_score_, 4))
print('test   :', round(accuracy_score(yte, gs.predict(Xte)), 4))

Best: C=0.1, gamma=10 → CV 0.9679, test 0.9583

Why: A complex kernel (high gamma) kept in check by a wide, regularized margin (small C) generalizes best here. The grid compared all 9 cells on the folds — the student would never guess this pairing by hand.

Cgamma5-fold CV acc
0.1100.9679
1010.9679
110.9643
10100.9500
0.10.10.8429

67. Watch it run: Grid search: tune C and gamma with CV

Pattern

Step through it

Step through Grid search: tune C and gamma with CV one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: C is 0.1
  2. Step 2: C is 10
  3. Step 3: C is 1
  4. Step 4: C is 10
  5. Step 5: C is 0.1

68. Something is wrong here: grid-searching on the test set

Anomaly

Predict first

A student writes this, and it looks reasonable:

To find the best C and gamma, try all 9 combinations and keep the one with the highest test accuracy — that's the real target metric.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise.

Use GridSearchCV with CV on the training split only; score the winner on the test set exactly once.

Why: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise. The test set must be touched exactly ONCE, after every hyperparameter is frozen.

69. Trap: grid-searching on the test set

Trap

The trap

To find the best C and gamma, try all 9 combinations and keep the one with the highest test accuracy — that's the real target metric.

Evaluate every combo on the test set, pick the winner ✗

Why: This is data leakage. The test set is now part of model selection, so its accuracy is optimistically biased — you tuned to that specific held-out noise. The test set must be touched exactly ONCE, after every hyperparameter is frozen.

The fix

Use GridSearchCV with CV on the training split only; score the winner on the test set exactly once.

gs.fit(Xtr, ytr); test = gs.score(Xte, yte) ✓

Why: Cross-validation resamples inside the training data — no test information leaks in. The single final test evaluation gives an honest generalization estimate. This is the same discipline as every tuning workflow in the course.

70. Which of these survive contact with Lesson 43: Support Vector Machines?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
We carry one dataset through the whole geometry: four points, two per class, chosen so every number is clean and checkable by hand and against sklearn.; The perpendicular distance from a point x₀ to the hyperplane wᵀx + b = 0 is the score at x₀, divided by the length of the normal:; Because scaling (w, b) doesn't move the boundary, we are free to fix the scale. The SVM convention pins it so the closest points score exactly ±1:
Breaks
The support points score ±1, so the margin — the distance to the boundary — is just the score 1. Or: w measures spread, so a bigger w means a wider margin.; Bigger C punishes violations harder, so the model fits the training data better — just crank C to 10000.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 43: Support Vector Machines puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

71. Multi-class & SVM vs LR

Section

Part 5 of 6

72. Multi-class SVM: OvO and OvR

Concept

An SVM is binary. For K classes we combine binary SVMs two ways: One-vs-Rest (OvR) trains K classifiers (each class vs all others); One-vs-One (OvO) trains K(K−1)/2 classifiers (every class pair).

strategy# classifiers (K=3)# classifiers (K=10)decision rule
OvR310highest score wins
OvO345majority vote over pairs

sklearn's SVC always fits OvO internally (each binary problem is smaller, so it's fast for moderate K). decision_function_shape only changes how the scores are reported, not which classifiers are trained.

73. Fill in: decision rule for Multi-class SVM: OvO and OvR

Comparison

Comparison matrix

From Multi-class SVM: OvO and OvR: refill the decision rule column from what you know. The rest of the table is as it appeared.

strategy# classifiers (K=3)# classifiers (K=10)decision rule
OvR310highest score wins
OvO345majority vote over pairs

74. What has to be given first: OvO vs OvR on Iris (K=3)

Missing information

Discussion prompt

Fit an RBF SVM on Iris (150 samples, 4 features, 3 classes) under both decision_function_shape settings. Predict: will accuracy differ? Standalone:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The underlying classifiers are the same OvO set; only the reported decision scores reshape. Predictions are identical. (For K=3 the shapes even coincide: K(K−1)/2 = 3 = K — a numerical accident, but the score VALUES still differ.)

75. OvO vs OvR on Iris (K=3)

Worked example

Fit an RBF SVM on Iris (150 samples, 4 features, 3 classes) under both decision_function_shape settings. Predict: will accuracy differ? Standalone:

import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
    test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
for shape in ['ovo', 'ovr']:
    m = SVC(kernel='rbf', C=1, gamma='scale',
            decision_function_shape=shape).fit(Xtr, ytr)
    print(shape, 'acc =', round(accuracy_score(yte, m.predict(Xte)), 4),
          ' n_sv =', m.n_support_)

Both give acc=0.9778, n_sv=[7,18,17] — identical fit

Why: The underlying classifiers are the same OvO set; only the reported decision scores reshape. Predictions are identical. (For K=3 the shapes even coincide: K(K−1)/2 = 3 = K — a numerical accident, but the score VALUES still differ.)

classn_support_vectors
setosa (0)7
versicolor (1)18
virginica (2)17
total42

76. Watch it run: OvO vs OvR on Iris (K=3)

Pattern

Step through it

Step through OvO vs OvR on Iris (K=3) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: class is setosa (0)
  2. Step 2: class is versicolor (1)
  3. Step 3: class is virginica (2)
  4. Step 4: class is total

77. SVM vs logistic regression

Concept

Both draw linear boundaries in their base form, but they optimize different losses and offer different extras. Pick by what the task needs:

criterionprefer SVMprefer Logistic Regression
non-linearityfree via kernel (RBF/poly)needs hand-built features
probabilitiesnot native (no P(y|x))native, calibrated log-odds
interpretabilityless (dual weights)more (coeff = log-odds)
scale (large n)slow past ~100k rowsfast; supports online SGD
margin / robustnessmaximizes the marginno explicit margin

Rule of thumb: start with LR (fast, interpretable, gives probabilities). Upgrade to an RBF SVM when training accuracy is high but test accuracy lags — the signature of a needed non-linear boundary.

78. By analogy: SVM vs logistic regression

Analogy

Discussion prompt

Explain SVM vs logistic regression by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Both draw linear boundaries in their base form, but they optimize different losses and offer different extras. Pick by what the task needs:

79. Guess the shape of the answer: SVM vs LR on moons — the non-linear gap

Estimation

Predict first

On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:

Commit before you compute: what does SVM vs LR on moons — the non-linear gap come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: LogReg 0.8083 vs RBF SVM 0.9500

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Logistic regression is stuck with a straight boundary (0.8083).

80. SVM vs LR on moons — the non-linear gap

Worked example

On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:

import numpy as np
from sklearn.datasets import make_moons
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
X, y = make_moons(n_samples=400, noise=0.2, random_state=42)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
lr = LogisticRegression().fit(Xtr, ytr)
sv = SVC(kernel='rbf', C=1, gamma='scale').fit(Xtr, ytr)
print('LogReg  acc =', round(accuracy_score(yte, lr.predict(Xte)), 4))
print('RBF SVM acc =', round(accuracy_score(yte, sv.predict(Xte)), 4))

LogReg 0.8083 vs RBF SVM 0.9500

Why: Logistic regression is stuck with a straight boundary (0.8083). The RBF SVM curves around the interleaving (0.9500) — a 14-point gap that is entirely the kernel. On linearly separable data the two would tie.

classifiermoons test accboundary
Logistic Regression0.8083linear
RBF SVM0.9500non-linear (kernel)

81. Work backwards from the answer: SVM vs LR on moons — the non-linear gap

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

LogReg 0.8083 vs RBF SVM 0.9500

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

On the same standardized moons split, pit plain logistic regression against an RBF SVM. Predict which wins and by how much. Standalone:

82. Rebuild the recipe: The SVM workflow recipe

Ranking

Put in order

These are the steps of The SVM workflow recipe, scrambled. Put them back in order before the next slide shows you.

  1. Scale first: StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)
  2. Start linear: SVC(kernel='linear', C=1) to test whether the data is linearly separable
  3. Kernelize: SVC(kernel='rbf', C=1, gamma='scale') for non-linear boundaries
  4. Grid-search: tune C and gamma with GridSearchCV(cv=5) on training data only
  5. Multi-class: SVC does OvO automatically; LinearSVC does OvR by default
  6. Evaluate once: report test accuracy only after every hyperparameter is frozen

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

83. The SVM workflow recipe

Pattern

  1. Scale first: StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)
  2. Start linear: SVC(kernel='linear', C=1) to test whether the data is linearly separable
  3. Kernelize: SVC(kernel='rbf', C=1, gamma='scale') for non-linear boundaries
  4. Grid-search: tune C and gamma with GridSearchCV(cv=5) on training data only
  5. Multi-class: SVC does OvO automatically; LinearSVC does OvR by default
  6. Evaluate once: report test accuracy only after every hyperparameter is frozen

84. Where does it stop working: The SVM workflow recipe

Edge cases

Discussion prompt

The SVM workflow recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Scale first: StandardScaler — SVM is not scale-invariant (both C and gamma depend on feature scale)
  2. Start linear: SVC(kernel='linear', C=1) to test whether the data is linearly separable
  3. Kernelize: SVC(kernel='rbf', C=1, gamma='scale') for non-linear boundaries
  4. Grid-search: tune C and gamma with GridSearchCV(cv=5) on training data only
  5. Multi-class: SVC does OvO automatically; LinearSVC does OvR by default
  6. Evaluate once: report test accuracy only after every hyperparameter is frozen

85. Rule out three: Check yourself — the margin formula

Elimination

Eliminate the wrong options

An SVM is scaled so the closest points score ±1 and the fitted weights have ‖w‖ = 0.25. What is the margin (full street width)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 8
  • B. 0.25
  • C. 4
  • D. 0.5

Survives elimination: A

Why: The margin is 2/‖w‖ = 2/0.25 = 8. Each support point sits at signed distance 1/‖w‖ = 4 from the boundary; the two half-widths add to 8.

86. Check yourself — the margin formula

Check

Reason from the derivation, not from memory.

Check your understanding

An SVM is scaled so the closest points score ±1 and the fitted weights have ‖w‖ = 0.25. What is the margin (full street width)?

  • A. 8 (correct)
  • B. 0.25
  • C. 4
  • D. 0.5

Answer: A

Why: The margin is 2/‖w‖ = 2/0.25 = 8. Each support point sits at signed distance 1/‖w‖ = 4 from the boundary; the two half-widths add to 8.

Why B tempts people
That is ‖w‖ itself, not the margin. ‖w‖ is inversely related to the margin — a small ‖w‖ means a WIDE margin.
Why C tempts people
That is the half-margin 1/‖w‖ = 1/0.25 = 4 (the distance to one margin plane), not the full street width, which doubles it to 8.
Why D tempts people
That looks like 2·‖w‖ = 0.5. The margin divides by ‖w‖, it does not multiply — larger ‖w‖ gives a narrower margin.

87. Answer it before you see the options: Check yourself — support vectors

Prediction

Predict first

You remove a training point that is NOT a support vector, then refit a hard-margin SVM. What happens to the boundary?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Nothing — non-support-vectors have αᵢ = 0 and don't affect w or b

Why: By complementary slackness, a non-support-vector satisfies yᵢ(wᵀxᵢ+b) > 1, so its αᵢ = 0. It contributes nothing to w = Σαᵢyᵢxᵢ, so removing it leaves w, b, and the margin unchanged.

88. Check yourself — support vectors

Check

Think about the KKT conditions and w = Σαᵢyᵢxᵢ.

Check your understanding

You remove a training point that is NOT a support vector, then refit a hard-margin SVM. What happens to the boundary?

  • A. Nothing — non-support-vectors have αᵢ = 0 and don't affect w or b (correct)
  • B. The margin widens because there's one fewer constraint
  • C. The boundary shifts toward the removed point's class
  • D. The model becomes non-linear

Answer: A

Why: By complementary slackness, a non-support-vector satisfies yᵢ(wᵀxᵢ+b) > 1, so its αᵢ = 0. It contributes nothing to w = Σαᵢyᵢxᵢ, so removing it leaves w, b, and the margin unchanged.

Why B tempts people
The margin is 2/‖w‖ and depends only on the support vectors. Dropping a point that already had αᵢ = 0 does not change ‖w‖.
Why C tempts people
A non-SV exerts zero influence via the dual (αᵢ = 0), so nothing pulls the boundary in any direction. Removing a SUPPORT vector could move it — but this point isn't one.
Why D tempts people
Removing a data point never changes the kernel; a linear SVM stays linear regardless of which points are present.

89. Answer it before you see the options: Check yourself — C and gamma

Prediction

Predict first

An RBF SVM overfits: train acc 0.99, test acc 0.75. Which change is most likely to help?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Decrease C and/or decrease gamma

Why: Overfitting means the model is too complex. Lowering C widens the margin (more regularization); lowering gamma broadens each RBF bump (smoother boundary). Both raise bias and cut variance — exactly the correction needed.

90. Check yourself — C and gamma

Check

Reason about bias and variance.

Check your understanding

An RBF SVM overfits: train acc 0.99, test acc 0.75. Which change is most likely to help?

  • A. Decrease C and/or decrease gamma (correct)
  • B. Increase C and increase gamma
  • C. Switch decision_function_shape from ovo to ovr
  • D. Remove the StandardScaler

Answer: A

Why: Overfitting means the model is too complex. Lowering C widens the margin (more regularization); lowering gamma broadens each RBF bump (smoother boundary). Both raise bias and cut variance — exactly the correction needed.

Why B tempts people
Higher C narrows the margin (less regularization) and higher gamma sharpens the kernel (more wiggle) — both make overfitting worse.
Why C tempts people
decision_function_shape only reshapes reported scores; it changes neither the fit nor the complexity, so it can't fix overfitting.
Why D tempts people
Dropping the scaler makes C and gamma scale-dependent and usually hurts — it does not address model complexity.

91. Rule out three: Check yourself — the kernel trick

Elimination

Eliminate the wrong options

The kernel trick works because the SVM's dual decision function depends on the data only through:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Pairwise dot products xᵢᵀx, which can be swapped for K(xᵢ, x)
  • B. The mean and variance of each feature
  • C. Gradient descent, which computes kernels via autograd
  • D. The number of support vectors, which equals the feature-space dimension

Survives elimination: A

Why: Substituting w = Σαᵢyᵢxᵢ gives f(x) = Σαᵢyᵢ(xᵢᵀx) + b — the inputs appear only inside dot products. Replacing each xᵢᵀx with K(xᵢ, x) = φ(xᵢ)ᵀφ(x) uses a high-dimensional map φ without ever computing it.

92. Check yourself — the kernel trick

Check

Why is the trick even possible?

Check your understanding

The kernel trick works because the SVM's dual decision function depends on the data only through:

  • A. Pairwise dot products xᵢᵀx, which can be swapped for K(xᵢ, x) (correct)
  • B. The mean and variance of each feature
  • C. Gradient descent, which computes kernels via autograd
  • D. The number of support vectors, which equals the feature-space dimension

Answer: A

Why: Substituting w = Σαᵢyᵢxᵢ gives f(x) = Σαᵢyᵢ(xᵢᵀx) + b — the inputs appear only inside dot products. Replacing each xᵢᵀx with K(xᵢ, x) = φ(xᵢ)ᵀφ(x) uses a high-dimensional map φ without ever computing it.

Why B tempts people
Kernels act on raw pairs (xᵢ, x), not on feature statistics. Mean/variance belong to preprocessing (StandardScaler), not the dual.
Why C tempts people
The SVM is solved as a QP, not by gradient descent. The trick is algebraic — it's about dot products in the dual, independent of the solver.
Why D tempts people
The support-vector count is a property of the solution; the RBF feature space is infinite-dimensional regardless of how many SVs there are.

93. Your turn: build it

Section

Part 6 of 6 — the project

94. Project: an SVM pipeline on Iris

Concept

Build the full workflow on Iris: scale → linear baseline → RBF with grid search → compare to logistic regression → report test accuracy once. You've derived every piece — now wire it together.

#milestonetool
1scale + linear SVM baselineStandardScaler + SVC(kernel='linear')
2RBF SVM with grid searchGridSearchCV, cv=5, C+gamma
3compare to LogisticRegressionsklearn.linear_model.LogisticRegression

Build rules: scale before fitting; call GridSearchCV.fit on the training split only; touch the test set exactly once at the very end.

95. Break it if you can: Project: an SVM pipeline on Iris

Counterexample

Discussion prompt

Build the full workflow on Iris: scale → linear baseline → RBF with grid search → compare to logistic regression → report test accuracy once. You've derived every piece — now wire it together.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: scale before fitting; call GridSearchCV.fit on the training split only; touch the test set exactly once at the very end.

96. Milestone 1 — linear baseline on Iris

Worked example

Your turn: scale Iris, fit a linear SVM at C=1, and record test accuracy and support-vector count per class. Predict the accuracy before you print.

Hint: split with random_state=0, test_size=0.3; StandardScaler().fit_transform(Xtr) then .transform(Xte) — never fit the scaler on test data.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
    test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
svm = SVC(kernel='linear', C=1).fit(Xtr, ytr)
print('acc :', round(accuracy_score(yte, svm.predict(Xte)), 4))
print('n_sv:', svm.n_support_)
metricvalue (verified)
test accuracy0.9778
n_sv setosa (0)2
n_sv versicolor (1)12
n_sv virginica (2)10

97. Milestone 2 — RBF SVM with grid search

Worked example

Your turn: grid-search C ∈ {0.1, 1, 10} × gamma ∈ {0.1, 1, 10} with 5-fold CV on the training split. Predict whether RBF will beat the linear baseline on Iris.

Hint: GridSearchCV(SVC(kernel='rbf'), {...}, cv=5).fit(Xtr, ytr), then read gs.best_params_ and gs.score(Xte, yte).

import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
    test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
    {'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
print('best params:', gs.best_params_)
print('CV  acc    :', round(gs.best_score_, 4))
print('test acc   :', round(gs.score(Xte, yte), 4))
settingCV acctest acc
linear SVM, C=1 (baseline)—0.9778
RBF, best {C:10, gamma:0.1}0.95240.9778

98. Milestone 3 — compare to logistic regression

Worked example

Your turn: fit logistic regression on the same scaled split and compare to the grid-searched SVM. Predict which wins on Iris and why.

Hint: reuse the exact Xtr, Xte from Milestone 2; LogisticRegression(C=1, max_iter=200, random_state=0). Iris is nearly linearly separable, so think about whether the kernel can even help.

import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
    test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
    {'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
lr = LogisticRegression(C=1, max_iter=200, random_state=0).fit(Xtr, ytr)
print('SVM test acc:', round(gs.score(Xte, yte), 4))
print('LR  test acc:', round(accuracy_score(yte, lr.predict(Xte)), 4))
classifiertest accprobabilities?
Logistic Regression0.9778yes (native)
RBF SVM (grid-searched)0.9778no (by default)

99. What each one costs: Milestone 3 — compare to logistic regression

Trade off

Comparison matrix

From Milestone 3 — compare to logistic regression: every row here is a choice with a cost. Fill the probabilities? column, then say which row you would actually pick and what you give up for it.

classifiertest accprobabilities?
Logistic Regression0.9778yes (native)
RBF SVM (grid-searched)0.9778no (by default)

100. The full program

Concept

import numpy as np
from sklearn.datasets import load_iris
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import accuracy_score
iris = load_iris()
Xtr, Xte, ytr, yte = train_test_split(iris.data, iris.target,
    test_size=0.3, random_state=0)
sc = StandardScaler(); Xtr = sc.fit_transform(Xtr); Xte = sc.transform(Xte)
gs = GridSearchCV(SVC(kernel='rbf'),
    {'C': [0.1, 1, 10], 'gamma': [0.1, 1, 10]}, cv=5).fit(Xtr, ytr)
lr = LogisticRegression(C=1, max_iter=200, random_state=0).fit(Xtr, ytr)
print('SVM best:', gs.best_params_, 'test acc:', round(gs.score(Xte, yte), 4))
print('LR  test acc:', round(accuracy_score(yte, lr.predict(Xte)), 4))
printed linevalue (verified)
SVM best params{'C': 10, 'gamma': 0.1}
SVM test acc0.9778
LR test acc0.9778
verdicttie on Iris — LR preferred for native probabilities

On Iris both hit 97.8% — the data is nearly linear, so the kernel earns nothing. The SVM's edge shows on moons (SVM 0.950 vs LR 0.808). Reach for LR when you need calibrated probabilities; reach for the RBF SVM when the boundary must bend.

101. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
SVM best params{'C': 10, 'gamma': 0.1}
SVM test acc0.9778
LR test acc0.9778
verdicttie on Iris — LR preferred for native probabilities

102. Show it off

Concept

Slides closed, out loud: explain (1) how margin = 2/‖w‖ falls out of the signed-distance formula and the ±1 scaling, (2) why the KKT conditions force αᵢ = 0 for every non-support-vector, and (3) how replacing xᵢᵀx with K(xᵢ, x) gives a non-linear boundary without ever computing φ(x).

Stretch (homework): plot the moons decision boundary for gamma ∈ {0.1, 1, 10} and watch it go from underfit to overfit; derive the full SVM dual objective Σαᵢ − ½Σᵢⱼαᵢαⱼyᵢyⱼ(xᵢᵀxⱼ); and confirm that dropping a support vector does move the boundary. Next up: decision trees and random forests.

103. Connect it up: Lesson 43: Support Vector Machines

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The maximum-margin hyperplane · The dual & support vectors · Soft margin: slack and C · The kernel trick & grid search · Multi-class & SVM vs LR · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

104. What you can do now

Recap

ideathe one thing to remember
margin2/‖w‖ — divide the ±1 score by ‖w‖; minimize ½‖w‖²
support vectorsonly αᵢ > 0 points define w; KKT zeroes the rest
soft margin Clarge C → narrow, overfits; small C → wide, regularized
kernel trickswap xᵢᵀx for K(xᵢ,x) — never compute φ(x)
workflowscale, grid-search on train CV, touch test once

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 43 (Week 15 — Support Vector Machines) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn SVC / SVM user guide (dual formulation, decision_function_shape)
  3. Cristianini & Shawe-Taylor, An Introduction to Support Vector Machines, Ch. 6 (margin, dual, KKT) — Cambridge University Press, 2000
  4. Every weight, margin, kernel value, accuracy, and support-vector count produced by real execution — scikit-learn 1.9.0 + numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108