USAAIO Lesson 24, from Week 8, fully worked, building constrained optimization from the ground up. One running toy problem - minimizing x² + y² on a line - carries the derivation of the Lagrangian gradient by gradient, and the multiplier is revealed as a shadow price. Inequality constraints then give the four KKT conditions, with complementary slackness proved on actual numbers. A four-point one-dimensional SVM is solved both as a primal problem and as a dual QP, giving α = [0, 0.5, 0.5, 0] with support vectors at x = 2 and x = 4, and the lesson closes strong duality and Slater's condition by matching the dual optimum to the primal. Every scipy and sklearn snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.
Subject: Machine Learning · 111 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 24 · Week 8 (Optimization)
Minimizing with rules you must obey. We derive Lagrange multipliers gradient by gradient, read the multiplier as a price, extend to inequalities with the four KKT conditions, and finish by solving a real SVM as both a primal and a dual — skipping no step, with every number from live execution.
Objectives
min f s.t. h=0, g≤0 and say why the unconstrained minimum usually fails∇f = −λ∇h from 'no improving feasible direction' — no step skippedλ as a shadow price: ∂f⋆/∂c = −λ, and verify it numericallyμᵢ gᵢ = 0α ≥ 0, and see why only support vectors survive0Warm-up
Discussion prompt
Before we open Lesson 24: Constrained Optimization & KKT: without looking back, what was the main idea of The Multivariate Gaussian, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the multivariate Gaussian N(μ,Σ), why Σ must be PSD, Mahalanobis distance, Gaussian marginals/conditionals, the diagonal-Σ independence fact, and Cholesky sampling x=μ+Lz — with the VAE latent-space connection. Sample a 2D Gaussian via Cholesky and recover its covariance.
Section
Part 1 of 6
Concept
Unconstrained optimization asks 'where is f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.
\[ \min_{x}\; f(x) \quad \text{s.t.} \quad h(x) = 0, \;\; g(x) \le 0 \]
feasible set — The set of all x that satisfy every constraint. Optimization now happens only inside this set — the true minimum of f may lie outside it, and then it is simply not allowed.
Counterexample
Discussion prompt
Unconstrained optimization asks 'where is f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:
\[ \min_{x,y}\; f(x,y) = x^2 + y^2 \quad \text{s.t.} \quad h(x,y) = x + y - 1 = 0 \]
f alone is minimized at the origin (0,0). But (0,0) is off the line x+y=1, so it is forbidden. We need the closest feasible point instead.
Analogy
Discussion prompt
Explain Our running example by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:
Intuition
Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.
The constraint is a straight line. We want the smallest circle that still touches the line. Too small and it misses the line entirely; too big and we overpaid.
The winning circle just barely kisses the line — touches it at exactly one point. That tangency is the whole secret, and the next slides turn 'tangent' into an equation.
Explain it
Discussion prompt
Explain The unconstrained answer is usually illegal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.
Picture it
Figure (svg): Concentric circles centered at the origin, a slanted line x+y=1, and the smallest circle touching the line at one point where the gradient of f and the gradient of h are drawn as parallel arrows.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
At the kissing point the circle and the line are tangent — they share the same tangent direction. So their perpendiculars (their gradients) point the same way.
Concept
At the kissing point the circle and the line are tangent — they share the same tangent direction. So their perpendiculars (their gradients) point the same way.
Figure (svg): Concentric circles centered at the origin, a slanted line x+y=1, and the smallest circle touching the line at one point where the gradient of f and the gradient of h are drawn as parallel arrows.
\[ \nabla f = -\lambda\, \nabla h \quad\text{for some scalar } \lambda \]
Intuition
Stand on the constraint line. You may only step along it. Moving lowers f only if ∇f has a component along your step direction.
If ∇f points straight across the line (perpendicular to it), then every legal step is sideways to ∇f — no direction along the line decreases f. You are stuck at the optimum.
∇h is exactly the direction perpendicular to the line. So 'stuck' means ∇f points along ∇h: ∇f = −λ∇h. That scalar λ is the Lagrange multiplier.
Concept
It helps to see the two regimes before the algebra. An equality h=0 traps you on a surface — you can never leave it, so the optimum is almost always where f's level set is tangent to that surface.
An inequality g≤0 gives you a whole region. If the free minimum is already inside, the constraint does nothing; if it's outside, you get pushed to the boundary and the inequality behaves like an equality there.
| constraint | optimum sits | multiplier sign |
|---|---|---|
| equality h=0 | on the surface (always) | λ any sign |
| inequality g≤0, active | on the boundary g=0 | μ ≥ 0 |
| inequality g≤0, inactive | in the interior g<0 | μ = 0 |
Comparison
Comparison matrix
From Two ways a constraint can bite: refill the optimum sits column from what you know. The rest of the table is as it appeared.
| constraint | optimum sits | multiplier sign |
|---|---|---|
| equality h=0 | on the surface (always) | λ any sign |
| inequality g≤0, active | on the boundary g=0 | μ ≥ 0 |
| inequality g≤0, inactive | in the interior g<0 | μ = 0 |
Section
Part 2 of 6 — derive it, then read λ
Concept
Rather than juggle ∇f = −λ∇h and h=0 separately, fold them into a single scalar function — the Lagrangian — whose stationary point reproduces both.
\[ \mathcal{L}(x, \lambda) = f(x) + \lambda\, h(x) \]
λ is a brand-new variable we optimize over too. Setting all its partial derivatives to zero recovers both the tangency condition and the constraint, as the next two slides show.
Ranking
Put in order
Put the moves of Stationarity in x recovers tangency into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. ∂/∂x of f(x)+λ h(x) is ∇f + λ∇h.
Worked example
Differentiate L with respect to the x-variables
Why: ∂/∂x of f(x)+λ h(x) is ∇f + λ∇h. λ is treated as a constant here.
\[ \nabla_x \mathcal{L} = \nabla f(x) + \lambda\, \nabla h(x) \]
Set it to zero
Why: At a stationary point the gradient vanishes.
\[ \nabla f(x) + \lambda\, \nabla h(x) = 0 \;\Longleftrightarrow\; \nabla f = -\lambda\, \nabla h \]
This is exactly the tangency condition
Why: The gradients-parallel picture from Part 1 falls out with no extra work — the Lagrangian's x-stationarity IS 'aligned gradients'.
Notation
Annotate
From Stationarity in x recovers tangency — read this one piece at a time. What is each part doing?
On: \( \nabla_x \mathcal{L} = \nabla f(x) + \lambda\, \nabla h(x) \)
Step zero
Discussion prompt
Stationarity in λ recovers the constraint — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Differentiate L with respect to λ
Answer:
Worked example
Differentiate L with respect to λ
Why: ∂/∂λ of f(x)+λ h(x) is just h(x) — f has no λ, and λ h(x) differentiates to h(x).
\[ \frac{\partial \mathcal{L}}{\partial \lambda} = h(x) \]
Set it to zero
Why: Stationarity in the λ direction.
\[ \frac{\partial \mathcal{L}}{\partial \lambda} = 0 \;\Longrightarrow\; h(x) = 0 \]
This is exactly the constraint
Why: So a single condition ∇L = 0 (over both x and λ) bundles BOTH tangency and feasibility. That is why the Lagrangian is worth building.
Translation
\( \frac{\partial \mathcal{L}}{\partial \lambda} = 0 \;\Longrightarrow\; h(x) = 0 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Ranking
Put in order
Put the moves of Solve our running example by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Both give the same relation to λ, so 2x = 2y ⇒ x = y.
Worked example
For f = x²+y² and h = x+y−1, write out L = x²+y²+λ(x+y−1) and take all three partials:
∂L/∂x = 2x + λ = 0 and ∂L/∂y = 2y + λ = 0
Why: Both give the same relation to λ, so 2x = 2y ⇒ x = y. The problem's symmetry appears automatically.
\[ 2x + \lambda = 0,\qquad 2y + \lambda = 0 \;\Longrightarrow\; x = y \]
∂L/∂λ = x + y − 1 = 0
Why: The constraint. Substitute x = y: 2x = 1, so x = y = 0.5.
\[ x + y = 1,\; x = y \;\Longrightarrow\; x = y = \tfrac12 \]
Back out λ and f⋆
Why: λ = −2x = −1; f⋆ = 0.5² + 0.5² = 0.5. Answer: (0.5, 0.5), f⋆ = 0.5, λ = −1.
\[ \boxed{\,(x^\star, y^\star) = (\tfrac12, \tfrac12), \quad f^\star = \tfrac12, \quad \lambda = -1\,} \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Back out λ and f⋆
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
For f = x²+y² and h = x+y−1, write out L = x²+y²+λ(x+y−1) and take all three partials:
Estimation
Predict first
Hand it the same problem numerically. scipy.optimize.minimize takes each constraint as a dict; 'eq' means fun(v)=0. Runnable as-is:
Commit before you compute: what does Verify the hand answer against scipy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: scipy lands on the same point
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Prints [0.5 0.5] 0.5 — identical to the by-hand (0.5, 0.5) with f⋆ = 0.5.
Worked example
Hand it the same problem numerically. scipy.optimize.minimize takes each constraint as a dict; 'eq' means fun(v)=0. Runnable as-is:
from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
constraints={'type': 'eq', 'fun': lambda v: v[0] + v[1] - 1})
print(res.x.round(4), round(res.fun, 4))scipy lands on the same point
Why: Prints [0.5 0.5] 0.5 — identical to the by-hand (0.5, 0.5) with f⋆ = 0.5. The Lagrangian answer is confirmed by an independent solver.
| quantity | by hand | scipy (verified) |
|---|---|---|
| x⋆, y⋆ | [0.5, 0.5] | [0.5, 0.5] |
| f⋆ | 0.5 | 0.5 |
| λ | −1 | — (implicit) |
Trade off
Comparison matrix
From Verify the hand answer against scipy: every row here is a choice with a cost. Fill the by hand column, then say which row you would actually pick and what you give up for it.
| quantity | by hand | scipy (verified) |
|---|---|---|
| x⋆, y⋆ | [0.5, 0.5] | [0.5, 0.5] |
| f⋆ | 0.5 | 0.5 |
| λ | −1 | — (implicit) |
Anomaly
Predict first
A student writes this, and it looks reasonable:
It's still x²+y², so minimize f directly and report the origin.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e.
Fold the constraint in with a multiplier and solve ∇L = 0 and h = 0 together.
Why: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e. infeasible. The unconstrained minimum is almost always outside the feasible set, which is the entire reason constrained optimization exists.
Trap
It's still x²+y², so minimize f directly and report the origin.
Answer (0, 0), f = 0
Why: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e. infeasible. The unconstrained minimum is almost always outside the feasible set, which is the entire reason constrained optimization exists.
Fold the constraint in with a multiplier and solve ∇L = 0 and h = 0 together.
Answer (0.5, 0.5), f = 0.5
Why: The multiplier λ = −1 balances the pull toward the origin against staying on x+y=1. The feasible optimum costs more than the illegal one — and that gap, exactly, is what λ prices.
Concept
The multiplier has a meaning: it is the shadow price of the constraint — how fast the optimal value changes when you loosen the constraint by one unit.
Generalize the constraint to x + y = c. Then the optimal cost f⋆(c) depends on c, and the envelope theorem says its derivative is minus the multiplier:
\[ \frac{d f^\star}{d c} = -\lambda \]
For our line, f⋆(c) = c²/2, so df⋆/dc = c = 1 at c=1, matching −λ = −(−1) = 1. Let's confirm that on real numbers.
Missing information
Discussion prompt
Solve the problem at c = 1.00 and c = 1.01, then estimate df⋆/dc by finite difference. It should come out near 1 = −λ:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Loosening the constraint to x+y=1.01 raised the optimum from 0.50000 to 0.51005; the rate 1.005 (a finite-difference estimate of the true 1.0) is exactly minus the multiplier. λ literally prices the constraint.
Worked example
Solve the problem at c = 1.00 and c = 1.01, then estimate df⋆/dc by finite difference. It should come out near 1 = −λ:
from scipy.optimize import minimize
def solve(c):
return minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
constraints={'type': 'eq',
'fun': lambda v: v[0] + v[1] - c}).fun
f1, f2 = solve(1.00), solve(1.01)
print('f*(1.00) =', round(f1, 5))
print('f*(1.01) =', round(f2, 5))
print('df*/dc ~', round((f2 - f1) / 0.01, 4), ' (= -lambda = 1)')The slope of f⋆ in c is ~1 = −λ
Why: Loosening the constraint to x+y=1.01 raised the optimum from 0.50000 to 0.51005; the rate 1.005 (a finite-difference estimate of the true 1.0) is exactly minus the multiplier. λ literally prices the constraint.
| c | f⋆(c) (verified) | meaning |
|---|---|---|
| 1.00 | 0.50000 | our original problem |
| 1.01 | 0.51005 | constraint loosened by 0.01 |
| Δf⋆/Δc | 1.005 ≈ 1 | = −λ, the shadow price |
Concept
Real problems mix equalities and inequalities. Give each equality hⱼ a free multiplier λⱼ and each inequality gᵢ a nonnegative multiplier μᵢ, and add them all in:
\[ \mathcal{L}(x, \lambda, \mu) = f(x) + \sum_j \lambda_j\, h_j(x) + \sum_i \mu_i\, g_i(x), \qquad \mu_i \ge 0 \]
Stationarity ∇ₓL = 0 now reads ∇f + Σλⱼ∇hⱼ + Σμᵢ∇gᵢ = 0. The only new subtlety versus Part 2 is the sign restriction on the μᵢ — which the next slides earn geometrically.
Section
Part 3 of 6 — four conditions
Concept
An inequality g(x) ≤ 0 carves out a whole region, not a curve. The optimum can sit inside the region (constraint irrelevant) or on its boundary (constraint binding).
active vs inactive — A constraint is ACTIVE at x⋆ if g(x⋆)=0 (the solution sits on the boundary, the constraint is pushing back). It is INACTIVE if g(x⋆)<0 (satisfied with room to spare, exerting no force).
We don't know in advance which case we're in — so we need conditions that cover both. Those are the KKT conditions.
Definition probe
Sort into buckets
Every line below is part of the definition of feasible set or of active vs inactive — one or the other, never both. Put each where it belongs.
Concept
For min f s.t. gᵢ(x) ≤ 0 (with multipliers μᵢ), a KKT point satisfies all four of these at once:
∇f + Σᵢ μᵢ ∇gᵢ = 0gᵢ(x) ≤ 0 (obey the rules)μᵢ ≥ 0 (inequality multipliers are one-signed)μᵢ gᵢ(x) = 0For convex problems these are not just necessary but sufficient: any point meeting all four is a global optimum. That is what makes them the workhorse of ML optimization.
Intuition
For an equality, λ could be any sign — you can be pushed either way along the constraint. For an inequality g ≤ 0, you can only be pushed inward, off the boundary.
So the constraint's force μ∇g can only point one way, which pins μ ≥ 0 (dual feasibility). A negative μ would mean the constraint is pulling the solution the wrong direction — impossible for a real barrier.
Sorting
Sort into buckets
These are the pieces of Lesson 24: Constrained Optimization & KKT, out of order. Put each one back under the part of the lesson it belongs to.
Intuition
μᵢ gᵢ = 0 says: for each constraint, at least one of the pair is zero. Either the constraint is active (gᵢ = 0) and may push (μᵢ > 0), or it is inactive (gᵢ < 0) and cannot push (μᵢ = 0).
Never both slack and pushing. A constraint you satisfy with room to spare exerts no force on the answer — remove it and nothing changes. This single fact is what will make the SVM sparse.
Pattern
Predict first
The table runs: x⋆, y⋆ | [1, 0] | on the boundary · g = 1 − x at x⋆ | 0 | ACTIVE
In An ACTIVE inequality, solved and checked, given the rows so far: what is the next one — the row where quantity is μ?
Correct: μ | 2 > 0 | dual-feasible, may push
| quantity | value (verified) | KKT reading |
|---|---|---|
| x⋆, y⋆ | [1, 0] | on the boundary |
| g = 1 − x at x⋆ | 0 | ACTIVE |
| μ | 2 > 0 | dual-feasible, may push |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Prints [1. -0.] 1.0 and active: True.
Worked example
Minimize f = x²+y² subject to x ≥ 1, i.e. g = 1−x ≤ 0. The unconstrained min (0,0) violates it, so the boundary must bind. Solve numerically — scipy's 'ineq' means fun(v) ≥ 0, so pass x−1:
from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
constraints={'type': 'ineq', 'fun': lambda v: v[0] - 1})
print(res.x.round(4), round(res.fun, 4))
print('active:', round(res.x[0] - 1, 4) == 0.0)Solution sits ON the boundary x = 1
Why: Prints [1. -0.] 1.0 and active: True. The constraint is tight (g = 0), so complementary slackness permits μ > 0 — by hand, stationarity 2x = μ gives μ = 2.
| quantity | value (verified) | KKT reading |
|---|---|---|
| x⋆, y⋆ | [1, 0] | on the boundary |
| g = 1 − x at x⋆ | 0 | ACTIVE |
| μ | 2 > 0 | dual-feasible, may push |
Worked example
Now relax the rule to x ≥ −1, i.e. g = −1−x ≤ 0. The unconstrained min (0,0) already obeys it with slack, so the constraint should exert no force and get μ = 0:
from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
constraints={'type': 'ineq', 'fun': lambda v: v[0] + 1})
print(res.x.round(4), round(res.fun, 4))
print('slack g = x+1 =', round(res.x[0] + 1, 4), '-> mu = 0')Solution ignores the boundary and sits at the origin
Why: Prints [-0. -0.] 0.0. Here g = x+1 = 1 > 0 (strictly satisfied), so complementary slackness μ·g = 0 forces μ = 0. Same problem shape, opposite KKT case.
| quantity | value (verified) | KKT reading |
|---|---|---|
| x⋆, y⋆ | [0, 0] | interior of feasible set |
| g = x + 1 at x⋆ | 1 > 0 | INACTIVE (slack) |
| μ | 0 | no force on the solution |
Fill the middle
Fill in the blanks
From All four KKT conditions on numbers — one line has had its right-hand side removed. Put it back.
import numpy as np
xs, ys, mu = 1.0, 0.0, 2.0 # solution + multiplier
grad_f = np.array([2xs, 2ys]) # grad of x^2+y^2
grad_g = np.array([-1.0, 0.0]) # g(x)=1-x, so grad_g=(-1,0)
print('stationarity :', grad_f + mu*grad_g) # want [0,0]
print('primal feas :', 1 - xs, '<= 0') # g <= 0
print('dual feas :', mu, '>= 0') # mu >= 0
print('slackness :', mu * (1 - xs), '= 0') # mu*g = 0
Why: grad_g is what everything below it consumes, so the wrong expression here fails later and somewhere else. Stationarity (2,0)+2(−1,0) = [0,0]; primal 1−1 = 0 ≤ 0; dual 2 ≥ 0; slackness 2·0 = 0.
Worked example
Take the active case (x⋆, y⋆) = (1, 0) with g = 1−x and μ = 2, and check each KKT condition explicitly. ∇f = (2x, 2y), ∇g = (−1, 0):
import numpy as np
xs, ys, mu = 1.0, 0.0, 2.0 # solution + multiplier
grad_f = np.array([2*xs, 2*ys]) # grad of x^2+y^2
grad_g = np.array([-1.0, 0.0]) # g(x)=1-x, so grad_g=(-1,0)
print('stationarity :', grad_f + mu*grad_g) # want [0,0]
print('primal feas :', 1 - xs, '<= 0') # g <= 0
print('dual feas :', mu, '>= 0') # mu >= 0
print('slackness :', mu * (1 - xs), '= 0') # mu*g = 0Every condition holds exactly
Why: Stationarity (2,0)+2(−1,0) = [0,0]; primal 1−1 = 0 ≤ 0; dual 2 ≥ 0; slackness 2·0 = 0. A complete, machine-checked KKT certificate for (1,0).
| KKT condition | expression | value |
|---|---|---|
| stationarity | ∇f + μ∇g | [0, 0] |
| primal feasibility | g = 1 − x | 0 (≤ 0) |
| dual feasibility | μ | 2 (≥ 0) |
| compl. slackness | μ · g | 0 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0.
Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).
Why: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.
Trap
There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.
Assume μᵢ > 0 for all i
Why: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.
Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).
Active ⇒ μᵢ > 0; inactive ⇒ μᵢ = 0
Why: Only boundary-touching constraints carry weight. This is precisely what makes an SVM sparse: most points are strictly inside their margin (inactive) and get α = 0, so they don't shape the boundary at all.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.
Section
Part 4 of 6 — KKT in the wild
Concept
A hard-margin support vector machine finds the separating boundary with the widest margin. Maximizing margin = 2/‖w‖ is the same as minimizing ½‖w‖², subject to every point being on the correct side by at least the margin:
\[ \min_{w, b}\; \tfrac12 \lVert w \rVert^2 \quad \text{s.t.} \quad y_i\,(w^\top x_i + b) \ge 1 \;\; \forall i \]
Each data point contributes one inequality constraint. This is a convex quadratic program with linear inequalities — exactly the KKT setting we just built.
Concept
Keep it 1-D so we can do it by hand. Two negatives at x=1, 2 and two positives at x=4, 5. The model is sign(w·x + b).
| point x | label y | class |
|---|---|---|
| 1 | −1 | negative |
| 2 | −1 | negative |
| 4 | +1 | positive |
| 5 | +1 | positive |
Intuitively the boundary should sit at x=3 (the midpoint of the gap), and the two inner points x=2 and x=4 should be the ones that pin it. Let's prove that.
Pattern
Step through it
Step through Our running SVM: four points on a line one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
Many lines separate the two classes. The one that leaves the widest empty corridor between them is the safest — it is farthest from every point, so small wiggles in the data won't flip a prediction.
That corridor's half-width is 1/‖w‖. Widening the corridor means shrinking ‖w‖, which is why the objective is min ½‖w‖². The constraints just forbid any point from entering the corridor on the wrong side.
Ranking
Put in order
Put the moves of Solve the SVM primal by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (w·4+b) − (w·2+b) = 1 − (−1) gives 2w = 2, so w = 1.
Worked example
The widest margin is set by the closest opposing pair, x=2 (neg) and x=4 (pos). Put them exactly on the margin: w·x + b = −1 at x=2 and = +1 at x=4.
Subtract the two margin equations
Why: (w·4+b) − (w·2+b) = 1 − (−1) gives 2w = 2, so w = 1.
\[ (4w + b) - (2w + b) = 1 - (-1) \;\Longrightarrow\; 2w = 2 \;\Longrightarrow\; w = 1 \]
Back-substitute for b
Why: From w·4 + b = 1 with w = 1: 4 + b = 1, so b = −3. Boundary w·x+b = 0 sits at x = 3, the midpoint — as expected.
\[ \boxed{\,w = 1, \quad b = -3, \quad \text{boundary at } x = 3\,} \]
Half-margin = 1/‖w‖ = 1
Why: The margin planes sit at x = 2 and x = 4, distance 1 either side of x = 3. Full margin 2/‖w‖ = 2.
Notation
Annotate
From Solve the SVM primal by hand — read this one piece at a time. What is each part doing?
On: \( \boxed{\,w = 1, \quad b = -3, \quad \text{boundary at } x = 3\,} \)
Estimation
Predict first
A linear SVC with huge C approximates the hard margin. It should recover w=1, b=−3 and flag x=2, 4 as the support vectors:
Commit before you compute: what does Confirm the primal with scikit-learn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: sklearn matches the hand answer exactly
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Prints w = [1.], b = [-3.], support vectors x = [2.
Worked example
A linear SVC with huge C approximates the hard margin. It should recover w=1, b=−3 and flag x=2, 4 as the support vectors:
import numpy as np
from sklearn.svm import SVC
X = np.array([1., 2., 4., 5.]).reshape(-1, 1)
y = np.array([-1., -1., 1., 1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print('w =', clf.coef_.ravel(), ' b =', clf.intercept_.round(4))
print('support vectors x =', clf.support_vectors_.ravel())
print('alpha*y =', clf.dual_coef_.ravel())sklearn matches the hand answer exactly
Why: Prints w = [1.], b = [-3.], support vectors x = [2. 4.], and signed multipliers [-0.5, 0.5]. The two INNER points are the support vectors; x=1 and x=5 are not.
| quantity | value (verified) |
|---|---|
| w | [1.] |
| b | [-3.] |
| support vectors (x) | [2., 4.] |
| dual_coef_ (αᵢ yᵢ) | [-0.5, 0.5] |
Concept
Attach a multiplier αᵢ ≥ 0 to each margin constraint 1 − yᵢ(w·xᵢ+b) ≤ 0 and form the Lagrangian:
\[ \mathcal{L}(w,b,\alpha) = \tfrac12\lVert w\rVert^2 + \sum_i \alpha_i\big(1 - y_i(w^\top x_i + b)\big) \]
Stationarity in w and b gives the two eliminations that turn the primal into a pure-α problem:
\[ \nabla_w \mathcal{L} = 0 \Rightarrow w = \sum_i \alpha_i y_i x_i, \qquad \frac{\partial \mathcal{L}}{\partial b} = 0 \Rightarrow \sum_i \alpha_i y_i = 0 \]
Concept
Substituting w = Σ αᵢ yᵢ xᵢ back into L cancels w and b and leaves a quadratic program in α alone — the SVM dual:
\[ \max_{\alpha \ge 0}\; \sum_i \alpha_i - \tfrac12 \sum_{i,j} \alpha_i \alpha_j\, y_i y_j\, x_i^\top x_j \quad \text{s.t.} \quad \sum_i \alpha_i y_i = 0 \]
Two payoffs: the data enters only through inner products xᵢᵀxⱼ (the door to kernels), and by complementary slackness most αᵢ will be 0. Let's solve it.
Fill the middle
Fill in the blanks
From Solve the SVM dual as a QP — one line has had its right-hand side removed. Put it back.
import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x) # G_ij = y_i y_j x_i x_j
obj = lambda a: 0.5 * a @ G @ a - a.sum() # minimize -dual
cons = (**(a * y) @ x**,) # sum a_i y_i = 0
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
constraints=cons, method='SLSQP').x
print('alpha =', a.round(4))
w = ___
k = a.argmax()
b = y[k] - w * x[k]
print('w =', round(w, 4), ' b =', round(b, 4))
Why: w is what everything below it consumes, so the wrong expression here fails later and somewhere else. Prints alpha = [0. 0.5 0.5 0.] and w = 1.0, b = -3.0 — identical to the primal and to sklearn.
Worked example
Build the Gram matrix Gᵢⱼ = yᵢ yⱼ xᵢᵀxⱼ, then minimize ½αᵀGα − Σαᵢ with α ≥ 0 and Σαᵢ yᵢ = 0. Recover w and b from the α:
import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x) # G_ij = y_i y_j x_i x_j
obj = lambda a: 0.5 * a @ G @ a - a.sum() # minimize -dual
cons = ({'type': 'eq', 'fun': lambda a: a @ y},) # sum a_i y_i = 0
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
constraints=cons, method='SLSQP').x
print('alpha =', a.round(4))
w = (a * y) @ x
k = a.argmax()
b = y[k] - w * x[k]
print('w =', round(w, 4), ' b =', round(b, 4))The dual reproduces the primal
Why: Prints alpha = [0. 0.5 0.5 0.] and w = 1.0, b = -3.0 — identical to the primal and to sklearn. Only x=2 and x=4 carry nonzero α.
| point x | label y | α (verified) |
|---|---|---|
| 1 | −1 | 0.0 |
| 2 | −1 | 0.5 ← support vector |
| 4 | +1 | 0.5 ← support vector |
| 5 | +1 | 0.0 |
Comparison
Comparison matrix
From Solve the SVM dual as a QP: refill the label y column from what you know. The rest of the table is as it appeared.
| point x | label y | α (verified) |
|---|---|---|
| 1 | −1 | 0.0 |
| 2 | −1 | 0.5 ← support vector |
| 4 | +1 | 0.5 ← support vector |
| 5 | +1 | 0.0 |
Pattern
Predict first
The table runs: 1 | 0.0 | 2.0 | 0.0 (inactive) · 2 | 0.5 | 1.0 | 0.0 (active SV) · 4 | 0.5 | 1.0 | 0.0 (active SV)
In Complementary slackness makes it sparse, given the rows so far: what is the next one — the row where x is 5?
Correct: 5 | 0.0 | 2.0 | 0.0 (inactive)
| x | α | margin y(wx+b) | α·(margin−1) |
|---|---|---|---|
| 1 | 0.0 | 2.0 | 0.0 (inactive) |
| 2 | 0.5 | 1.0 | 0.0 (active SV) |
| 4 | 0.5 | 1.0 | 0.0 (active SV) |
| 5 | 0.0 | 2.0 | 0.0 (inactive) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Points x=2, 4 sit exactly ON the margin (margin = 1, constraint ACTIVE) and carry α = 0.5.
Worked example
For each point, complementary slackness says αᵢ·(yᵢ(w·xᵢ+b) − 1) = 0. Compute the margin yᵢ(w·xᵢ+b) for every point and watch the pattern:
import numpy as np
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
alpha = np.array([0., 0.5, 0.5, 0.])
w, b = 1.0, -3.0
margin = y * (w * x + b) # y_i (w x_i + b)
for xi, ai, mi in zip(x, alpha, margin):
print(f'x={xi:.0f} alpha={ai:.1f} margin={mi:.1f} alpha*(margin-1)={ai*(mi-1):.1f}')Nonzero α only where the margin equals 1
Why: Points x=2, 4 sit exactly ON the margin (margin = 1, constraint ACTIVE) and carry α = 0.5. Points x=1, 5 are past the margin (margin = 2, INACTIVE) and carry α = 0. Every product α(margin−1) = 0 — complementary slackness holds.
| x | α | margin y(wx+b) | α·(margin−1) |
|---|---|---|---|
| 1 | 0.0 | 2.0 | 0.0 (inactive) |
| 2 | 0.5 | 1.0 | 0.0 (active SV) |
| 4 | 0.5 | 1.0 | 0.0 (active SV) |
| 5 | 0.0 | 2.0 | 0.0 (inactive) |
Missing information
Discussion prompt
The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
w = 0.5·(−1)·2 + 0.5·(+1)·4 = 1.0, matching the primal w. And Σαᵢyᵢ = −0.5 + 0.5 = 0, the b-stationarity condition. The dual's derivation is faithful on real numbers.
Worked example
The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:
import numpy as np
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
alpha = np.array([0., 0.5, 0.5, 0.])
w = (alpha * y) @ x # w = sum alpha_i y_i x_i
print('w = sum a_i y_i x_i =', round(w, 4))
print('sum a_i y_i =', round((alpha * y).sum(), 4))Both identities check out
Why: w = 0.5·(−1)·2 + 0.5·(+1)·4 = 1.0, matching the primal w. And Σαᵢyᵢ = −0.5 + 0.5 = 0, the b-stationarity condition. The dual's derivation is faithful on real numbers.
| identity | computed | expected |
|---|---|---|
| w = Σ αᵢ yᵢ xᵢ | 1.0 | 1.0 (primal w) |
| Σ αᵢ yᵢ | 0.0 | 0 (b-stationarity) |
| contributing points | x=2, x=4 | the support vectors |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Both identities check out
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:
Concept
Look again at the dual objective: the features appear only inside xᵢᵀxⱼ. The prediction too — w·x + b = Σᵢ αᵢ yᵢ (xᵢᵀx) + b — never needs the raw w, just inner products with the support vectors.
So you can replace every xᵢᵀxⱼ with a kernel K(xᵢ, xⱼ) — an inner product in some richer feature space — and get nonlinear boundaries for free. That is the kernel trick, and it exists purely because Lagrangian duality rewrote the SVM in terms of inner products.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The extreme points x=1 and x=5 are the most 'characteristic' of each class, so they must be the support vectors.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0.
Support vectors are the points whose margin constraint is active (margin = 1).
Why: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0. They could be deleted and the boundary would not move.
Trap
The extreme points x=1 and x=5 are the most 'characteristic' of each class, so they must be the support vectors.
Call x=1 and x=5 the support vectors
Why: Wrong: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0. They could be deleted and the boundary would not move.
Support vectors are the points whose margin constraint is active (margin = 1).
Support vectors are x=2 and x=4
Why: The two INNER points sit exactly on the margin, so their constraints bind and α = 0.5 > 0. Only these shape the boundary — delete any other point and w, b are unchanged. 'Support vector' = active constraint, not 'extreme point'.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.; One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:; Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.x²+y², so minimize f directly and report the origin.; There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.Section
Part 5 of 6 — why the dual is safe
Intuition
Here's the trick that spawns the dual. Write the primal as min_x max_{μ≥0} L(x,μ). The inner max over μ is a referee: if x breaks a constraint (g>0), it sends μ→∞ and the value blows up; if x is feasible, the best μ is 0 and you're left with f(x).
So min_x max_μ L is exactly the original constrained problem. The dual simply swaps the order to max_μ min_x L. Swapping min and max can only shrink the value — which is why the dual is always a lower bound (weak duality).
Concept
The dual function g(λ) is the Lagrangian minimized over x. For any λ, it lower-bounds the primal optimum p⋆ — that is weak duality, and it always holds:
\[ g(\lambda) = \min_x \mathcal{L}(x,\lambda) \;\le\; p^\star \]
For our running problem, minimizing L = x²+y²+λ(x+y−1) over x,y gives x=y=−λ/2, so g(λ) = −λ²/2 − λ. The best lower bound is max_λ g(λ) — the dual problem.
Explain it
Discussion prompt
Explain The dual function and weak duality to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The dual function g(λ) is the Lagrangian minimized over x. For any λ, it lower-bounds the primal optimum p⋆ — that is weak duality, and it always holds:
Estimation
Predict first
Evaluate g(λ) = −λ²/2 − λ and maximize it. Its peak should equal the primal optimum p⋆ = 0.5 — a zero duality gap:
Commit before you compute: what does Strong duality: the gap is zero come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Dual optimum = primal optimum
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. g peaks at λ = −1 with g(−1) = 0.5, exactly the primal f⋆ = 0.5.
Worked example
Evaluate g(λ) = −λ²/2 − λ and maximize it. Its peak should equal the primal optimum p⋆ = 0.5 — a zero duality gap:
def g(lam): # inf_x L = -lam^2/2 - lam
return -lam**2 / 2 - lam
for lam in [-2., -1., -0.5, 0.]:
print(f'lam={lam:+.1f} g(lam)={g(lam):.3f}')
print('max g =', g(-1.0), ' = primal f* = 0.5 -> gap = 0')Dual optimum = primal optimum
Why: g peaks at λ = −1 with g(−1) = 0.5, exactly the primal f⋆ = 0.5. The lower bound is tight — strong duality holds, so solving the dual solves the primal.
| λ | g(λ) (verified) | vs p⋆ = 0.5 |
|---|---|---|
| −2.0 | 0.000 | below |
| −1.0 | 0.500 | equal ← optimum |
| −0.5 | 0.375 | below |
| 0.0 | 0.000 | below |
Discrimination
Sort into buckets
Sort these by g(λ) (verified), from memory, without looking back at Strong duality: the gap is zero. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Strong duality is not automatic — it needs a certificate. Slater's condition is the usual one: if the problem is convex and there exists a strictly feasible point (every inequality satisfied with strict slack, gᵢ(x) < 0), then the duality gap is zero.
The SVM primal is convex (quadratic objective, linear constraints) and, for separable data, strictly feasible — so Slater applies. That is precisely why solving the SVM dual yields the true primal optimum, as we just saw on numbers.
Intuition
Weak duality is free, but the gap need not be zero. For a non-convex problem the dual can sit strictly below the primal — solving the dual then only gives a lower bound, not the answer.
Slater's certificate needs both ingredients: convexity and a strictly feasible point. Miss either and the gap can open. Fortunately the SVM, least squares, and most convex ML objectives satisfy Slater, so their duals are exact — which is why duality is a workhorse and not a curiosity.
Section
Part 6 of 6
Ranking
Put in order
These are the steps of The constrained-optimization recipe, scrambled. Put them back in order before the next slide shows you.
L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)∇ₓ L = 0 for the primal variablesh=0, g≤0 and dual μᵢ ≥ 0μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)λ as the shadow priceWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)∇ₓ L = 0 for the primal variablesh=0, g≤0 and dual μᵢ ≥ 0μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)λ as the shadow priceEdge cases
Discussion prompt
The constrained-optimization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)∇ₓ L = 0 for the primal variablesh=0, g≤0 and dual μᵢ ≥ 0μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)λ as the shadow priceElimination
Eliminate the wrong options
At the solution of min f s.t. h(x)=0, stationarity gives ∇f + λ∇h = 0. What does this say geometrically?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: ∇f = −λ∇h means ∇f points straight across the constraint surface (along ∇h). Every legal step runs along the surface, perpendicular to ∇f, so no feasible move decreases f — that is optimality on the constraint.
Check
Picture the smallest level set kissing the constraint.
Check your understanding
At the solution of min f s.t. h(x)=0, stationarity gives ∇f + λ∇h = 0. What does this say geometrically?
Answer: A
Why: ∇f = −λ∇h means ∇f points straight across the constraint surface (along ∇h). Every legal step runs along the surface, perpendicular to ∇f, so no feasible move decreases f — that is optimality on the constraint.
Prediction
Predict first
For min x²+y² s.t. x+y=c, the multiplier is λ=−1 at c=1. If you loosen the constraint to c=1.01, the optimal value f⋆ changes by approximately:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: +0.01, because df⋆/dc = −λ = 1
Why: The envelope theorem gives df⋆/dc = −λ = 1, so a step of Δc = 0.01 raises f⋆ by about 0.01. Numerically f⋆ went 0.50000 → 0.51005, a rise of ≈ 0.01.
Check
Recall f⋆(c) for the constraint x+y=c.
Check your understanding
For min x²+y² s.t. x+y=c, the multiplier is λ=−1 at c=1. If you loosen the constraint to c=1.01, the optimal value f⋆ changes by approximately:
Answer: A
Why: The envelope theorem gives df⋆/dc = −λ = 1, so a step of Δc = 0.01 raises f⋆ by about 0.01. Numerically f⋆ went 0.50000 → 0.51005, a rise of ≈ 0.01.
Commit first
Predict first
An inequality constraint gᵢ(x) ≤ 0 is strictly satisfied at the optimum, gᵢ(x⋆) < 0. Its KKT multiplier μᵢ is:
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: 0 — the constraint is inactive and exerts no force
Why: Complementary slackness μᵢ gᵢ = 0 with gᵢ < 0 forces μᵢ = 0. A constraint satisfied with room to spare pushes on nothing — delete it and the optimum is unchanged.
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Active or inactive?
Check your understanding
An inequality constraint gᵢ(x) ≤ 0 is strictly satisfied at the optimum, gᵢ(x⋆) < 0. Its KKT multiplier μᵢ is:
Answer: A
Why: Complementary slackness μᵢ gᵢ = 0 with gᵢ < 0 forces μᵢ = 0. A constraint satisfied with room to spare pushes on nothing — delete it and the optimum is unchanged.
Prediction
Predict first
For our SVM the dual optimum equals the primal optimum. Which statement is correct in general?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Weak duality (dual ≤ primal) always holds; the gap is zero when the problem is convex and Slater's condition holds
Why: Weak duality g(λ) ≤ p⋆ is automatic (min-max ≥ max-min). The gap closes — strong duality — for convex problems with a strictly feasible point (Slater). The SVM is convex and strictly feasible on separable data, so its dual is exact.
Check
Weak vs strong — mind the gap.
Check your understanding
For our SVM the dual optimum equals the primal optimum. Which statement is correct in general?
Answer: A
Why: Weak duality g(λ) ≤ p⋆ is automatic (min-max ≥ max-min). The gap closes — strong duality — for convex problems with a strictly feasible point (Slater). The SVM is convex and strictly feasible on separable data, so its dual is exact.
Elimination
Eliminate the wrong options
In our 4-point SVM (neg at x=1,2; pos at x=4,5) the dual gives α = [0, 0.5, 0.5, 0]. Why do exactly x=2 and x=4 have α > 0?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: x=2 and x=4 have margin y(wx+b) = 1 exactly — their constraints are active — so complementary slackness α(margin−1)=0 permits α > 0. They are the support vectors. x=1 and x=5 have margin 2 (inactive), forcing α = 0.
Check
Which points survive with α > 0?
Check your understanding
In our 4-point SVM (neg at x=1,2; pos at x=4,5) the dual gives α = [0, 0.5, 0.5, 0]. Why do exactly x=2 and x=4 have α > 0?
Answer: A
Why: x=2 and x=4 have margin y(wx+b) = 1 exactly — their constraints are active — so complementary slackness α(margin−1)=0 permits α > 0. They are the support vectors. x=1 and x=5 have margin 2 (inactive), forcing α = 0.
Section
The project
Concept
Build the whole chain yourself: an equality problem by Lagrange and scipy, an inequality with an active constraint, then the 4-point SVM dual with its support vectors. You've derived every piece — now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | Equality: min x²+y² s.t. x+y=1 | Lagrange + minimize(eq) |
| 2 | Inequality x≥1 → find the active constraint | minimize(ineq) |
| 3 | SVM dual → α and the support vectors | Gram + minimize(SLSQP) |
Build rules: type every line yourself, pass each scipy constraint as a dict ('type':'eq'/'ineq'), and after each solve, check which constraints are active at the answer.
Analogy
Discussion prompt
Explain Project: from Lagrange to the SVM dual by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Build rules: type every line yourself, pass each scipy constraint as a dict ('type':'eq'/'ineq'), and after each solve, check which constraints are active at the answer.
Worked example
Your turn: set ∇L=0 for f=x²+y² with x+y=1, predict (x,y), then confirm with scipy. Say the answer aloud before you print.
Hint: 2x+λ=0, 2y+λ=0 ⇒ x=y; substitute into x+y=1. For scipy use constraints={'type':'eq','fun':lambda v: v[0]+v[1]-1}.
from scipy.optimize import minimize
# by hand: 2x=-lam, 2y=-lam -> x=y ; x+y=1 -> x=y=0.5, f=0.5, lam=-1
res = minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
constraints={'type': 'eq', 'fun': lambda v: v[0] + v[1] - 1})
print(res.x.round(4), round(res.fun, 4))| source | x⋆, y⋆ | f⋆ | λ |
|---|---|---|---|
| by hand | [0.5, 0.5] | 0.5 | −1 |
| scipy | [0.5, 0.5] | 0.5 | — |
Worked example
Your turn: minimize x²+y² subject to x ≥ 1. Predict where the solution lands and whether the constraint binds, then verify.
Hint: scipy's 'ineq' means fun ≥ 0, so pass x−1. Test activeness by checking whether x⋆−1 == 0 at the solution.
from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
constraints={'type': 'ineq', 'fun': lambda v: v[0] - 1})
print(res.x.round(4), round(res.fun, 4))
print('active:', round(res.x[0] - 1, 4) == 0.0)| quantity | value |
|---|---|
| x⋆, y⋆ | [1, 0] |
| f⋆ | 1.0 |
| active (x=1)? | True |
Pattern
Step through it
Step through Milestone 2 — an active inequality one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: solve the dual QP for the 4 points, then recover w, b, and read off which points have α > 0. Predict the support vectors first.
Hint: G = np.outer(y,y)*np.outer(x,x); minimize 0.5*a@G@a - a.sum() with bounds=[(0,None)]*4 and equality a@y==0. Then w=(a*y)@x and b=y[k]-w*x[k] for any support-vector index k.
import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x)
obj = lambda a: 0.5 * a @ G @ a - a.sum()
cons = ({'type': 'eq', 'fun': lambda a: a @ y},)
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
constraints=cons, method='SLSQP').x
w = (a * y) @ x
k = a.argmax()
b = y[k] - w * x[k]
print('alpha =', a.round(4), ' w =', round(w, 4), ' b =', round(b, 4))| output | value |
|---|---|
| alpha | [0. 0.5 0.5 0.] |
| support vectors (x) | 2 and 4 |
| w, b | 1.0, -3.0 |
Trade off
Comparison matrix
From Milestone 3 — the SVM dual and its support vectors: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| output | value |
|---|---|
| alpha | [0. 0.5 0.5 0.] |
| support vectors (x) | 2 and 4 |
| w, b | 1.0, -3.0 |
Concept
import numpy as np
from scipy.optimize import minimize
# 1. equality min x^2+y^2 s.t. x+y=1 -> (0.5, 0.5)
eq = minimize(lambda v: v[0]**2+v[1]**2, [0.,0.],
constraints={'type':'eq','fun':lambda v: v[0]+v[1]-1})
# 2. inequality min x^2+y^2 s.t. x>=1 -> (1, 0), active
ineq = minimize(lambda v: v[0]**2+v[1]**2, [2.,2.],
constraints={'type':'ineq','fun':lambda v: v[0]-1})
# 3. SVM dual -> alpha=[0,0.5,0.5,0], w=1, b=-3
x = np.array([1.,2.,4.,5.]); y = np.array([-1.,-1.,1.,1.])
G = np.outer(y,y)*np.outer(x,x)
a = minimize(lambda a: 0.5*a@G@a - a.sum(), np.zeros(4),
bounds=[(0,None)]*4,
constraints=({'type':'eq','fun':lambda a: a@y},)).x
w = (a*y)@x; b = y[a.argmax()] - w*x[a.argmax()]
print('equality :', eq.x.round(3), round(eq.fun,3))
print('inequality:', ineq.x.round(3), round(ineq.fun,3))
print('svm alpha :', a.round(3), ' w,b =', round(w,3), round(b,3))| printed line | value (verified) |
|---|---|
| equality : | [0.5 0.5] 0.5 |
| inequality: | [1. -0.] 1.0 |
| svm alpha : | [0. 0.5 0.5 0.] w,b = 1.0 -3.0 |
If the equality lands at (0.5, 0.5), the inequality binds at (1, 0), and the SVM gives α = [0, 0.5, 0.5, 0] with w=1, b=−3 — you can read the KKT conditions off a real SVM.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| equality : | [0.5 0.5] 0.5 |
| inequality: | [1. -0.] 1.0 |
| svm alpha : | [0. 0.5 0.5 0.] w,b = 1.0 -3.0 |
Concept
Slides closed, out loud: explain (1) why ∇f ∥ ∇h at an equality optimum, (2) what complementary slackness says about active vs inactive constraints, and (3) why only support vectors have α > 0.
Stretch (homework): use Lagrange multipliers to show the max-entropy distribution with fixed mean and variance is the Gaussian, and re-derive the SVM dual from the primal Lagrangian yourself. KKT returns for kernels and soft margins (Week 25) and for RLHF's KL constraint (Week 40).
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why ∇f ∥ ∇h at an equality optimum, (2) what complementary slackness says about active vs inactive constraints, and (3) why only support vectors have α > 0.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The problem: rules to obey · The Lagrangian · Inequalities & KKT · The SVM, primal and dual · Duality & Slater · Recipe, checks & your turn. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
min f s.t. h=0, g≤0 and explain why the unconstrained min is usually infeasible∇f = −λ∇h from x-stationarity, and recover the constraint from λ-stationarityλ as a shadow price ∂f⋆/∂c = −λ — confirmed at f⋆(1)=0.5, f⋆(1.01)=0.51005α=[0,0.5,0.5,0], support vectors x=2,4), and know Slater ⇒ zero duality gap| idea | the one thing to remember |
|---|---|
| Lagrangian | ∇f = −λ∇h at the optimum (aligned gradients) |
| multiplier | λ is a shadow price: ∂f⋆/∂c = −λ |
| KKT | stationarity + primal/dual feasibility + slackness |
| slackness | μᵢ gᵢ = 0: active ⇒ μ>0, inactive ⇒ μ=0 |
| SVM dual | QP in α≥0; support vectors are the active constraints |
| duality | convex + Slater ⇒ strong duality, gap = 0 |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.