Lesson 24: Constrained Optimization & KKT

USAAIO Lesson 24, from Week 8, fully worked, building constrained optimization from the ground up. One running toy problem - minimizing x² + y² on a line - carries the derivation of the Lagrangian gradient by gradient, and the multiplier is revealed as a shadow price. Inequality constraints then give the four KKT conditions, with complementary slackness proved on actual numbers. A four-point one-dimensional SVM is solved both as a primal problem and as a dual QP, giving α = [0, 0.5, 0.5, 0] with support vectors at x = 2 and x = 4, and the lesson closes strong duality and Slater's condition by matching the dual optimum to the primal. Every scipy and sklearn snippet runs standalone, and every number came from real execution. The lesson runs to 62 slides.

Subject: Machine Learning · 111 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Constrained Optimization & the KKT Conditions

Title

USAAIO · Lesson 24 · Week 8 (Optimization)

Minimizing with rules you must obey. We derive Lagrange multipliers gradient by gradient, read the multiplier as a price, extend to inequalities with the four KKT conditions, and finish by solving a real SVM as both a primal and a dual — skipping no step, with every number from live execution.

2. By the end of this lesson you can

Objectives

  1. Set up a constrained problem min f s.t. h=0, g≤0 and say why the unconstrained minimum usually fails
  2. Derive the Lagrangian condition ∇f = −λ∇h from 'no improving feasible direction' — no step skipped
  3. Read the multiplier λ as a shadow price: ∂f⋆/∂c = −λ, and verify it numerically
  4. State and check all four KKT conditions, including complementary slackness μᵢ gᵢ = 0
  5. Solve a 4-point SVM by hand as a primal, then as a dual QP in α ≥ 0, and see why only support vectors survive
  6. Explain strong duality and Slater's condition, and confirm the duality gap is 0

3. What survived from The Multivariate Gaussian?

Warm-up

Discussion prompt

Before we open Lesson 24: Constrained Optimization & KKT: without looking back, what was the main idea of The Multivariate Gaussian, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the multivariate Gaussian N(μ,Σ), why Σ must be PSD, Mahalanobis distance, Gaussian marginals/conditionals, the diagonal-Σ independence fact, and Cholesky sampling x=μ+Lz — with the VAE latent-space connection. Sample a 2D Gaussian via Cholesky and recover its covariance.

4. The problem: rules to obey

Section

Part 1 of 6

5. The constrained problem

Concept

Unconstrained optimization asks 'where is f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.

\[ \min_{x}\; f(x) \quad \text{s.t.} \quad h(x) = 0, \;\; g(x) \le 0 \]

feasible set — The set of all x that satisfy every constraint. Optimization now happens only inside this set — the true minimum of f may lie outside it, and then it is simply not allowed.

6. Break it if you can: The constrained problem

Counterexample

Discussion prompt

Unconstrained optimization asks 'where is f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Our running example

Concept

One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:

\[ \min_{x,y}\; f(x,y) = x^2 + y^2 \quad \text{s.t.} \quad h(x,y) = x + y - 1 = 0 \]

f alone is minimized at the origin (0,0). But (0,0) is off the line x+y=1, so it is forbidden. We need the closest feasible point instead.

8. By analogy: Our running example

Analogy

Discussion prompt

Explain Our running example by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:

9. The unconstrained answer is usually illegal

Intuition

Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.

The constraint is a straight line. We want the smallest circle that still touches the line. Too small and it misses the line entirely; too big and we overpaid.

The winning circle just barely kisses the line — touches it at exactly one point. That tangency is the whole secret, and the next slides turn 'tangent' into an equation.

10. Teach it back: The unconstrained answer is usually illegal

Explain it

Discussion prompt

Explain The unconstrained answer is usually illegal to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.

11. Picture it first: Tangency = aligned gradients

Picture it

Figure (svg): Concentric circles centered at the origin, a slanted line x+y=1, and the smallest circle touching the line at one point where the gradient of f and the gradient of h are drawn as parallel arrows.

The smallest circle that reaches the line touches it once; there ∇f and ∇h are parallel.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

At the kissing point the circle and the line are tangent — they share the same tangent direction. So their perpendiculars (their gradients) point the same way.

12. Tangency = aligned gradients

Concept

At the kissing point the circle and the line are tangent — they share the same tangent direction. So their perpendiculars (their gradients) point the same way.

Figure (svg): Concentric circles centered at the origin, a slanted line x+y=1, and the smallest circle touching the line at one point where the gradient of f and the gradient of h are drawn as parallel arrows.

The smallest circle that reaches the line touches it once; there ∇f and ∇h are parallel.

\[ \nabla f = -\lambda\, \nabla h \quad\text{for some scalar } \lambda \]

13. Why parallel gradients mean 'stuck'

Intuition

Stand on the constraint line. You may only step along it. Moving lowers f only if ∇f has a component along your step direction.

If ∇f points straight across the line (perpendicular to it), then every legal step is sideways to ∇f — no direction along the line decreases f. You are stuck at the optimum.

∇h is exactly the direction perpendicular to the line. So 'stuck' means ∇f points along ∇h: ∇f = −λ∇h. That scalar λ is the Lagrange multiplier.

14. Two ways a constraint can bite

Concept

It helps to see the two regimes before the algebra. An equality h=0 traps you on a surface — you can never leave it, so the optimum is almost always where f's level set is tangent to that surface.

An inequality g≤0 gives you a whole region. If the free minimum is already inside, the constraint does nothing; if it's outside, you get pushed to the boundary and the inequality behaves like an equality there.

constraintoptimum sitsmultiplier sign
equality h=0on the surface (always)λ any sign
inequality g≤0, activeon the boundary g=0μ ≥ 0
inequality g≤0, inactivein the interior g<0μ = 0

15. Fill in: optimum sits for Two ways a constraint can bite

Comparison

Comparison matrix

From Two ways a constraint can bite: refill the optimum sits column from what you know. The rest of the table is as it appeared.

constraintoptimum sitsmultiplier sign
equality h=0on the surface (always)λ any sign
inequality g≤0, activeon the boundary g=0μ ≥ 0
inequality g≤0, inactivein the interior g<0μ = 0

16. The Lagrangian

Section

Part 2 of 6 — derive it, then read λ

17. Package it into one function

Concept

Rather than juggle ∇f = −λ∇h and h=0 separately, fold them into a single scalar function — the Lagrangian — whose stationary point reproduces both.

\[ \mathcal{L}(x, \lambda) = f(x) + \lambda\, h(x) \]

λ is a brand-new variable we optimize over too. Setting all its partial derivatives to zero recovers both the tangency condition and the constraint, as the next two slides show.

18. What has to happen first: Stationarity in x recovers tangency

Ranking

Put in order

Put the moves of Stationarity in x recovers tangency into the order they have to happen.

  1. Differentiate L with respect to the x-variables
  2. Set it to zero
  3. This is exactly the tangency condition

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. ∂/∂x of f(x)+λ h(x) is ∇f + λ∇h.

19. Stationarity in x recovers tangency

Worked example

Differentiate L with respect to the x-variables

Why: ∂/∂x of f(x)+λ h(x) is ∇f + λ∇h. λ is treated as a constant here.

\[ \nabla_x \mathcal{L} = \nabla f(x) + \lambda\, \nabla h(x) \]

Set it to zero

Why: At a stationary point the gradient vanishes.

\[ \nabla f(x) + \lambda\, \nabla h(x) = 0 \;\Longleftrightarrow\; \nabla f = -\lambda\, \nabla h \]

This is exactly the tangency condition

Why: The gradients-parallel picture from Part 1 falls out with no extra work — the Lagrangian's x-stationarity IS 'aligned gradients'.

20. Decode the notation: Stationarity in x recovers tangency

Notation

Annotate

From Stationarity in x recovers tangency — read this one piece at a time. What is each part doing?

On: \( \nabla_x \mathcal{L} = \nabla f(x) + \lambda\, \nabla h(x) \)

  • ∂/∂x of f(x)+λ h(x) is ∇f + λ∇h. λ is treated as a constant here.
  • At a stationary point the gradient vanishes.
  • The gradients-parallel picture from Part 1 falls out with no extra work — the Lagrangian's x-stationarity IS 'aligned gradients'.

21. Plan first: Stationarity in λ recovers the constraint

Step zero

Discussion prompt

Stationarity in λ recovers the constraint — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Differentiate L with respect to λ

Answer:

  1. Differentiate L with respect to λ
  2. Set it to zero
  3. This is exactly the constraint

22. Stationarity in λ recovers the constraint

Worked example

Differentiate L with respect to λ

Why: ∂/∂λ of f(x)+λ h(x) is just h(x) — f has no λ, and λ h(x) differentiates to h(x).

\[ \frac{\partial \mathcal{L}}{\partial \lambda} = h(x) \]

Set it to zero

Why: Stationarity in the λ direction.

\[ \frac{\partial \mathcal{L}}{\partial \lambda} = 0 \;\Longrightarrow\; h(x) = 0 \]

This is exactly the constraint

Why: So a single condition ∇L = 0 (over both x and λ) bundles BOTH tangency and feasibility. That is why the Lagrangian is worth building.

23. Say it in words: Stationarity in λ recovers the constraint

Translation

\( \frac{\partial \mathcal{L}}{\partial \lambda} = 0 \;\Longrightarrow\; h(x) = 0 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

24. What has to happen first: Solve our running example by hand

Ranking

Put in order

Put the moves of Solve our running example by hand into the order they have to happen.

  1. ∂L/∂x = 2x + λ = 0 and ∂L/∂y = 2y + λ = 0
  2. ∂L/∂λ = x + y − 1 = 0
  3. Back out λ and f⋆

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Both give the same relation to λ, so 2x = 2y ⇒ x = y.

25. Solve our running example by hand

Worked example

For f = x²+y² and h = x+y−1, write out L = x²+y²+λ(x+y−1) and take all three partials:

∂L/∂x = 2x + λ = 0 and ∂L/∂y = 2y + λ = 0

Why: Both give the same relation to λ, so 2x = 2y ⇒ x = y. The problem's symmetry appears automatically.

\[ 2x + \lambda = 0,\qquad 2y + \lambda = 0 \;\Longrightarrow\; x = y \]

∂L/∂λ = x + y − 1 = 0

Why: The constraint. Substitute x = y: 2x = 1, so x = y = 0.5.

\[ x + y = 1,\; x = y \;\Longrightarrow\; x = y = \tfrac12 \]

Back out λ and f⋆

Why: λ = −2x = −1; f⋆ = 0.5² + 0.5² = 0.5. Answer: (0.5, 0.5), f⋆ = 0.5, λ = −1.

\[ \boxed{\,(x^\star, y^\star) = (\tfrac12, \tfrac12), \quad f^\star = \tfrac12, \quad \lambda = -1\,} \]

26. Work backwards from the answer: Solve our running example by hand

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Back out λ and f⋆

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

For f = x²+y² and h = x+y−1, write out L = x²+y²+λ(x+y−1) and take all three partials:

27. Guess the shape of the answer: Verify the hand answer against scipy

Estimation

Predict first

Hand it the same problem numerically. scipy.optimize.minimize takes each constraint as a dict; 'eq' means fun(v)=0. Runnable as-is:

Commit before you compute: what does Verify the hand answer against scipy come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: scipy lands on the same point

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Prints [0.5 0.5] 0.5 — identical to the by-hand (0.5, 0.5) with f⋆ = 0.5.

28. Verify the hand answer against scipy

Worked example

Hand it the same problem numerically. scipy.optimize.minimize takes each constraint as a dict; 'eq' means fun(v)=0. Runnable as-is:

from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
               constraints={'type': 'eq', 'fun': lambda v: v[0] + v[1] - 1})
print(res.x.round(4), round(res.fun, 4))

scipy lands on the same point

Why: Prints [0.5 0.5] 0.5 — identical to the by-hand (0.5, 0.5) with f⋆ = 0.5. The Lagrangian answer is confirmed by an independent solver.

quantityby handscipy (verified)
x⋆, y⋆[0.5, 0.5][0.5, 0.5]
f⋆0.50.5
λ−1— (implicit)

29. What each one costs: Verify the hand answer against scipy

Trade off

Comparison matrix

From Verify the hand answer against scipy: every row here is a choice with a cost. Fill the by hand column, then say which row you would actually pick and what you give up for it.

quantityby handscipy (verified)
x⋆, y⋆[0.5, 0.5][0.5, 0.5]
f⋆0.50.5
λ−1— (implicit)

30. Something is wrong here: just ignore the constraint

Anomaly

Predict first

A student writes this, and it looks reasonable:

It's still x²+y², so minimize f directly and report the origin.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e.

Fold the constraint in with a multiplier and solve ∇L = 0 and h = 0 together.

Why: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e. infeasible. The unconstrained minimum is almost always outside the feasible set, which is the entire reason constrained optimization exists.

31. Trap: just ignore the constraint

Trap

The trap

It's still x²+y², so minimize f directly and report the origin.

Answer (0, 0), f = 0

Why: But (0,0) gives x+y = 0 ≠ 1 — it is OFF the line, i.e. infeasible. The unconstrained minimum is almost always outside the feasible set, which is the entire reason constrained optimization exists.

The fix

Fold the constraint in with a multiplier and solve ∇L = 0 and h = 0 together.

Answer (0.5, 0.5), f = 0.5

Why: The multiplier λ = −1 balances the pull toward the origin against staying on x+y=1. The feasible optimum costs more than the illegal one — and that gap, exactly, is what λ prices.

32. λ is a price, not just a bookkeeping symbol

Concept

The multiplier has a meaning: it is the shadow price of the constraint — how fast the optimal value changes when you loosen the constraint by one unit.

Generalize the constraint to x + y = c. Then the optimal cost f⋆(c) depends on c, and the envelope theorem says its derivative is minus the multiplier:

\[ \frac{d f^\star}{d c} = -\lambda \]

For our line, f⋆(c) = c²/2, so df⋆/dc = c = 1 at c=1, matching −λ = −(−1) = 1. Let's confirm that on real numbers.

33. What has to be given first: The multiplier as a measured shadow price

Missing information

Discussion prompt

Solve the problem at c = 1.00 and c = 1.01, then estimate df⋆/dc by finite difference. It should come out near 1 = −λ:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Loosening the constraint to x+y=1.01 raised the optimum from 0.50000 to 0.51005; the rate 1.005 (a finite-difference estimate of the true 1.0) is exactly minus the multiplier. λ literally prices the constraint.

34. The multiplier as a measured shadow price

Worked example

Solve the problem at c = 1.00 and c = 1.01, then estimate df⋆/dc by finite difference. It should come out near 1 = −λ:

from scipy.optimize import minimize
def solve(c):
    return minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
                    constraints={'type': 'eq',
                                 'fun': lambda v: v[0] + v[1] - c}).fun
f1, f2 = solve(1.00), solve(1.01)
print('f*(1.00) =', round(f1, 5))
print('f*(1.01) =', round(f2, 5))
print('df*/dc   ~', round((f2 - f1) / 0.01, 4), '  (= -lambda = 1)')

The slope of f⋆ in c is ~1 = −λ

Why: Loosening the constraint to x+y=1.01 raised the optimum from 0.50000 to 0.51005; the rate 1.005 (a finite-difference estimate of the true 1.0) is exactly minus the multiplier. λ literally prices the constraint.

cf⋆(c) (verified)meaning
1.000.50000our original problem
1.010.51005constraint loosened by 0.01
Δf⋆/Δc1.005 ≈ 1= −λ, the shadow price

35. The general Lagrangian with mixed constraints

Concept

Real problems mix equalities and inequalities. Give each equality hⱼ a free multiplier λⱼ and each inequality gᵢ a nonnegative multiplier μᵢ, and add them all in:

\[ \mathcal{L}(x, \lambda, \mu) = f(x) + \sum_j \lambda_j\, h_j(x) + \sum_i \mu_i\, g_i(x), \qquad \mu_i \ge 0 \]

Stationarity ∇ₓL = 0 now reads ∇f + Σλⱼ∇hⱼ + Σμᵢ∇gᵢ = 0. The only new subtlety versus Part 2 is the sign restriction on the μᵢ — which the next slides earn geometrically.

36. Inequalities & KKT

Section

Part 3 of 6 — four conditions

37. Inequalities change the game

Concept

An inequality g(x) ≤ 0 carves out a whole region, not a curve. The optimum can sit inside the region (constraint irrelevant) or on its boundary (constraint binding).

active vs inactive — A constraint is ACTIVE at x⋆ if g(x⋆)=0 (the solution sits on the boundary, the constraint is pushing back). It is INACTIVE if g(x⋆)<0 (satisfied with room to spare, exerting no force).

We don't know in advance which case we're in — so we need conditions that cover both. Those are the KKT conditions.

38. Take the definitions apart: feasible set vs active vs inactive

Definition probe

Sort into buckets

Every line below is part of the definition of feasible set or of active vs inactive — one or the other, never both. Put each where it belongs.

feasible set
The set of all x that satisfy every constraint.; Optimization now happens only inside this set; the true minimum of f may lie outside it, and then it is simply not allowed.
active vs inactive
A constraint is ACTIVE at x⋆ if g(x⋆)=0 (the solution sits on the boundary, the constraint is pushing back).; It is INACTIVE if g(x⋆)<0 (satisfied with room to spare, exerting no force).
b1
The set of all x that satisfy every constraint. Optimization now happens only inside this set — the true minimum of f may lie outside it, and then it is simply not allowed.
b2
A constraint is ACTIVE at x⋆ if g(x⋆)=0 (the solution sits on the boundary, the constraint is pushing back). It is INACTIVE if g(x⋆)<0 (satisfied with room to spare, exerting no force).

39. The four KKT conditions

Concept

For min f s.t. gᵢ(x) ≤ 0 (with multipliers μᵢ), a KKT point satisfies all four of these at once:

  1. Stationarity: ∇f + Σᵢ μᵢ ∇gᵢ = 0
  2. Primal feasibility: gᵢ(x) ≤ 0 (obey the rules)
  3. Dual feasibility: μᵢ ≥ 0 (inequality multipliers are one-signed)
  4. Complementary slackness: μᵢ gᵢ(x) = 0

For convex problems these are not just necessary but sufficient: any point meeting all four is a global optimum. That is what makes them the workhorse of ML optimization.

40. Why the multiplier is one-signed now

Intuition

For an equality, λ could be any sign — you can be pushed either way along the constraint. For an inequality g ≤ 0, you can only be pushed inward, off the boundary.

So the constraint's force μ∇g can only point one way, which pins μ ≥ 0 (dual feasibility). A negative μ would mean the constraint is pulling the solution the wrong direction — impossible for a real barrier.

41. Where does each piece belong: Lesson 24: Constrained Optimization & KKT

Sorting

Sort into buckets

These are the pieces of Lesson 24: Constrained Optimization & KKT, out of order. Put each one back under the part of the lesson it belongs to.

The problem: rules to obey
The constrained problem; Our running example; The unconstrained answer is usually illegal
The Lagrangian
Package it into one function; Stationarity in x recovers tangency; Stationarity in λ recovers the constraint
Inequalities & KKT
Inequalities change the game; The four KKT conditions; Why the multiplier is one-signed now
s1
The problem: rules to obey is where Lesson 24: Constrained Optimization & KKT puts The constrained problem, Our running example, The unconstrained answer is usually illegal. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The Lagrangian is where Lesson 24: Constrained Optimization & KKT puts Package it into one function, Stationarity in x recovers tangency, Stationarity in λ recovers the constraint. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Inequalities & KKT is where Lesson 24: Constrained Optimization & KKT puts Inequalities change the game, The four KKT conditions, Why the multiplier is one-signed now. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

42. Complementary slackness in one sentence

Intuition

μᵢ gᵢ = 0 says: for each constraint, at least one of the pair is zero. Either the constraint is active (gᵢ = 0) and may push (μᵢ > 0), or it is inactive (gᵢ < 0) and cannot push (μᵢ = 0).

Never both slack and pushing. A constraint you satisfy with room to spare exerts no force on the answer — remove it and nothing changes. This single fact is what will make the SVM sparse.

43. Predict the next row: An ACTIVE inequality, solved and checked

Pattern

Predict first

The table runs: x⋆, y⋆ | [1, 0] | on the boundary · g = 1 − x at x⋆ | 0 | ACTIVE

In An ACTIVE inequality, solved and checked, given the rows so far: what is the next one — the row where quantity is μ?

Correct: μ | 2 > 0 | dual-feasible, may push

quantityvalue (verified)KKT reading
x⋆, y⋆[1, 0]on the boundary
g = 1 − x at x⋆0ACTIVE
μ2 > 0dual-feasible, may push

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Prints [1. -0.] 1.0 and active: True.

44. An ACTIVE inequality, solved and checked

Worked example

Minimize f = x²+y² subject to x ≥ 1, i.e. g = 1−x ≤ 0. The unconstrained min (0,0) violates it, so the boundary must bind. Solve numerically — scipy's 'ineq' means fun(v) ≥ 0, so pass x−1:

from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
               constraints={'type': 'ineq', 'fun': lambda v: v[0] - 1})
print(res.x.round(4), round(res.fun, 4))
print('active:', round(res.x[0] - 1, 4) == 0.0)

Solution sits ON the boundary x = 1

Why: Prints [1. -0.] 1.0 and active: True. The constraint is tight (g = 0), so complementary slackness permits μ > 0 — by hand, stationarity 2x = μ gives μ = 2.

quantityvalue (verified)KKT reading
x⋆, y⋆[1, 0]on the boundary
g = 1 − x at x⋆0ACTIVE
μ2 > 0dual-feasible, may push

45. The same problem, INACTIVE constraint

Worked example

Now relax the rule to x ≥ −1, i.e. g = −1−x ≤ 0. The unconstrained min (0,0) already obeys it with slack, so the constraint should exert no force and get μ = 0:

from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
               constraints={'type': 'ineq', 'fun': lambda v: v[0] + 1})
print(res.x.round(4), round(res.fun, 4))
print('slack g = x+1 =', round(res.x[0] + 1, 4), '-> mu = 0')

Solution ignores the boundary and sits at the origin

Why: Prints [-0. -0.] 0.0. Here g = x+1 = 1 > 0 (strictly satisfied), so complementary slackness μ·g = 0 forces μ = 0. Same problem shape, opposite KKT case.

quantityvalue (verified)KKT reading
x⋆, y⋆[0, 0]interior of feasible set
g = x + 1 at x⋆1 > 0INACTIVE (slack)
μ0no force on the solution

46. Restore the missing line: All four KKT conditions on numbers

Fill the middle

Fill in the blanks

From All four KKT conditions on numbers — one line has had its right-hand side removed. Put it back.

import numpy as np
xs, ys, mu = 1.0, 0.0, 2.0 # solution + multiplier
grad_f = np.array([2xs, 2ys]) # grad of x^2+y^2
grad_g = np.array([-1.0, 0.0]) # g(x)=1-x, so grad_g=(-1,0)
print('stationarity :', grad_f + mu*grad_g) # want [0,0]
print('primal feas :', 1 - xs, '<= 0') # g <= 0
print('dual feas :', mu, '>= 0') # mu >= 0
print('slackness :', mu * (1 - xs), '= 0') # mu*g = 0

Why: grad_g is what everything below it consumes, so the wrong expression here fails later and somewhere else. Stationarity (2,0)+2(−1,0) = [0,0]; primal 1−1 = 0 ≤ 0; dual 2 ≥ 0; slackness 2·0 = 0.

47. All four KKT conditions on numbers

Worked example

Take the active case (x⋆, y⋆) = (1, 0) with g = 1−x and μ = 2, and check each KKT condition explicitly. ∇f = (2x, 2y), ∇g = (−1, 0):

import numpy as np
xs, ys, mu = 1.0, 0.0, 2.0                    # solution + multiplier
grad_f = np.array([2*xs, 2*ys])               # grad of x^2+y^2
grad_g = np.array([-1.0, 0.0])                # g(x)=1-x, so grad_g=(-1,0)
print('stationarity :', grad_f + mu*grad_g)   # want [0,0]
print('primal feas  :', 1 - xs, '<= 0')       # g <= 0
print('dual feas    :', mu, '>= 0')           # mu >= 0
print('slackness    :', mu * (1 - xs), '= 0') # mu*g = 0

Every condition holds exactly

Why: Stationarity (2,0)+2(−1,0) = [0,0]; primal 1−1 = 0 ≤ 0; dual 2 ≥ 0; slackness 2·0 = 0. A complete, machine-checked KKT certificate for (1,0).

KKT conditionexpressionvalue
stationarity∇f + μ∇g[0, 0]
primal feasibilityg = 1 − x0 (≤ 0)
dual feasibilityμ2 (≥ 0)
compl. slacknessμ · g0

48. Something is wrong here: every constraint gets a positive multiplier

Anomaly

Predict first

A student writes this, and it looks reasonable:

There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0.

Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).

Why: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.

49. Trap: every constraint gets a positive multiplier

Trap

The trap

There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.

Assume μᵢ > 0 for all i

Why: An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.

The fix

Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).

Active ⇒ μᵢ > 0; inactive ⇒ μᵢ = 0

Why: Only boundary-touching constraints carry weight. This is precisely what makes an SVM sparse: most points are strictly inside their margin (inactive) and get α = 0, so they don't shape the boundary at all.

50. Break it on purpose: every constraint gets a positive multiplier

Break the constraint

Discussion prompt

The rule this trap just fixed:

Let complementary slackness sort constraints into active (μ>0) and inactive (μ=0).

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

An INACTIVE constraint (gᵢ < 0, slack) has μᵢ = 0 — complementary slackness μᵢ gᵢ = 0 with gᵢ ≠ 0 FORCES μᵢ = 0. Assuming all μᵢ > 0 gives an over-determined, inconsistent system.

51. The SVM, primal and dual

Section

Part 4 of 6 — KKT in the wild

52. The SVM as a constrained problem

Concept

A hard-margin support vector machine finds the separating boundary with the widest margin. Maximizing margin = 2/‖w‖ is the same as minimizing ½‖w‖², subject to every point being on the correct side by at least the margin:

\[ \min_{w, b}\; \tfrac12 \lVert w \rVert^2 \quad \text{s.t.} \quad y_i\,(w^\top x_i + b) \ge 1 \;\; \forall i \]

Each data point contributes one inequality constraint. This is a convex quadratic program with linear inequalities — exactly the KKT setting we just built.

53. Our running SVM: four points on a line

Concept

Keep it 1-D so we can do it by hand. Two negatives at x=1, 2 and two positives at x=4, 5. The model is sign(w·x + b).

point xlabel yclass
1−1negative
2−1negative
4+1positive
5+1positive

Intuitively the boundary should sit at x=3 (the midpoint of the gap), and the two inner points x=2 and x=4 should be the ones that pin it. Let's prove that.

54. Watch it run: Our running SVM: four points on a line

Pattern

Step through it

Step through Our running SVM: four points on a line one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: point x is 1
  2. Step 2: point x is 2
  3. Step 3: point x is 4
  4. Step 4: point x is 5

55. Why maximize the margin at all

Intuition

Many lines separate the two classes. The one that leaves the widest empty corridor between them is the safest — it is farthest from every point, so small wiggles in the data won't flip a prediction.

That corridor's half-width is 1/‖w‖. Widening the corridor means shrinking ‖w‖, which is why the objective is min ½‖w‖². The constraints just forbid any point from entering the corridor on the wrong side.

56. What has to happen first: Solve the SVM primal by hand

Ranking

Put in order

Put the moves of Solve the SVM primal by hand into the order they have to happen.

  1. Subtract the two margin equations
  2. Back-substitute for b
  3. Half-margin = 1/‖w‖ = 1

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. (w·4+b) − (w·2+b) = 1 − (−1) gives 2w = 2, so w = 1.

57. Solve the SVM primal by hand

Worked example

The widest margin is set by the closest opposing pair, x=2 (neg) and x=4 (pos). Put them exactly on the margin: w·x + b = −1 at x=2 and = +1 at x=4.

Subtract the two margin equations

Why: (w·4+b) − (w·2+b) = 1 − (−1) gives 2w = 2, so w = 1.

\[ (4w + b) - (2w + b) = 1 - (-1) \;\Longrightarrow\; 2w = 2 \;\Longrightarrow\; w = 1 \]

Back-substitute for b

Why: From w·4 + b = 1 with w = 1: 4 + b = 1, so b = −3. Boundary w·x+b = 0 sits at x = 3, the midpoint — as expected.

\[ \boxed{\,w = 1, \quad b = -3, \quad \text{boundary at } x = 3\,} \]

Half-margin = 1/‖w‖ = 1

Why: The margin planes sit at x = 2 and x = 4, distance 1 either side of x = 3. Full margin 2/‖w‖ = 2.

58. Decode the notation: Solve the SVM primal by hand

Notation

Annotate

From Solve the SVM primal by hand — read this one piece at a time. What is each part doing?

On: \( \boxed{\,w = 1, \quad b = -3, \quad \text{boundary at } x = 3\,} \)

  • (w·4+b) − (w·2+b) = 1 − (−1) gives 2w = 2, so w = 1.
  • From w·4 + b = 1 with w = 1: 4 + b = 1, so b = −3. Boundary w·x+b = 0 sits at x = 3, the midpoint — as expected.
  • The margin planes sit at x = 2 and x = 4, distance 1 either side of x = 3. Full margin 2/‖w‖ = 2.

59. Guess the shape of the answer: Confirm the primal with scikit-learn

Estimation

Predict first

A linear SVC with huge C approximates the hard margin. It should recover w=1, b=−3 and flag x=2, 4 as the support vectors:

Commit before you compute: what does Confirm the primal with scikit-learn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: sklearn matches the hand answer exactly

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Prints w = [1.], b = [-3.], support vectors x = [2.

60. Confirm the primal with scikit-learn

Worked example

A linear SVC with huge C approximates the hard margin. It should recover w=1, b=−3 and flag x=2, 4 as the support vectors:

import numpy as np
from sklearn.svm import SVC
X = np.array([1., 2., 4., 5.]).reshape(-1, 1)
y = np.array([-1., -1., 1., 1.])
clf = SVC(kernel='linear', C=1e6).fit(X, y)
print('w =', clf.coef_.ravel(), ' b =', clf.intercept_.round(4))
print('support vectors x =', clf.support_vectors_.ravel())
print('alpha*y =', clf.dual_coef_.ravel())

sklearn matches the hand answer exactly

Why: Prints w = [1.], b = [-3.], support vectors x = [2. 4.], and signed multipliers [-0.5, 0.5]. The two INNER points are the support vectors; x=1 and x=5 are not.

quantityvalue (verified)
w[1.]
b[-3.]
support vectors (x)[2., 4.]
dual_coef_ (αᵢ yᵢ)[-0.5, 0.5]

61. Build the Lagrangian, then eliminate w and b

Concept

Attach a multiplier αᵢ ≥ 0 to each margin constraint 1 − yᵢ(w·xᵢ+b) ≤ 0 and form the Lagrangian:

\[ \mathcal{L}(w,b,\alpha) = \tfrac12\lVert w\rVert^2 + \sum_i \alpha_i\big(1 - y_i(w^\top x_i + b)\big) \]

Stationarity in w and b gives the two eliminations that turn the primal into a pure-α problem:

\[ \nabla_w \mathcal{L} = 0 \Rightarrow w = \sum_i \alpha_i y_i x_i, \qquad \frac{\partial \mathcal{L}}{\partial b} = 0 \Rightarrow \sum_i \alpha_i y_i = 0 \]

62. The dual problem

Concept

Substituting w = Σ αᵢ yᵢ xᵢ back into L cancels w and b and leaves a quadratic program in α alone — the SVM dual:

\[ \max_{\alpha \ge 0}\; \sum_i \alpha_i - \tfrac12 \sum_{i,j} \alpha_i \alpha_j\, y_i y_j\, x_i^\top x_j \quad \text{s.t.} \quad \sum_i \alpha_i y_i = 0 \]

Two payoffs: the data enters only through inner products xᵢᵀxⱼ (the door to kernels), and by complementary slackness most αᵢ will be 0. Let's solve it.

63. Restore the missing line: Solve the SVM dual as a QP

Fill the middle

Fill in the blanks

From Solve the SVM dual as a QP — one line has had its right-hand side removed. Put it back.

import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x) # G_ij = y_i y_j x_i x_j
obj = lambda a: 0.5 * a @ G @ a - a.sum() # minimize -dual
cons = (**(a * y) @ x**,) # sum a_i y_i = 0
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
constraints=cons, method='SLSQP').x
print('alpha =', a.round(4))
w = ___
k = a.argmax()
b = y[k] - w * x[k]
print('w =', round(w, 4), ' b =', round(b, 4))

Why: w is what everything below it consumes, so the wrong expression here fails later and somewhere else. Prints alpha = [0. 0.5 0.5 0.] and w = 1.0, b = -3.0 — identical to the primal and to sklearn.

64. Solve the SVM dual as a QP

Worked example

Build the Gram matrix Gᵢⱼ = yᵢ yⱼ xᵢᵀxⱼ, then minimize ½αᵀGα − Σαᵢ with α ≥ 0 and Σαᵢ yᵢ = 0. Recover w and b from the α:

import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x)          # G_ij = y_i y_j x_i x_j
obj = lambda a: 0.5 * a @ G @ a - a.sum()    # minimize -dual
cons = ({'type': 'eq', 'fun': lambda a: a @ y},)   # sum a_i y_i = 0
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
             constraints=cons, method='SLSQP').x
print('alpha =', a.round(4))
w = (a * y) @ x
k = a.argmax()
b = y[k] - w * x[k]
print('w =', round(w, 4), ' b =', round(b, 4))

The dual reproduces the primal

Why: Prints alpha = [0. 0.5 0.5 0.] and w = 1.0, b = -3.0 — identical to the primal and to sklearn. Only x=2 and x=4 carry nonzero α.

point xlabel yα (verified)
1−10.0
2−10.5 ← support vector
4+10.5 ← support vector
5+10.0

65. Fill in: label y for Solve the SVM dual as a QP

Comparison

Comparison matrix

From Solve the SVM dual as a QP: refill the label y column from what you know. The rest of the table is as it appeared.

point xlabel yα (verified)
1−10.0
2−10.5 ← support vector
4+10.5 ← support vector
5+10.0

66. Predict the next row: Complementary slackness makes it sparse

Pattern

Predict first

The table runs: 1 | 0.0 | 2.0 | 0.0 (inactive) · 2 | 0.5 | 1.0 | 0.0 (active SV) · 4 | 0.5 | 1.0 | 0.0 (active SV)

In Complementary slackness makes it sparse, given the rows so far: what is the next one — the row where x is 5?

Correct: 5 | 0.0 | 2.0 | 0.0 (inactive)

xαmargin y(wx+b)α·(margin−1)
10.02.00.0 (inactive)
20.51.00.0 (active SV)
40.51.00.0 (active SV)
50.02.00.0 (inactive)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Points x=2, 4 sit exactly ON the margin (margin = 1, constraint ACTIVE) and carry α = 0.5.

67. Complementary slackness makes it sparse

Worked example

For each point, complementary slackness says αᵢ·(yᵢ(w·xᵢ+b) − 1) = 0. Compute the margin yᵢ(w·xᵢ+b) for every point and watch the pattern:

import numpy as np
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
alpha = np.array([0., 0.5, 0.5, 0.])
w, b = 1.0, -3.0
margin = y * (w * x + b)                      # y_i (w x_i + b)
for xi, ai, mi in zip(x, alpha, margin):
    print(f'x={xi:.0f}  alpha={ai:.1f}  margin={mi:.1f}  alpha*(margin-1)={ai*(mi-1):.1f}')

Nonzero α only where the margin equals 1

Why: Points x=2, 4 sit exactly ON the margin (margin = 1, constraint ACTIVE) and carry α = 0.5. Points x=1, 5 are past the margin (margin = 2, INACTIVE) and carry α = 0. Every product α(margin−1) = 0 — complementary slackness holds.

xαmargin y(wx+b)α·(margin−1)
10.02.00.0 (inactive)
20.51.00.0 (active SV)
40.51.00.0 (active SV)
50.02.00.0 (inactive)

68. What has to be given first: Check the two stationarity identities on…

Missing information

Discussion prompt

The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

w = 0.5·(−1)·2 + 0.5·(+1)·4 = 1.0, matching the primal w. And Σαᵢyᵢ = −0.5 + 0.5 = 0, the b-stationarity condition. The dual's derivation is faithful on real numbers.

69. Check the two stationarity identities on numbers

Worked example

The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:

import numpy as np
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
alpha = np.array([0., 0.5, 0.5, 0.])
w = (alpha * y) @ x                 # w = sum alpha_i y_i x_i
print('w = sum a_i y_i x_i =', round(w, 4))
print('sum a_i y_i        =', round((alpha * y).sum(), 4))

Both identities check out

Why: w = 0.5·(−1)·2 + 0.5·(+1)·4 = 1.0, matching the primal w. And Σαᵢyᵢ = −0.5 + 0.5 = 0, the b-stationarity condition. The dual's derivation is faithful on real numbers.

identitycomputedexpected
w = Σ αᵢ yᵢ xᵢ1.01.0 (primal w)
Σ αᵢ yᵢ0.00 (b-stationarity)
contributing pointsx=2, x=4the support vectors

70. Work backwards from the answer: Check the two stationarity identities on…

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Both identities check out

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The dual came from two eliminations: w = Σ αᵢ yᵢ xᵢ and Σ αᵢ yᵢ = 0. Verify both hold at the solved α = [0, 0.5, 0.5, 0]:

71. The payoff: data enters only as inner products

Concept

Look again at the dual objective: the features appear only inside xᵢᵀxⱼ. The prediction too — w·x + b = Σᵢ αᵢ yᵢ (xᵢᵀx) + b — never needs the raw w, just inner products with the support vectors.

So you can replace every xᵢᵀxⱼ with a kernel K(xᵢ, xⱼ) — an inner product in some richer feature space — and get nonlinear boundaries for free. That is the kernel trick, and it exists purely because Lagrangian duality rewrote the SVM in terms of inner products.

72. Something is wrong here: reading support vectors off the wrong points

Anomaly

Predict first

A student writes this, and it looks reasonable:

The extreme points x=1 and x=5 are the most 'characteristic' of each class, so they must be the support vectors.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0.

Support vectors are the points whose margin constraint is active (margin = 1).

Why: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0. They could be deleted and the boundary would not move.

73. Trap: reading support vectors off the wrong points

Trap

The trap

The extreme points x=1 and x=5 are the most 'characteristic' of each class, so they must be the support vectors.

Call x=1 and x=5 the support vectors

Why: Wrong: those are the points FARTHEST from the boundary, with margin 2 > 1 — their constraints are inactive, so α = 0. They could be deleted and the boundary would not move.

The fix

Support vectors are the points whose margin constraint is active (margin = 1).

Support vectors are x=2 and x=4

Why: The two INNER points sit exactly on the margin, so their constraints bind and α = 0.5 > 0. Only these shape the boundary — delete any other point and w, b are unchanged. 'Support vector' = active constraint, not 'extreme point'.

74. Which of these survive contact with Lesson 24: Constrained Optimization & KKT?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Unconstrained optimization asks 'where is f smallest?'. Constrained optimization adds rules the answer must respect — equalities h(x)=0 and inequalities g(x)≤0.; One tiny problem will carry the entire Lagrange story. Minimize the squared distance to the origin, but you must stay on a line:; Picture concentric circles x²+y²=r² growing out from the origin — level sets of f. Small circles are cheap, big circles are expensive.
Breaks
It's still x²+y², so minimize f directly and report the origin.; There are constraints in the problem, so give each one a positive multiplier μᵢ > 0.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 24: Constrained Optimization & KKT puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

75. Duality & Slater

Section

Part 5 of 6 — why the dual is safe

76. Primal as a min-max game

Intuition

Here's the trick that spawns the dual. Write the primal as min_x max_{μ≥0} L(x,μ). The inner max over μ is a referee: if x breaks a constraint (g>0), it sends μ→∞ and the value blows up; if x is feasible, the best μ is 0 and you're left with f(x).

So min_x max_μ L is exactly the original constrained problem. The dual simply swaps the order to max_μ min_x L. Swapping min and max can only shrink the value — which is why the dual is always a lower bound (weak duality).

77. The dual function and weak duality

Concept

The dual function g(λ) is the Lagrangian minimized over x. For any λ, it lower-bounds the primal optimum p⋆ — that is weak duality, and it always holds:

\[ g(\lambda) = \min_x \mathcal{L}(x,\lambda) \;\le\; p^\star \]

For our running problem, minimizing L = x²+y²+λ(x+y−1) over x,y gives x=y=−λ/2, so g(λ) = −λ²/2 − λ. The best lower bound is max_λ g(λ) — the dual problem.

78. Teach it back: The dual function and weak duality

Explain it

Discussion prompt

Explain The dual function and weak duality to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The dual function g(λ) is the Lagrangian minimized over x. For any λ, it lower-bounds the primal optimum p⋆ — that is weak duality, and it always holds:

79. Guess the shape of the answer: Strong duality: the gap is zero

Estimation

Predict first

Evaluate g(λ) = −λ²/2 − λ and maximize it. Its peak should equal the primal optimum p⋆ = 0.5 — a zero duality gap:

Commit before you compute: what does Strong duality: the gap is zero come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Dual optimum = primal optimum

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. g peaks at λ = −1 with g(−1) = 0.5, exactly the primal f⋆ = 0.5.

80. Strong duality: the gap is zero

Worked example

Evaluate g(λ) = −λ²/2 − λ and maximize it. Its peak should equal the primal optimum p⋆ = 0.5 — a zero duality gap:

def g(lam):                                   # inf_x L = -lam^2/2 - lam
    return -lam**2 / 2 - lam
for lam in [-2., -1., -0.5, 0.]:
    print(f'lam={lam:+.1f}   g(lam)={g(lam):.3f}')
print('max g =', g(-1.0), ' = primal f* = 0.5  ->  gap = 0')

Dual optimum = primal optimum

Why: g peaks at λ = −1 with g(−1) = 0.5, exactly the primal f⋆ = 0.5. The lower bound is tight — strong duality holds, so solving the dual solves the primal.

λg(λ) (verified)vs p⋆ = 0.5
−2.00.000below
−1.00.500equal ← optimum
−0.50.375below
0.00.000below

81. Which is which, by g(λ) (verified)

Discrimination

Sort into buckets

Sort these by g(λ) (verified), from memory, without looking back at Strong duality: the gap is zero. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.000
−2.0; 0.0
0.500
−1.0
0.375
−0.5
g1
g(λ) (verified) is "0.000" for −2.0, 0.0 — that is what the table on "Strong duality: the gap is zero" records, and it is the single property separating this group from the rest.
g2
g(λ) (verified) is "0.500" for −1.0 — that is what the table on "Strong duality: the gap is zero" records, and it is the single property separating this group from the rest.
g3
g(λ) (verified) is "0.375" for −0.5 — that is what the table on "Strong duality: the gap is zero" records, and it is the single property separating this group from the rest.

82. Slater's condition

Concept

Strong duality is not automatic — it needs a certificate. Slater's condition is the usual one: if the problem is convex and there exists a strictly feasible point (every inequality satisfied with strict slack, gᵢ(x) < 0), then the duality gap is zero.

The SVM primal is convex (quadratic objective, linear constraints) and, for separable data, strictly feasible — so Slater applies. That is precisely why solving the SVM dual yields the true primal optimum, as we just saw on numbers.

83. When strong duality can fail

Intuition

Weak duality is free, but the gap need not be zero. For a non-convex problem the dual can sit strictly below the primal — solving the dual then only gives a lower bound, not the answer.

Slater's certificate needs both ingredients: convexity and a strictly feasible point. Miss either and the gap can open. Fortunately the SVM, least squares, and most convex ML objectives satisfy Slater, so their duals are exact — which is why duality is a workhorse and not a curiosity.

84. Recipe, checks & your turn

Section

Part 6 of 6

85. Rebuild the recipe: The constrained-optimization recipe

Ranking

Put in order

These are the steps of The constrained-optimization recipe, scrambled. Put them back in order before the next slide shows you.

  1. Form the Lagrangian L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)
  2. Stationarity: solve ∇ₓ L = 0 for the primal variables
  3. Feasibility: enforce primal h=0, g≤0 and dual μᵢ ≥ 0
  4. Complementary slackness: μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)
  5. If convex + Slater, solve the dual instead — same optimum, and it exposes structure (support vectors, kernels)
  6. Sanity-check with scipy/sklearn and read λ as the shadow price

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

86. The constrained-optimization recipe

Pattern

  1. Form the Lagrangian L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)
  2. Stationarity: solve ∇ₓ L = 0 for the primal variables
  3. Feasibility: enforce primal h=0, g≤0 and dual μᵢ ≥ 0
  4. Complementary slackness: μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)
  5. If convex + Slater, solve the dual instead — same optimum, and it exposes structure (support vectors, kernels)
  6. Sanity-check with scipy/sklearn and read λ as the shadow price

87. Where does it stop working: The constrained-optimization recipe

Edge cases

Discussion prompt

The constrained-optimization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Form the Lagrangian L = f + Σⱼ λⱼ hⱼ + Σᵢ μᵢ gᵢ (one multiplier per constraint)
  2. Stationarity: solve ∇ₓ L = 0 for the primal variables
  3. Feasibility: enforce primal h=0, g≤0 and dual μᵢ ≥ 0
  4. Complementary slackness: μᵢ gᵢ = 0 — split constraints into active (μ>0) and inactive (μ=0)
  5. If convex + Slater, solve the dual instead — same optimum, and it exposes structure (support vectors, kernels)
  6. Sanity-check with scipy/sklearn and read λ as the shadow price

88. Rule out three: Check yourself — the Lagrangian condition

Elimination

Eliminate the wrong options

At the solution of min f s.t. h(x)=0, stationarity gives ∇f + λ∇h = 0. What does this say geometrically?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. ∇f is parallel to ∇h, so no feasible (along-the-constraint) step lowers f
  • B. ∇f = 0 at the solution
  • C. The constraint can be dropped once you know λ
  • D. λ must equal 0

Survives elimination: A

Why: ∇f = −λ∇h means ∇f points straight across the constraint surface (along ∇h). Every legal step runs along the surface, perpendicular to ∇f, so no feasible move decreases f — that is optimality on the constraint.

89. Check yourself — the Lagrangian condition

Check

Picture the smallest level set kissing the constraint.

Check your understanding

At the solution of min f s.t. h(x)=0, stationarity gives ∇f + λ∇h = 0. What does this say geometrically?

  • A. ∇f is parallel to ∇h, so no feasible (along-the-constraint) step lowers f (correct)
  • B. ∇f = 0 at the solution
  • C. The constraint can be dropped once you know λ
  • D. λ must equal 0

Answer: A

Why: ∇f = −λ∇h means ∇f points straight across the constraint surface (along ∇h). Every legal step runs along the surface, perpendicular to ∇f, so no feasible move decreases f — that is optimality on the constraint.

Why B tempts people
∇f need NOT vanish at a constrained optimum — it is balanced by λ∇h, not zero. ∇f = 0 only if the constraint happens to be inactive (then λ = 0).
Why C tempts people
The λ∇h term IS the constraint's influence; dropping the constraint returns you to the (infeasible) unconstrained minimum at the origin.
Why D tempts people
λ is generally nonzero for a binding constraint — here λ = −1. λ = 0 would mean the constraint exerts no force.

90. Answer it before you see the options: Check yourself — the shadow price

Prediction

Predict first

For min x²+y² s.t. x+y=c, the multiplier is λ=−1 at c=1. If you loosen the constraint to c=1.01, the optimal value f⋆ changes by approximately:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: +0.01, because df⋆/dc = −λ = 1

Why: The envelope theorem gives df⋆/dc = −λ = 1, so a step of Δc = 0.01 raises f⋆ by about 0.01. Numerically f⋆ went 0.50000 → 0.51005, a rise of ≈ 0.01.

91. Check yourself — the shadow price

Check

Recall f⋆(c) for the constraint x+y=c.

Check your understanding

For min x²+y² s.t. x+y=c, the multiplier is λ=−1 at c=1. If you loosen the constraint to c=1.01, the optimal value f⋆ changes by approximately:

  • A. +0.01, because df⋆/dc = −λ = 1 (correct)
  • B. −0.01, because loosening a constraint always lowers the cost
  • C. 0, because λ does not affect the optimal value
  • D. +1, because the shadow price is added directly to f⋆

Answer: A

Why: The envelope theorem gives df⋆/dc = −λ = 1, so a step of Δc = 0.01 raises f⋆ by about 0.01. Numerically f⋆ went 0.50000 → 0.51005, a rise of ≈ 0.01.

Why B tempts people
Loosening does not always lower cost — here the feasible line moves farther from the origin, so the minimum squared distance RISES. Sign of the change is set by λ, not by a blanket rule.
Why C tempts people
λ is exactly the sensitivity of the optimal value to the constraint level — that is the whole point of the shadow-price interpretation.
Why D tempts people
The shadow price is a RATE (df⋆/dc), not an amount added to f⋆. Multiply it by the step Δc = 0.01 to get the ≈ 0.01 change.

92. How sure are you: Check yourself — complementary slackness

Commit first

Predict first

An inequality constraint gᵢ(x) ≤ 0 is strictly satisfied at the optimum, gᵢ(x⋆) < 0. Its KKT multiplier μᵢ is:

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: 0 — the constraint is inactive and exerts no force

Why: Complementary slackness μᵢ gᵢ = 0 with gᵢ < 0 forces μᵢ = 0. A constraint satisfied with room to spare pushes on nothing — delete it and the optimum is unchanged.

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

93. Check yourself — complementary slackness

Check

Active or inactive?

Check your understanding

An inequality constraint gᵢ(x) ≤ 0 is strictly satisfied at the optimum, gᵢ(x⋆) < 0. Its KKT multiplier μᵢ is:

  • A. 0 — the constraint is inactive and exerts no force (correct)
  • B. strictly positive
  • C. negative
  • D. equal to gᵢ(x⋆)

Answer: A

Why: Complementary slackness μᵢ gᵢ = 0 with gᵢ < 0 forces μᵢ = 0. A constraint satisfied with room to spare pushes on nothing — delete it and the optimum is unchanged.

Why B tempts people
A positive multiplier requires an ACTIVE constraint (gᵢ = 0). With gᵢ < 0 the product μᵢ gᵢ could not be zero unless μᵢ = 0.
Why C tempts people
Dual feasibility requires μᵢ ≥ 0; negative inequality multipliers are never allowed — a constraint can only push inward.
Why D tempts people
μᵢ and gᵢ are different objects (a price vs a slack). Slackness sets μᵢ = 0, not μᵢ = gᵢ.

94. Answer it before you see the options: Check yourself — strong duality

Prediction

Predict first

For our SVM the dual optimum equals the primal optimum. Which statement is correct in general?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Weak duality (dual ≤ primal) always holds; the gap is zero when the problem is convex and Slater's condition holds

Why: Weak duality g(λ) ≤ p⋆ is automatic (min-max ≥ max-min). The gap closes — strong duality — for convex problems with a strictly feasible point (Slater). The SVM is convex and strictly feasible on separable data, so its dual is exact.

95. Check yourself — strong duality

Check

Weak vs strong — mind the gap.

Check your understanding

For our SVM the dual optimum equals the primal optimum. Which statement is correct in general?

  • A. Weak duality (dual ≤ primal) always holds; the gap is zero when the problem is convex and Slater's condition holds (correct)
  • B. The dual always equals the primal, for every optimization problem
  • C. The dual is always an upper bound on the primal
  • D. Strong duality requires the problem to be non-convex

Answer: A

Why: Weak duality g(λ) ≤ p⋆ is automatic (min-max ≥ max-min). The gap closes — strong duality — for convex problems with a strictly feasible point (Slater). The SVM is convex and strictly feasible on separable data, so its dual is exact.

Why B tempts people
For non-convex problems a strictly positive duality gap is common; equality is not guaranteed. It holds here only because the SVM is convex and satisfies Slater.
Why C tempts people
The dual is a LOWER bound (dual ≤ primal), because swapping min and max can only decrease the value. It is never guaranteed to be an upper bound.
Why D tempts people
Backwards: strong duality is guaranteed for CONVEX problems (with Slater). Non-convexity is exactly what can OPEN a gap.

96. Rule out three: Check yourself — the SVM dual

Elimination

Eliminate the wrong options

In our 4-point SVM (neg at x=1,2; pos at x=4,5) the dual gives α = [0, 0.5, 0.5, 0]. Why do exactly x=2 and x=4 have α > 0?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Their margin constraints are active (they sit exactly on the margin), so slackness allows α > 0
  • B. They are the points farthest from the boundary
  • C. They are chosen at random by the QP solver
  • D. Every point in each class contributes equally, and rounding zeroed two of them

Survives elimination: A

Why: x=2 and x=4 have margin y(wx+b) = 1 exactly — their constraints are active — so complementary slackness α(margin−1)=0 permits α > 0. They are the support vectors. x=1 and x=5 have margin 2 (inactive), forcing α = 0.

97. Check yourself — the SVM dual

Check

Which points survive with α > 0?

Check your understanding

In our 4-point SVM (neg at x=1,2; pos at x=4,5) the dual gives α = [0, 0.5, 0.5, 0]. Why do exactly x=2 and x=4 have α > 0?

  • A. Their margin constraints are active (they sit exactly on the margin), so slackness allows α > 0 (correct)
  • B. They are the points farthest from the boundary
  • C. They are chosen at random by the QP solver
  • D. Every point in each class contributes equally, and rounding zeroed two of them

Answer: A

Why: x=2 and x=4 have margin y(wx+b) = 1 exactly — their constraints are active — so complementary slackness α(margin−1)=0 permits α > 0. They are the support vectors. x=1 and x=5 have margin 2 (inactive), forcing α = 0.

Why B tempts people
The opposite: x=1 and x=5 are the FARTHEST (margin 2), and they get α = 0. Support vectors are the CLOSEST opposing points, on the margin.
Why C tempts people
The nonzero set is geometrically determined by which constraints are active, not random — rerunning gives the same [0, 0.5, 0.5, 0].
Why D tempts people
The values are not equal or rounded: x=1 and x=5 have exact α = 0 because their constraints are strictly slack, which is genuine sparsity, not a rounding artifact.

98. Your turn: solve constrained

Section

The project

99. Project: from Lagrange to the SVM dual

Concept

Build the whole chain yourself: an equality problem by Lagrange and scipy, an inequality with an active constraint, then the 4-point SVM dual with its support vectors. You've derived every piece — now assemble it.

#requirementtool
1Equality: min x²+y² s.t. x+y=1Lagrange + minimize(eq)
2Inequality x≥1 → find the active constraintminimize(ineq)
3SVM dual → α and the support vectorsGram + minimize(SLSQP)

Build rules: type every line yourself, pass each scipy constraint as a dict ('type':'eq'/'ineq'), and after each solve, check which constraints are active at the answer.

100. By analogy: Project: from Lagrange to the SVM dual

Analogy

Discussion prompt

Explain Project: from Lagrange to the SVM dual by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Build rules: type every line yourself, pass each scipy constraint as a dict ('type':'eq'/'ineq'), and after each solve, check which constraints are active at the answer.

101. Milestone 1 — equality by Lagrange, then scipy

Worked example

Your turn: set ∇L=0 for f=x²+y² with x+y=1, predict (x,y), then confirm with scipy. Say the answer aloud before you print.

Hint: 2x+λ=0, 2y+λ=0 ⇒ x=y; substitute into x+y=1. For scipy use constraints={'type':'eq','fun':lambda v: v[0]+v[1]-1}.

from scipy.optimize import minimize
# by hand: 2x=-lam, 2y=-lam -> x=y ; x+y=1 -> x=y=0.5, f=0.5, lam=-1
res = minimize(lambda v: v[0]**2 + v[1]**2, [0., 0.],
               constraints={'type': 'eq', 'fun': lambda v: v[0] + v[1] - 1})
print(res.x.round(4), round(res.fun, 4))
sourcex⋆, y⋆f⋆λ
by hand[0.5, 0.5]0.5−1
scipy[0.5, 0.5]0.5—

102. Milestone 2 — an active inequality

Worked example

Your turn: minimize x²+y² subject to x ≥ 1. Predict where the solution lands and whether the constraint binds, then verify.

Hint: scipy's 'ineq' means fun ≥ 0, so pass x−1. Test activeness by checking whether x⋆−1 == 0 at the solution.

from scipy.optimize import minimize
res = minimize(lambda v: v[0]**2 + v[1]**2, [2., 2.],
               constraints={'type': 'ineq', 'fun': lambda v: v[0] - 1})
print(res.x.round(4), round(res.fun, 4))
print('active:', round(res.x[0] - 1, 4) == 0.0)
quantityvalue
x⋆, y⋆[1, 0]
f⋆1.0
active (x=1)?True

103. Watch it run: Milestone 2 — an active inequality

Pattern

Step through it

Step through Milestone 2 — an active inequality one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is x⋆, y⋆
  2. Step 2: quantity is f⋆
  3. Step 3: quantity is active (x=1)?

104. Milestone 3 — the SVM dual and its support vectors

Worked example

Your turn: solve the dual QP for the 4 points, then recover w, b, and read off which points have α > 0. Predict the support vectors first.

Hint: G = np.outer(y,y)*np.outer(x,x); minimize 0.5*a@G@a - a.sum() with bounds=[(0,None)]*4 and equality a@y==0. Then w=(a*y)@x and b=y[k]-w*x[k] for any support-vector index k.

import numpy as np
from scipy.optimize import minimize
x = np.array([1., 2., 4., 5.])
y = np.array([-1., -1., 1., 1.])
G = np.outer(y, y) * np.outer(x, x)
obj = lambda a: 0.5 * a @ G @ a - a.sum()
cons = ({'type': 'eq', 'fun': lambda a: a @ y},)
a = minimize(obj, np.zeros(4), bounds=[(0, None)]*4,
             constraints=cons, method='SLSQP').x
w = (a * y) @ x
k = a.argmax()
b = y[k] - w * x[k]
print('alpha =', a.round(4), ' w =', round(w, 4), ' b =', round(b, 4))
outputvalue
alpha[0. 0.5 0.5 0.]
support vectors (x)2 and 4
w, b1.0, -3.0

105. What each one costs: Milestone 3 — the SVM dual and its support vectors

Trade off

Comparison matrix

From Milestone 3 — the SVM dual and its support vectors: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

outputvalue
alpha[0. 0.5 0.5 0.]
support vectors (x)2 and 4
w, b1.0, -3.0

106. The full program

Concept

import numpy as np
from scipy.optimize import minimize

# 1. equality  min x^2+y^2 s.t. x+y=1  ->  (0.5, 0.5)
eq = minimize(lambda v: v[0]**2+v[1]**2, [0.,0.],
              constraints={'type':'eq','fun':lambda v: v[0]+v[1]-1})

# 2. inequality  min x^2+y^2 s.t. x>=1  ->  (1, 0), active
ineq = minimize(lambda v: v[0]**2+v[1]**2, [2.,2.],
                constraints={'type':'ineq','fun':lambda v: v[0]-1})

# 3. SVM dual  ->  alpha=[0,0.5,0.5,0], w=1, b=-3
x = np.array([1.,2.,4.,5.]); y = np.array([-1.,-1.,1.,1.])
G = np.outer(y,y)*np.outer(x,x)
a = minimize(lambda a: 0.5*a@G@a - a.sum(), np.zeros(4),
             bounds=[(0,None)]*4,
             constraints=({'type':'eq','fun':lambda a: a@y},)).x
w = (a*y)@x; b = y[a.argmax()] - w*x[a.argmax()]
print('equality  :', eq.x.round(3),  round(eq.fun,3))
print('inequality:', ineq.x.round(3), round(ineq.fun,3))
print('svm alpha :', a.round(3), ' w,b =', round(w,3), round(b,3))
printed linevalue (verified)
equality :[0.5 0.5] 0.5
inequality:[1. -0.] 1.0
svm alpha :[0. 0.5 0.5 0.] w,b = 1.0 -3.0

If the equality lands at (0.5, 0.5), the inequality binds at (1, 0), and the SVM gives α = [0, 0.5, 0.5, 0] with w=1, b=−3 — you can read the KKT conditions off a real SVM.

107. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
equality :[0.5 0.5] 0.5
inequality:[1. -0.] 1.0
svm alpha :[0. 0.5 0.5 0.] w,b = 1.0 -3.0

108. Show it off

Concept

Slides closed, out loud: explain (1) why ∇f ∥ ∇h at an equality optimum, (2) what complementary slackness says about active vs inactive constraints, and (3) why only support vectors have α > 0.

Stretch (homework): use Lagrange multipliers to show the max-entropy distribution with fixed mean and variance is the Gaussian, and re-derive the SVM dual from the primal Lagrangian yourself. KKT returns for kernels and soft margins (Week 25) and for RLHF's KL constraint (Week 40).

109. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why ∇f ∥ ∇h at an equality optimum, (2) what complementary slackness says about active vs inactive constraints, and (3) why only support vectors have α > 0.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

110. Connect it up: Lesson 24: Constrained Optimization & KKT

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The problem: rules to obey · The Lagrangian · Inequalities & KKT · The SVM, primal and dual · Duality & Slater · Recipe, checks & your turn. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

111. What you can do now

Recap

ideathe one thing to remember
Lagrangian∇f = −λ∇h at the optimum (aligned gradients)
multiplierλ is a shadow price: ∂f⋆/∂c = −λ
KKTstationarity + primal/dual feasibility + slackness
slacknessμᵢ gᵢ = 0: active ⇒ μ>0, inactive ⇒ μ=0
SVM dualQP in α≥0; support vectors are the active constraints
dualityconvex + Slater ⇒ strong duality, gap = 0

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 24 (Week 8 — Lagrangian Optimization) — Barron · USAAIO Round 2 Preparation, 2026
  2. Boyd & Vandenberghe, Convex Optimization, Ch. 5 (Duality) & KKT conditions
  3. scipy.optimize.minimize (SLSQP) constraint dicts
  4. scikit-learn SVC (linear kernel) — support vectors & dual coefficients
  5. Every multiplier, alpha, and margin produced by real execution — numpy 2.2.6 + scipy 1.16 + scikit-learn 1.9, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108