Lesson 11: Convex Optimization

USAAIO Lesson 11, from Week 4 on optimization, fully worked. It proves the chord definition of convexity on a running quadratic, then derives the first-order tangent condition and the second-order condition that the Hessian be positive semi-definite, step by step, and dissects a non-convex counterexample. It shows why every stationary point of a convex function is global, writes gradient descent as a contraction with the update factor derived from scratch, proves the stable-step bound eta < 2/L, and covers the speed limit set by the condition number along with the classes of convergence rate. It ends with a from-scratch gradient-descent experiment verified against real execution. Every number was produced by running the code. The lesson runs to 61 slides.

Subject: Machine Learning · 114 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Convex Optimization

Title

USAAIO · Lesson 11 · Week 4 (Optimization)

Why some loss surfaces are easy and others aren't. We prove convexity three equivalent ways, show why a convex bowl has no traps, then derive gradient descent as a contraction and the exact step-size bound η < 2/L — every algebra move shown, every number produced by running the code.

2. By the end of this lesson you can

Objectives

  1. State the chord definition of a convex function and verify it numerically on a quadratic
  2. Move between the three equivalent tests — chord, first-order tangent, second-order Hessian-PSD — and know which to reach for
  3. Dissect a non-convex function g(t)=t⁴−3t² and locate exactly where the Hessian goes negative
  4. Prove that for a convex f, ∇f(x)=0 forces a global minimum — no local traps
  5. Derive the gradient-descent update as a contraction and read off the stable-step bound η < 2/L
  6. Relate the condition number κ = L/μ to convergence speed and quote the O(1/t) / linear / O(1/t²) rate classes
  7. Implement a GD learning-rate sweep in NumPy and classify each run as converging or diverging against real output

3. What survived from Object Detection & YOLO?

Warm-up

Discussion prompt

Before we open Lesson 11: Convex Optimization: without looking back, what was the main idea of Object Detection & YOLO, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

bounding-box parameterisation (corner vs centre-offset formats), IoU from scratch, non-maximum suppression (NMS), YOLO grid/anchor math (S=13, B=5, C=80 → 17 745 raw predictions), anchor-offset decoding with sigmoid/exp, Feature Pyramid Networks (FPN) for multi-scale detection, and mAP@0.5 calculation.

4. What makes a problem convex

Section

Part 1 of 6 — the definition

5. Our running example: a quadratic bowl

Concept

We carry one function through the whole lesson so every slide compounds. It is a 2-variable quadratic in w = (w₁, w₂):

\[ f(w) = \tfrac{1}{2}\big(w_1^2 + 4\,w_2^2\big) \]

Its lowest point is clearly w = (0, 0), where f = 0. The 4 on w₂ makes the bowl steeper in the w₂ direction than the w₁ direction — that asymmetry becomes the whole story of Part 5.

For the pure learning-rate arithmetic we also keep a 1-D shadow, f(x) = x², because its gradient-descent step is a single clean multiplication.

6. Break it if you can: Our running example: a quadratic bowl

Counterexample

Discussion prompt

We carry one function through the whole lesson so every slide compounds. It is a 2-variable quadratic in w = (w₁, w₂):

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Its lowest point is clearly w = (0, 0), where f = 0. The 4 on w₂ makes the bowl steeper in the w₂ direction than the w₁ direction — that asymmetry becomes the whole story of Part 5.

7. Convex = one bowl, no dents

Intuition

Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.

If you stretch a string between any two points on the surface, the string never dips below the surface — it stays on or above it. That taut-string picture is exactly the formal definition we write next.

Why obsess over this? On a convex bowl, walking downhill always reaches the one true bottom. On a dented surface, downhill can dump you into the wrong valley — the difference between 'training just works' and 'training is an art'.

8. By analogy: Convex = one bowl, no dents

Analogy

Discussion prompt

Explain Convex = one bowl, no dents by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.

9. The chord definition

Concept

Take any two inputs x and y. The straight line between the two surface points (x, f(x)) and (y, f(y)) is the chord. Convexity says the function sits on or below its chords:

\[ f\big(\alpha x + (1-\alpha)y\big) \;\le\; \alpha f(x) + (1-\alpha)f(y), \qquad \alpha \in [0,1] \]

convex function — A function whose value at any weighted average of two points is at most the same weighted average of the two function values. Geometrically: the graph never rises above a chord. α sweeps the chord: α=1 is the x end, α=0 is the y end, α=½ is the midpoint.

10. Guess the shape of the answer: Verify the chord inequality on our bowl

Estimation

Predict first

Claim: f(w)=½(w₁²+4w₂²) is convex. Let's not take it on faith — pick x=(4,1) and y=(−2,3), sweep α, and check the chord inequality numerically. Runnable as written:

Commit before you compute: what does Verify the chord inequality on our bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: At every α, f(mix) ≤ chord — the inequality holds

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The surface height never exceeds the chord height.

11. Verify the chord inequality on our bowl

Worked example

Claim: f(w)=½(w₁²+4w₂²) is convex. Let's not take it on faith — pick x=(4,1) and y=(−2,3), sweep α, and check the chord inequality numerically. Runnable as written:

import numpy as np
def f(w): return 0.5*(w[0]**2 + 4*w[1]**2)
x = np.array([4., 1.]); y = np.array([-2., 3.])
for a in [0., 0.25, 0.5, 0.75, 1.]:
    mix   = a*x + (1-a)*y          # point on the segment
    chord = a*f(x) + (1-a)*f(y)    # height of the chord
    print(a, round(f(mix),4), round(chord,4), f(mix) <= chord)

At every α, f(mix) ≤ chord — the inequality holds

Why: The surface height never exceeds the chord height. A single violation would disprove convexity; none appears, consistent with (but not yet a proof of) convexity.

αf(mix)chord heightf(mix) ≤ chord?
0.0020.000020.0000True (y endpoint)
0.2512.625017.5000True
0.508.500015.0000True
0.757.625012.5000True
1.0010.000010.0000True (x endpoint)

12. The picture: surface below its chord

Concept

Here is the chord condition drawn for a 1-D convex slice. The straight chord connects two surface points; the curve sags below it. That sag is convexity.

Figure (svg): A U-shaped convex curve with a straight chord drawn between two points on it; the curve lies entirely below the chord, and a vertical gap at the midpoint is labelled.

Convex: the curve never rises above the chord. The midpoint gap is exactly f(mix) ≤ chord.

Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.

13. Teach it back: The picture: surface below its chord

Explain it

Discussion prompt

Explain The picture: surface below its chord to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.

14. One example can't prove it — we need a certificate

Intuition

The table checked one pair of points at five values of α. Convexity demands the inequality for every pair and every α — infinitely many checks. No finite table can settle it.

So we need a certificate: a condition we can check once, using derivatives, that guarantees the chord inequality for all pairs at once. That is what the first- and second-order conditions give us.

15. First-order condition: tangents lie below

Concept

For a differentiable f, convexity is equivalent to: the tangent line (plane) at any point stays below the function everywhere. The linear approximation always underestimates.

\[ f(y) \;\ge\; f(x) + \nabla f(x)^\top (y - x) \qquad \forall\, x, y \]

Read it as: the true value f(y) is at least what the tangent at x predicts. This single inequality is the engine behind 'no local traps' — we cash it in shortly.

16. What has to be given first: Verify the tangent lies below

Missing information

Discussion prompt

Test the first-order condition on our bowl. Anchor the tangent at x=(2,1), where f(x)=4 and ∇f(x)=[2,4], then check f(y) ≥ tangent(y) at several y. Runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The tangent anchored at x underestimates f everywhere, touching only at y=x (both read 4). That is the first-order condition holding, and it is what forces a zero-gradient point to be global. Verified by execution.

17. Verify the tangent lies below

Worked example

Test the first-order condition on our bowl. Anchor the tangent at x=(2,1), where f(x)=4 and ∇f(x)=[2,4], then check f(y) ≥ tangent(y) at several y. Runnable:

import numpy as np
def f(w):    return 0.5*(w[0]**2 + 4*w[1]**2)
def grad(w): return np.array([w[0], 4*w[1]])
x = np.array([2., 1.])
for y in [[0.,0.], [3.,2.], [-1.,1.], [2.,1.]]:
    y = np.array(y)
    tang = f(x) + grad(x) @ (y - x)     # tangent-plane height
    print(y.tolist(), round(f(y),3), round(tang,3), f(y) >= tang - 1e-9)

At every y, f(y) ≥ tangent — the plane stays under the surface

Why: The tangent anchored at x underestimates f everywhere, touching only at y=x (both read 4). That is the first-order condition holding, and it is what forces a zero-gradient point to be global. Verified by execution.

yf(y)tangent at x=(2,1)f(y) ≥ tangent?
[0, 0]0.0−4.0True
[3, 2]12.510.0True
[−1, 1]2.5−2.0True
[2, 1]4.04.0True (touches at y=x)

18. Which is which, by f(y) ≥ tangent?

Discrimination

Sort into buckets

Sort these by f(y) ≥ tangent?, from memory, without looking back at Verify the tangent lies below. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

True
[0, 0]; [3, 2]; [−1, 1]
True (touches at y=x)
[2, 1]
g1
f(y) ≥ tangent? is "True" for [0, 0], [3, 2], [−1, 1] — that is what the table on "Verify the tangent lies below" records, and it is the single property separating this group from the rest.
g2
f(y) ≥ tangent? is "True (touches at y=x)" for [2, 1] — that is what the table on "Verify the tangent lies below" records, and it is the single property separating this group from the rest.

19. Second-order condition: Hessian PSD everywhere

Concept

For a twice-differentiable f, convexity is equivalent to the Hessian (the matrix of second partials) being positive semidefinite at every point of the domain:

\[ f \text{ convex} \iff \nabla^2 f(x) \succeq 0 \quad \forall\, x \]

positive semidefinite (PSD) — A symmetric matrix H with zᵀHz ≥ 0 for every vector z — equivalently, all eigenvalues ≥ 0. It means non-negative curvature in every direction: the surface never bends downward. This is the easiest of the three tests to check, since it is just an eigenvalue sign check.

20. Take the definitions apart: convex function vs positive semidefinite…

Definition probe

Sort into buckets

Every line below is part of the definition of convex function or of positive semidefinite (PSD) — one or the other, never both. Put each where it belongs.

convex function
A function whose value at any weighted average of two points is at most the same weighted average of the two function values.; the graph never rises above a chord.; α=1 is the x end, α=0 is the y end, α=½ is the midpoint.
positive semidefinite (PSD)
A symmetric matrix H with zᵀHz ≥ 0 for every vector z; equivalently, all eigenvalues ≥ 0.; It means non-negative curvature in every direction
b1
A function whose value at any weighted average of two points is at most the same weighted average of the two function values. Geometrically: the graph never rises above a chord. α sweeps the chord: α=1 is the x end, α=0 is the y end, α=½ is the midpoint.
b2
A symmetric matrix H with zᵀHz ≥ 0 for every vector z — equivalently, all eigenvalues ≥ 0. It means non-negative curvature in every direction: the surface never bends downward. This is the easiest of the three tests to check, since it is just an eigenvalue sign check.

21. What has to happen first: Compute the Hessian of our bowl by hand

Ranking

Put in order

Put the moves of Compute the Hessian of our bowl by hand into the order they have to happen.

  1. First partials
  2. Second partials
  3. Verify: eigenvalues are 1 and 4, both ≥ 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. ∂f/∂w₁ = ½·2w₁ = w₁ and ∂f/∂w₂ = ½·8w₂ = 4w₂.

22. Compute the Hessian of our bowl by hand

Worked example

The Hessian collects all second partial derivatives. For f(w)=½(w₁²+4w₂²), take the partials one at a time — nothing skipped.

First partials

Why: ∂f/∂w₁ = ½·2w₁ = w₁ and ∂f/∂w₂ = ½·8w₂ = 4w₂. So ∇f = [w₁, 4w₂].

\[ \nabla f(w) = \begin{bmatrix} w_1 \\ 4 w_2 \end{bmatrix} \]

Second partials

Why: ∂²f/∂w₁² = 1, ∂²f/∂w₂² = 4, and the mixed partial ∂²f/∂w₁∂w₂ = 0 since w₁ and w₂ never multiply. The Hessian is constant — it does not depend on w.

\[ \nabla^2 f = \begin{bmatrix} 1 & 0 \\ 0 & 4 \end{bmatrix} \]

Verify: eigenvalues are 1 and 4, both ≥ 0

Why: A diagonal matrix's eigenvalues are its diagonal entries. Both positive ⇒ PSD (in fact positive definite) at every w ⇒ f is convex everywhere. Certificate obtained.

23. Decode the notation: Compute the Hessian of our bowl by hand

Notation

Annotate

From Compute the Hessian of our bowl by hand — read this one piece at a time. What is each part doing?

On: \( \nabla f(w) = \begin{bmatrix} w_1 \\ 4 w_2 \end{bmatrix} \)

  • ∂f/∂w₁ = ½·2w₁ = w₁ and ∂f/∂w₂ = ½·8w₂ = 4w₂. So ∇f = [w₁, 4w₂].
  • ∂²f/∂w₁² = 1, ∂²f/∂w₂² = 4, and the mixed partial ∂²f/∂w₁∂w₂ = 0 since w₁ and w₂ never multiply. The Hessian is constant — it does not depend on w.
  • A diagonal matrix's eigenvalues are its diagonal entries. Both positive ⇒ PSD (in fact positive definite) at every w ⇒ f is convex everywhere. Certificate obtained.

24. Predict the next row: Confirm the Hessian test in code

Pattern

Predict first

The table runs: Hessian ∇²f | [[1, 0], [0, 4]] (constant in w) · eigenvalues | [1.0, 4.0]

In Confirm the Hessian test in code, given the rows so far: what is the next one — the row where quantity is all ≥ 0??

Correct: all ≥ 0? | True → PSD → convex everywhere

quantityvalue (verified)
Hessian ∇²f[[1, 0], [0, 4]] (constant in w)
eigenvalues[1.0, 4.0]
all ≥ 0?True → PSD → convex everywhere

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The Hessian is constant, so this one eigenvalue check certifies convexity over the ENTIRE domain — no need to re-check at other points.

25. Confirm the Hessian test in code

Worked example

Ask NumPy for the eigenvalues of the Hessian and read off the sign. eigvalsh is the symmetric-matrix eigenvalue routine — the right tool for a Hessian.

import numpy as np
H = np.array([[1., 0.],
              [0., 4.]])       # Hessian of 0.5(w1^2 + 4 w2^2)
eig = np.linalg.eigvalsh(H)
print(eig)                     # [1. 4.]
print('convex everywhere?', np.all(eig >= 0))

eig(H) = [1, 4], all ≥ 0 → PSD → convex

Why: The Hessian is constant, so this one eigenvalue check certifies convexity over the ENTIRE domain — no need to re-check at other points. Verified by execution.

quantityvalue (verified)
Hessian ∇²f[[1, 0], [0, 4]] (constant in w)
eigenvalues[1.0, 4.0]
all ≥ 0?True → PSD → convex everywhere

26. Fill in: value (verified) for Confirm the Hessian test in code

Comparison

Comparison matrix

From Confirm the Hessian test in code: refill the value (verified) column from what you know. The rest of the table is as it appeared.

quantityvalue (verified)
Hessian ∇²f[[1, 0], [0, 4]] (constant in w)
eigenvalues[1.0, 4.0]
all ≥ 0?True → PSD → convex everywhere

27. Something is wrong here: checking convexity only at the minimum

Anomaly

Predict first

A student writes this, and it looks reasonable:

To test convexity, find the critical point and check the Hessian there. If ∇²f ⪰ 0 at that point, call f convex.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: A PSD Hessian at ONE point only certifies a LOCAL minimum.

Convexity is a global property. The Hessian must be PSD at every point in the domain, not just at the minimum.

Why: A PSD Hessian at ONE point only certifies a LOCAL minimum. The function is free to curve downward somewhere else and still pass this local test — so you can wrongly certify a non-convex function.

28. Trap: checking convexity only at the minimum

Trap

The trap

To test convexity, find the critical point and check the Hessian there. If ∇²f ⪰ 0 at that point, call f convex.

Test ∇²f ⪰ 0 only at the stationary point

Why: A PSD Hessian at ONE point only certifies a LOCAL minimum. The function is free to curve downward somewhere else and still pass this local test — so you can wrongly certify a non-convex function.

The fix

Convexity is a global property. The Hessian must be PSD at every point in the domain, not just at the minimum.

Verify ∇²f ⪰ 0 for ALL w

Why: For our bowl the Hessian is CONSTANT (= diag(1,4)), so one check covers all w. When the Hessian depends on w you must show it stays PSD throughout — as the next slide's counterexample makes painfully clear.

29. When convexity fails

Section

Part 2 of 6 — a counterexample

30. A function that is not convex

Concept

To feel what convexity buys us, break it. Consider the 1-D function

\[ g(t) = t^4 - 3t^2 \]

It has a hump in the middle and two dips on the sides — a classic double well. Downhill from the left leads to one valley, downhill from the right to another. That is exactly the trap a convex function forbids.

31. Plan first: Its second derivative changes sign

Step zero

Discussion prompt

Its second derivative changes sign — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: First derivative

Answer:

  1. First derivative
  2. Second derivative
  3. Find where g'' < 0

32. Its second derivative changes sign

Worked example

Differentiate g(t)=t⁴−3t² twice, one move per line:

First derivative

Why: g'(t) = 4t³ − 6t by the power rule, term by term.

\[ g'(t) = 4t^3 - 6t \]

Second derivative

Why: g''(t) = 12t² − 6, again by the power rule. This is the 1-D Hessian.

\[ g''(t) = 12t^2 - 6 \]

Find where g'' < 0

Why: 12t² − 6 < 0 ⇔ t² < ½ ⇔ |t| < 1/√2 ≈ 0.707. On that whole interval the curvature is NEGATIVE — the graph bends downward — so g is not convex.

\[ g''(t) < 0 \iff -\tfrac{1}{\sqrt2} < t < \tfrac{1}{\sqrt2} \]

33. Draw the shape of it: Its second derivative changes sign

Blank canvas

Draw it

Draw what Its second derivative changes sign just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

34. Trace the curvature and catch the chord violation

Worked example

Tabulate g'' across the domain, then test the chord between t=−1.2 and t=+1.2 at the midpoint t=0. Runnable:

import numpy as np
def g(t):   return t**4 - 3*t**2
def gpp(t): return 12*t**2 - 6      # second derivative
for t in [-1.5, -1., -0.5, 0., 0.5, 1., 1.5]:
    print(t, gpp(t), 'convex' if gpp(t) >= 0 else 'CONCAVE')
mid = g(0.0); chord = 0.5*(g(-1.2) + g(1.2))
print('g(0) =', mid, ' chord =', round(chord,4), ' convex?', mid <= chord)

g''< 0 on the middle rows; the chord test fails at t=0

Why: g(0) = 0 but the chord between t=±1.2 sits at −2.2464, which is BELOW 0. The surface rises above its chord — the definitional violation, matching the negative-curvature interval exactly.

tg''(t) = 12t²−6curvatureverdict
−1.521.0positiveconvex here
−1.06.0positiveconvex here
−0.5−3.0negativeCONCAVE
0.0−6.0negativeCONCAVE
0.5−3.0negativeCONCAVE
1.06.0positiveconvex here
1.521.0positiveconvex here

35. Which is which, by curvature

Discrimination

Sort into buckets

Sort these by curvature, from memory, without looking back at Trace the curvature and catch the chord violation. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

positive
−1.5; −1.0; 1.0; 1.5
negative
−0.5; 0.0; 0.5
g1
curvature is "positive" for −1.5, −1.0, 1.0, 1.5 — that is what the table on "Trace the curvature and catch the chord…" records, and it is the single property separating this group from the rest.
g2
curvature is "negative" for −0.5, 0.0, 0.5 — that is what the table on "Trace the curvature and catch the chord…" records, and it is the single property separating this group from the rest.

36. Two minima — the trap made concrete

Concept

Set g'(t)=4t³−6t = 2t(2t²−3) = 0. The stationary points are t=0 and t=±√(3/2) ≈ ±1.2247.

stationary tg(t)g''(t)type
−1.2247−2.2512 > 0local (global) minimum
00−6 < 0local MAXIMUM (the hump)
+1.2247−2.2512 > 0local (global) minimum

Two separate minima. Gradient descent started at t=−0.3 slides left; started at t=+0.3 it slides right. Where you start decides where you land — the pathology convexity rules out.

37. What each one costs: Two minima — the trap made concrete

Trade off

Comparison matrix

From Two minima — the trap made concrete: every row here is a choice with a cost. Fill the type column, then say which row you would actually pick and what you give up for it.

stationary tg(t)g''(t)type
−1.2247−2.2512 > 0local (global) minimum
00−6 < 0local MAXIMUM (the hump)
+1.2247−2.2512 > 0local (global) minimum

38. Why convex ⇒ only global minima

Section

Part 3 of 6 — the payoff

39. The one theorem that pays for everything

Intuition

Everything so far — chords, tangents, Hessians — was setup. Here is the payoff that makes convex optimization the well-behaved corner of machine learning:

On a convex function, a flat point is the best point. Find any w where the gradient is zero and you are done — it is the global minimum, not merely a local one. No restarts, no luck, no wondering if a better valley exists.

The next slide proves it in two lines using only the first-order (tangent-below) condition — the whole reason we bothered to state it.

40. Plan first: Prove: ∇f(x)=0 forces a global minimum

Step zero

Discussion prompt

Prove: ∇f(x)=0 forces a global minimum — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Start from the first-order condition

Answer:

  1. Start from the first-order condition
  2. Substitute ∇f(x) = 0
  3. Verify the conclusion holds for every y

41. Prove: ∇f(x)=0 forces a global minimum

Worked example

This is the theorem that makes convex problems easy. Suppose f is convex and differentiable and x is a stationary point, ∇f(x)=0. We show f(x) is the global minimum in two lines.

Start from the first-order condition

Why: Convexity gives the tangent-below inequality at x, valid for EVERY y.

\[ f(y) \ge f(x) + \nabla f(x)^\top (y - x) \]

Substitute ∇f(x) = 0

Why: The stationary condition kills the entire linear term, since 0ᵀ(y−x) = 0.

\[ f(y) \ge f(x) + 0 = f(x) \]

Verify the conclusion holds for every y

Why: f(y) ≥ f(x) for ALL y means x is a GLOBAL minimum. No 'other valley' can be lower — the tangent inequality forbids it. This is why a convex problem has no local traps.

\[ \boxed{\,f(y) \ge f(x)\ \ \forall y \;\Longrightarrow\; x \text{ is a global minimizer}\,} \]

42. Work backwards from the answer: Prove: ∇f(x)=0 forces a global minimum

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Verify the conclusion holds for every y

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

This is the theorem that makes convex problems easy. Suppose f is convex and differentiable and x is a stationary point, ∇f(x)=0. We show f(x) is the global minimum in two lines.

43. What the proof means for training

Intuition

For a convex loss, any point where the gradient vanishes is already the best possible fit. You never have to worry 'is this just a local minimum?' — there is only one.

Linear regression, ridge, logistic regression, SVMs — all convex. That is why we can trust gradient descent to find the answer, and why Lesson 7's normal equations gave the unique global optimum.

Neural nets are not convex. Their loss is a landscape of many valleys, so initialization, momentum, and luck all matter. Convexity is the dividing line between 'solved' and 'an art'.

44. Where does each piece belong: Lesson 11: Convex Optimization

Sorting

Sort into buckets

These are the pieces of Lesson 11: Convex Optimization, out of order. Put each one back under the part of the lesson it belongs to.

What makes a problem convex
Our running example: a quadratic bowl; Convex = one bowl, no dents; The chord definition
When convexity fails
A function that is not convex; Its second derivative changes sign; Trace the curvature and catch the chord violation
Why convex ⇒ only global minima
The one theorem that pays for everything; Prove: ∇f(x)=0 forces a global minimum; What the proof means for training
s1
What makes a problem convex is where Lesson 11: Convex Optimization puts Our running example: a quadratic bowl, Convex = one bowl, no dents, The chord definition. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
When convexity fails is where Lesson 11: Convex Optimization puts A function that is not convex, Its second derivative changes sign, Trace the curvature and catch the chord violation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Why convex ⇒ only global minima is where Lesson 11: Convex Optimization puts The one theorem that pays for everything, Prove: ∇f(x)=0 forces a global minimum, What the proof means for training. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

45. Something is wrong here: 'convex functions can still get stuck in local minima'

Anomaly

Predict first

A student writes this, and it looks reasonable:

Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This treats a convex bowl like a neural-net landscape.

A convex function has no local minimum other than the global one. Convergence is about speed, not about getting stuck.

Why: This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.

46. Trap: 'convex functions can still get stuck in local minima'

Trap

The trap

Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.

Restart MSE from many seeds to escape local minima

Why: This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.

The fix

A convex function has no local minimum other than the global one. Convergence is about speed, not about getting stuck.

Every stationary point of a convex f is global

Why: We just proved it: ∇f(x)=0 ⇒ f(y) ≥ f(x) for all y. So the only thing left to tune is the learning rate — which controls how FAST you reach the one minimum, the subject of Parts 4–5.

47. Break it on purpose: 'convex functions can still get stuck in…

Break the constraint

Discussion prompt

The rule this trap just fixed:

We just proved it: ∇f(x)=0 ⇒ f(y) ≥ f(x) for all y. So the only thing left to tune is the learning rate — which controls how FAST you reach the one minimum, the subject of Parts 4–5.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.

48. Restore the missing line: The MSE loss is convex — tying back to Lesson 7

Fill the middle

Fill in the blanks

From The MSE loss is convex — tying back to Lesson 7 — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1.,2.,3.,4.,5.]); y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
H = **2 * X.T @ X** # Hessian of ||Xw - y||^2
print(H) # [[ 10. 30.] [ 30. 110.]]
print(np.linalg.eigvalsh(H)) # [ 1.690481 118.309519]

Why: H is what everything below it consumes, so the wrong expression here fails later and somewhere else. zᵀ(2XᵀX)z = 2‖Xz‖² ≥ 0 for any z, so 2XᵀX is PSD by construction — and it is CONSTANT in w.

49. The MSE loss is convex — tying back to Lesson 7

Worked example

The least-squares loss L(w)=‖Xw−y‖² from Lesson 7 is convex, and now we can prove it in one line: its Hessian is 2XᵀX, which is PSD. On the 5-student data:

import numpy as np
x = np.array([1.,2.,3.,4.,5.]); y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
H = 2 * X.T @ X               # Hessian of ||Xw - y||^2
print(H)                      # [[ 10.  30.] [ 30. 110.]]
print(np.linalg.eigvalsh(H))  # [  1.690481 118.309519]

eig(2XᵀX) = [1.69, 118.31], both ≥ 0 → convex everywhere

Why: zᵀ(2XᵀX)z = 2‖Xz‖² ≥ 0 for any z, so 2XᵀX is PSD by construction — and it is CONSTANT in w. Convex everywhere ⇒ the normal-equation solution is the unique GLOBAL optimum. Verified by execution.

quantityvalue (verified)
Hessian2XᵀX = [[10, 30], [30, 110]] (constant)
eigenvalues[1.6905, 118.3095]
convex?yes — both ≥ 0, PSD everywhere

50. Strict convexity ⇒ a unique minimizer

Concept

Convex (∇²f ⪰ 0, eigenvalues ≥ 0) guarantees any minimum is global, but there could be a flat valley of equally-good minima. Strict convexity (∇²f ≻ 0, eigenvalues > 0) rules that out.

\[ \nabla^2 f \succ 0 \;\Longrightarrow\; \text{the global minimizer is unique} \]

Our bowl (eigenvalues 1, 4 > 0) and the full-rank MSE Hessian (1.69, 118.3 > 0) are both strictly convex — one point, not a valley. A rank-deficient X gives a zero eigenvalue: still convex, but with infinitely many optimal w — exactly the collinearity case ridge fixed in Lesson 7.

51. What has to happen first: Prove 2XᵀX is PSD without eigenvalues

Ranking

Put in order

Put the moves of Prove 2XᵀX is PSD without eigenvalues into the order they have to happen.

  1. Write out zᵀ(2XᵀX)z
  2. Recognize a squared norm
  3. Verify against the numbers

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Pull the constant 2 out front and group the middle: zᵀXᵀXz = (Xz)ᵀ(Xz).

52. Prove 2XᵀX is PSD without eigenvalues

Worked example

The eigenvalue check is numeric; here is the airtight algebraic reason 2XᵀX ⪰ 0 for any X. Test the PSD definition zᵀHz ≥ 0 directly.

Write out zᵀ(2XᵀX)z

Why: Pull the constant 2 out front and group the middle: zᵀXᵀXz = (Xz)ᵀ(Xz).

\[ z^\top (2 X^\top X)\, z = 2\,(Xz)^\top (Xz) \]

Recognize a squared norm

Why: (Xz)ᵀ(Xz) is the dot product of the vector Xz with itself — that is ‖Xz‖², which can never be negative.

\[ = 2\,\lVert Xz \rVert^2 \;\ge\; 0 \quad \forall z \]

Verify against the numbers

Why: zᵀHz ≥ 0 for all z is precisely the PSD definition — so 2XᵀX is convex-certifying for any data X, matching the eigenvalues [1.69, 118.31] ≥ 0 we computed. No noise, no exceptions.

stepexpressionsign
groupzᵀ(2XᵀX)z = 2(Xz)ᵀ(Xz)regrouped
norm= 2‖Xz‖²square
conclude≥ 0 for every zPSD ✓

53. Say it in words: Prove 2XᵀX is PSD without eigenvalues

Translation

\( = 2\,\lVert Xz \rVert^2 \;\ge\; 0 \quad \forall z \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

54. Gradient descent as a contraction

Section

Part 4 of 6 — the dynamics

55. Why not just solve ∇f = 0?

Intuition

For our quadratic we can solve ∇f=0 in closed form — that is exactly the normal equations of Lesson 7. So why iterate at all?

Because for most losses ∇f(w)=0 is a giant nonlinear system with no closed form. Logistic regression, neural nets — you cannot write the answer down. The only tool that scales is: compute the gradient, take a step downhill, repeat.

Convexity is what makes that humble loop trustworthy: on a convex loss, following the gradient down always arrives at the one global minimum. We study the loop on the quadratic because there we know the true answer and can check the iterates against it.

56. The gradient-descent update

Concept

Gradient descent takes repeated downhill steps. From the current point wₜ, step against the gradient by a learning rate η > 0:

\[ w_{t+1} = w_t - \eta\,\nabla f(w_t) \]

The gradient points uphill (steepest ascent), so the minus sign heads downhill. η sets the stride. The entire question of Parts 4–5 is: how big can η be before the strides overshoot and blow up?

57. The gradient is the steepest-ascent direction

Concept

Why step against ∇f specifically? Among all unit directions u, the directional derivative ∇f(w)ᵀu — the instantaneous slope — is largest when u points along ∇f, and most negative when u = −∇f.

\[ \min_{\lVert u\rVert = 1} \nabla f(w)^\top u = -\lVert \nabla f(w)\rVert \;\text{ at }\; u = -\frac{\nabla f(w)}{\lVert \nabla f(w)\rVert} \]

So −∇f is the locally fastest way down — the greediest possible descent direction. Gradient descent is just 'walk the steepest way down, one stride at a time'.

58. Plan first: The 1-D shadow: derive the update factor

Step zero

Discussion prompt

The 1-D shadow: derive the update factor — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Plug f'(x)=2x into the update

Answer:

  1. Plug f'(x)=2x into the update
  2. Factor out xₜ
  3. Unroll to a closed form

59. The 1-D shadow: derive the update factor

Worked example

Use the clean shadow f(x)=x². Its gradient is f'(x)=2x. Substitute into the update and simplify — this reveals GD as pure multiplication.

Plug f'(x)=2x into the update

Why: The general rule x_{t+1} = x_t − η f'(x_t) becomes x_{t+1} = x_t − η·2x_t.

\[ x_{t+1} = x_t - \eta\,(2 x_t) \]

Factor out xₜ

Why: Both terms carry x_t, so pull it out. Each step just MULTIPLIES the current point by the fixed number (1 − 2η).

\[ x_{t+1} = (1 - 2\eta)\, x_t \]

Unroll to a closed form

Why: Applying the same factor t times gives a geometric sequence. Whether it shrinks or grows is decided entirely by |1 − 2η|.

\[ x_t = (1 - 2\eta)^t\, x_0 \]

60. Draw the shape of it: The 1-D shadow: derive the update factor

Blank canvas

Draw it

Draw what The 1-D shadow: derive the update factor just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

61. Contraction: the whole game is |1 − 2η|

Concept

xₜ = (1−2η)ᵗ x₀ is a geometric sequence. Its fate hinges on one number, the contraction factor |1 − 2η|:

Solve |1 − 2η| < 1: it means −1 < 1 − 2η < 1, i.e. 0 < η < 1. So for f(x)=x², gradient descent converges exactly when 0 < η < 1. Keep that interval — the next slide watches it break.

62. Guess the shape of the answer: Four learning rates on f(x)=x²

Estimation

Predict first

Start at x₀=2 and run six steps for η ∈ {0.1, 0.5, 0.9, 1.1}. Watch the contraction factor decide each fate. Runnable:

Commit before you compute: what does Four learning rates on f(x)=x² come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: η=0.5 lands at 0 in ONE step; η=1.1 blows up

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. η=0.5 gives factor 0 — the perfect step for this f.

63. Four learning rates on f(x)=x²

Worked example

Start at x₀=2 and run six steps for η ∈ {0.1, 0.5, 0.9, 1.1}. Watch the contraction factor decide each fate. Runnable:

for lr in [0.1, 0.5, 0.9, 1.1]:
    x = 2.0; traj = [x]
    for _ in range(6):
        x = x - lr*(2*x)        # gradient step, grad = 2x
        traj.append(round(x, 4))
    print(lr, 'factor', round(1-2*lr,2), traj)

η=0.5 lands at 0 in ONE step; η=1.1 blows up

Why: η=0.5 gives factor 0 — the perfect step for this f. η=0.9 gives factor −0.8, so it converges but flips sign each step. η=1.1 gives factor −1.2 (>1 in magnitude), so |x| grows every step: divergence.

ηfactor 1−2ηx trajectory from 2.0fate
0.1+0.802 → 1.6 → 1.28 → 1.024 → 0.819 → 0.655 → 0.524converges (slow)
0.50.002 → 0 → 0 → 0 → 0 → 0 → 0converges (1 step)
0.9−0.802 → −1.6 → 1.28 → −1.024 → 0.819 → −0.655 → 0.524converges (oscillating)
1.1−1.202 → −2.4 → 2.88 → −3.456 → 4.147 → −4.977 → 5.972DIVERGES

64. Restore the missing line: Full trace: loss and gradient at η = 0.1

Fill the middle

Fill in the blanks

From Full trace: loss and gradient at η = 0.1 — one line has had its right-hand side removed. Put it back.

x, lr = 2.0, 0.1
for t in range(6):
print(t, round(x,5), round(xx,6), round(2x,5)) # t, x, f=x^2, grad=2x
x = *x - lr(2x)* # step downhill

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Near the minimum the gradient fades, so each step gets smaller on its own — GD naturally decelerates as it homes in.

65. Full trace: loss and gradient at η = 0.1

Worked example

Zoom in on the safe run η=0.1 from x₀=2 and log every quantity — position, loss, and gradient — pass by pass. This is the trace to internalize.

x, lr = 2.0, 0.1
for t in range(6):
    print(t, round(x,5), round(x*x,6), round(2*x,5))  # t, x, f=x^2, grad=2x
    x = x - lr*(2*x)                                    # step downhill

As x → 0, both the loss x² and the gradient 2x shrink toward 0

Why: Near the minimum the gradient fades, so each step gets smaller on its own — GD naturally decelerates as it homes in. The loss falls by the factor (1−2η)² = 0.64 per step, from 4.0 down.

step txf(x) = x²gradient 2x
02.000004.0000004.00000
11.600002.5600003.20000
21.280001.6384002.56000
31.024001.0485762.04800
40.819200.6710891.63840
50.655360.4294971.31072

66. Watch it run: Full trace: loss and gradient at η = 0.1

Pattern

Step through it

Step through Full trace: loss and gradient at η = 0.1 one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step t is 0
  2. Step 2: step t is 1
  3. Step 3: step t is 2
  4. Step 4: step t is 3
  5. Step 5: step t is 4
  6. Step 6: step t is 5

67. Something is wrong here: is a bigger learning rate always faster?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Bigger steps cover more ground, so cranking η up must reach the minimum sooner. Use η = 1.1 on f(x)=x².

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound.

There is a ceiling. Stay inside 0 < η < 1 (here), and the fastest step is the one that makes the factor smallest — η = 0.5 (factor 0).

Why: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound. Bigger is not faster — bigger is broken.

68. Trap: is a bigger learning rate always faster?

Trap

The trap

Bigger steps cover more ground, so cranking η up must reach the minimum sooner. Use η = 1.1 on f(x)=x².

η = 1.1 → factor |1−2η| = 1.2 > 1

Why: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound. Bigger is not faster — bigger is broken.

The fix

There is a ceiling. Stay inside 0 < η < 1 (here), and the fastest step is the one that makes the factor smallest — η = 0.5 (factor 0).

η = 0.5 → factor 0 → one-step convergence

Why: Within the stable range the error contracts every step; right at η = 1/L the factor hits its minimum. The 'sweet spot' is a balance the smoothness constant L pins down exactly — derived next.

69. Which of these survive contact with Lesson 11: Convex Optimization?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
We carry one function through the whole lesson so every slide compounds. It is a 2-variable quadratic in w = (w₁, w₂):; Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.; Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.
Breaks
To test convexity, find the critical point and check the Hessian there. If ∇²f ⪰ 0 at that point, call f convex.; Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 11: Convex Optimization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

70. The step-size bound η < 2/L

Section

Part 5 of 6 — smoothness & conditioning

71. Smoothness L bounds the curvature

Concept

L-smoothness — A function is L-smooth if its gradient changes no faster than L: ‖∇f(a) − ∇f(b)‖ ≤ L‖a − b‖. For a twice-differentiable f this L is the LARGEST eigenvalue of the Hessian — the steepest curvature anywhere.

L is the tightest speed limit on the gradient. Along the steepest-curvature direction, the same step size travels the most — so L is what caps how big a step you can safely take.

\[ L = \lambda_{\max}\!\big(\nabla^2 f\big) \]

72. Teach it back: Smoothness L bounds the curvature

Explain it

Discussion prompt

Explain Smoothness L bounds the curvature to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

L is the tightest speed limit on the gradient. Along the steepest-curvature direction, the same step size travels the most — so L is what caps how big a step you can safely take.

73. What has to happen first: Derive the stable-step bound η < 2/L

Ranking

Put in order

Put the moves of Derive the stable-step bound η < 2/L into the order they have to happen.

  1. Require the contraction factor < 1 in magnitude
  2. Unfold the absolute value
  3. Verify: the tightest direction is λ = L

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Convergence in a direction with curvature λ needs |1 − ηλ| < 1, exactly the 1-D condition applied per eigen-direction.

74. Derive the stable-step bound η < 2/L

Worked example

On a quadratic with Hessian eigenvalue λ in some direction, GD multiplies the error in that direction by (1 − ηλ) each step — the vector version of the (1−2η) we found. Stability needs that factor below 1 in magnitude.

Require the contraction factor < 1 in magnitude

Why: Convergence in a direction with curvature λ needs |1 − ηλ| < 1, exactly the 1-D condition applied per eigen-direction.

\[ |1 - \eta \lambda| < 1 \]

Unfold the absolute value

Why: |1 − ηλ| < 1 means −1 < 1 − ηλ < 1. The right inequality gives ηλ > 0 (always true for η,λ > 0); the left gives ηλ < 2.

\[ 0 < \eta \lambda < 2 \;\Longrightarrow\; \eta < \frac{2}{\lambda} \]

Verify: the tightest direction is λ = L

Why: The bound must hold for EVERY eigen-direction at once, so the smallest ceiling wins — the one from the largest λ, namely L. Hence the single stable-step rule below.

\[ \boxed{\,0 < \eta < \frac{2}{L}, \qquad L = \lambda_{\max}(\nabla^2 f)\,} \]

75. Decode the notation: Derive the stable-step bound η < 2/L

Notation

Annotate

From Derive the stable-step bound η < 2/L — read this one piece at a time. What is each part doing?

On: \( \boxed{\,0 < \eta < \frac{2}{L}, \qquad L = \lambda_{\max}(\nabla^2 f)\,} \)

  • Convergence in a direction with curvature λ needs |1 − ηλ| < 1, exactly the 1-D condition applied per eigen-direction.
  • |1 − ηλ| < 1 means −1 < 1 − ηλ < 1. The right inequality gives ηλ > 0 (always true for η,λ > 0); the left gives ηλ < 2.
  • The bound must hold for EVERY eigen-direction at once, so the smallest ceiling wins — the one from the largest λ, namely L. Hence the single stable-step rule below.

76. Predict the next row: The bound on our 2-D bowl

Pattern

Predict first

The table runs: 0.2 | [1.049, 0.000] | +0.20 | converges · 0.5 | [0.062, 1.000] | −1.00 | w₂ stalls (|factor|=1)

In The bound on our 2-D bowl, given the rows so far: what is the next one — the row where η is 0.6?

Correct: 0.6 | [0.016, 7.530] | −1.40 | DIVERGES in w₂

ηw after 6 stepsw₂ factor 1−4ηfate
0.2[1.049, 0.000]+0.20converges
0.5[0.062, 1.000]−1.00w₂ stalls (|factor|=1)
0.6[0.016, 7.530]−1.40DIVERGES in w₂

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The w₂ curvature is L=4, so its factor is 1−0.6·4 = −1.4, magnitude > 1: w₂ explodes (1 → −1.4 → 1.96 → …) even though w₁'s factor 1−0.6·1 = 0.4 is fine.

77. The bound on our 2-D bowl

Worked example

Back to f(w)=½(w₁²+4w₂²), Hessian diag(1,4), so L=4, μ=1, and the bound is η < 2/4 = 0.5. Run η=0.2 (safe), η=0.5 (edge), η=0.6 (over). Runnable:

import numpy as np
def grad(w): return np.array([w[0], 4*w[1]])   # Hessian diag(1,4)
for eta in [0.2, 0.5, 0.6]:
    w = np.array([4., 1.])
    for _ in range(6):
        w = w - eta*grad(w)
    print(eta, np.round(w, 3), 'w2 factor', round(1-eta*4, 2))

η=0.6 > 2/L=0.5 → the w₂ direction diverges

Why: The w₂ curvature is L=4, so its factor is 1−0.6·4 = −1.4, magnitude > 1: w₂ explodes (1 → −1.4 → 1.96 → …) even though w₁'s factor 1−0.6·1 = 0.4 is fine. ONE bad direction dooms the whole run.

ηw after 6 stepsw₂ factor 1−4ηfate
0.2[1.049, 0.000]+0.20converges
0.5[0.062, 1.000]−1.00w₂ stalls (|factor|=1)
0.6[0.016, 7.530]−1.40DIVERGES in w₂

78. Conditioning κ = L/μ sets the pace

Concept

The stable step is capped by the largest curvature L, but progress along the smallest-curvature direction μ = λ_min is what's slow. The gap between them is the trouble.

condition number κ — κ = L/μ = λ_max / λ_min of the Hessian. κ=1 is a perfectly round bowl (GD converges in ~1 step); large κ is a long narrow ravine where GD zig-zags. Convex GD needs about O(κ log 1/ε) steps.

Our bowl has κ = 4/1 = 4 — mild. The Lesson 7 MSE Hessian had κ ≈ 118.3/1.69 ≈ 70 — the SAME number that warned us off inverting XᵀX. Big κ is slow GD and fragile linear algebra; rescaling features fixes both.

79. By analogy: Conditioning κ = L/μ sets the pace

Analogy

Discussion prompt

Explain Conditioning κ = L/μ sets the pace by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The stable step is capped by the largest curvature L, but progress along the smallest-curvature direction μ = λ_min is what's slow. The gap between them is the trouble.

80. What has to be given first: Strong convexity gives geometric decay

Missing information

Discussion prompt

When μ > 0 (a curvature floor — 'strongly convex'), the loss gap shrinks by a constant ratio every step. On our bowl at η = 1/L = 0.25, the per-step factor settles to (1 − μ/L)² = (1 − 0.25)² = 0.5625. Runnable:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Geometric (a.k.a. LINEAR) convergence: the gap is multiplied by a constant < 1 each step, so it reaches ε in O(log 1/ε) steps. The w₂ term dies instantly (factor 0), leaving the slow w₁ direction to set the 0.5625 rate. Verified by execution.

81. Strong convexity gives geometric decay

Worked example

When μ > 0 (a curvature floor — 'strongly convex'), the loss gap shrinks by a constant ratio every step. On our bowl at η = 1/L = 0.25, the per-step factor settles to (1 − μ/L)² = (1 − 0.25)² = 0.5625. Runnable:

import numpy as np
def f(w):    return 0.5*(w[0]**2 + 4*w[1]**2)
def grad(w): return np.array([w[0], 4*w[1]])
w = np.array([4., 1.]); eta = 0.25; prev = None
for t in range(6):
    if prev: print(t, round(f(w),6), 'ratio', round(f(w)/prev,4))
    prev = f(w); w = w - eta*grad(w)

The loss ratio locks onto 0.5625 = (1−μ/L)²

Why: Geometric (a.k.a. LINEAR) convergence: the gap is multiplied by a constant < 1 each step, so it reaches ε in O(log 1/ε) steps. The w₂ term dies instantly (factor 0), leaving the slow w₁ direction to set the 0.5625 rate. Verified by execution.

step tf(wₜ)ratio f(wₜ)/f(wₜ₋₁)
14.5000000.4500
22.5312500.5625
31.4238280.5625
40.8009030.5625
50.4505080.5625

82. Watch it run: Strong convexity gives geometric decay

Pattern

Step through it

Step through Strong convexity gives geometric decay one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step t is 1
  2. Step 2: step t is 2
  3. Step 3: step t is 3
  4. Step 4: step t is 4
  5. Step 5: step t is 5

83. What 'O(1/t)' actually costs you

Intuition

The rate classes sound abstract until you price them. O(1/t) (sublinear) means to halve the error you roughly double the steps — get to 1e-3, and 1e-6 costs a thousand times more work. Painful for high precision.

O(cᵗ) (linear/geometric) means each fixed block of steps multiplies the error by a constant — so reaching 1e-6 costs only about twice the steps of reaching 1e-3. That is the prize strong convexity (μ>0) hands you, and why our 0.5625-ratio bowl felt fast.

84. The convergence-rate classes

Concept

How fast GD closes the gap to the minimum depends on the curvature class. Memorize this table — it is exam bread and butter:

objective classGD error after t stepsmeaning
convex + L-smoothO(1/t)sublinear — halving error costs ~2× steps
strongly convex (μ>0)O(cᵗ), c<1linear/geometric — like our 0.5625 ratio
convex + Nesterov momentumO(1/t²)accelerated — the optimal first-order rate

A curvature floor μ>0 upgrades O(1/t) to exponential decay. Momentum buys the 1/t² rate without needing that floor — it's the free lunch of Week 6's optimizers.

85. Fill in: meaning for The convergence-rate classes

Comparison

Comparison matrix

From The convergence-rate classes: refill the meaning column from what you know. The rest of the table is as it appeared.

objective classGD error after t stepsmeaning
convex + L-smoothO(1/t)sublinear — halving error costs ~2× steps
strongly convex (μ>0)O(cᵗ), c<1linear/geometric — like our 0.5625 ratio
convex + Nesterov momentumO(1/t²)accelerated — the optimal first-order rate

86. Ill-conditioning makes GD zig-zag

Intuition

Picture a long narrow valley — steep walls (large L), gentle floor (small μ). Plain GD bounces across the steep walls while barely inching along the floor toward the minimum. That is the κ=L/μ slowdown made visual.

Momentum fixes it by accumulating velocity: v ← βv − η∇f, then w ← w + v. The back-and-forth wall-bounces cancel while the steady floor-ward motion builds up — so you glide down the valley instead of rattling across it.

87. Guess the shape of the answer: Momentum beats plain GD on our bowl

Estimation

Predict first

Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:

Commit before you compute: what does Momentum beats plain GD on our bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Momentum reaches 1e-8 in 12 steps vs plain GD's 21

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Heavy-ball with η=0.4444, β=0.1111 nearly halves the step count on this mild κ=4 bowl.

88. Momentum beats plain GD on our bowl

Worked example

Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:

import numpy as np
f    = lambda w: 0.5*(w[0]**2 + 4*w[1]**2)
grad = lambda w: np.array([w[0], 4.*w[1]])
L, mu = 4., 1.
def plain():
    w = np.array([4.,1.]); eta = 2/(L+mu)
    k = 0
    while f(w) >= 1e-8: w = w - eta*grad(w); k += 1
    return k
def heavy_ball():
    w = np.array([4.,1.]); v = np.zeros(2)
    eta  = 4/((np.sqrt(L)+np.sqrt(mu))**2)          # 0.4444
    beta = ((np.sqrt(L/mu)-1)/(np.sqrt(L/mu)+1))**2 # 0.1111
    k = 0
    while f(w) >= 1e-8: v = beta*v - eta*grad(w); w = w + v; k += 1
    return k
print('plain GD steps:', plain(), ' heavy-ball steps:', heavy_ball())

Momentum reaches 1e-8 in 12 steps vs plain GD's 21

Why: Heavy-ball with η=0.4444, β=0.1111 nearly halves the step count on this mild κ=4 bowl. The bigger κ is, the larger the gap — momentum's O(√κ) beats plain GD's O(κ). Verified by execution.

methodparamssteps to loss < 1e-8
plain GDη = 2/(L+μ) = 0.421
heavy-ball momentumη = 0.4444, β = 0.111112
momentum trace f(w)10 → 3.68 → 0.87 → 0.17 → …geometric, faster ratio

89. Work backwards from the answer: Momentum beats plain GD on our bowl

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Momentum reaches 1e-8 in 12 steps vs plain GD's 21

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:

90. Without one step: Diagnosing any optimization problem

Constraint

Discussion prompt

Run Diagnosing any optimization problem with this step confiscated:

Conditioning κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Convex? Compute the Hessian; check it is PSD everywhere (all eigenvalues ≥ 0). If so, any stationary point is the global minimum.
  2. Smoothness L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.
  3. Conditioning κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.
  4. Diverging? the iterate grows → η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.
  5. Non-convex (neural nets): expect multiple minima → use SGD, momentum, and good initialization; accept you find a good minimum, not the one.

91. Diagnosing any optimization problem

Pattern

  1. Convex? Compute the Hessian; check it is PSD everywhere (all eigenvalues ≥ 0). If so, any stationary point is the global minimum.
  2. Smoothness L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.
  3. Conditioning κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.
  4. Diverging? the iterate grows → η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.
  5. Non-convex (neural nets): expect multiple minima → use SGD, momentum, and good initialization; accept you find a good minimum, not the one.

92. Where does it stop working: Diagnosing any optimization problem

Edge cases

Discussion prompt

Diagnosing any optimization problem works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Convex? Compute the Hessian; check it is PSD everywhere (all eigenvalues ≥ 0). If so, any stationary point is the global minimum.
  2. Smoothness L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.
  3. Conditioning κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.
  4. Diverging? the iterate grows → η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.
  5. Non-convex (neural nets): expect multiple minima → use SGD, momentum, and good initialization; accept you find a good minimum, not the one.

93. Rule out three: Check yourself — the certificate for convexity

Elimination

Eliminate the wrong options

A twice-differentiable f is convex if and only if its Hessian is…

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. positive semidefinite at every point in the domain
  • B. positive semidefinite at the minimum
  • C. invertible everywhere
  • D. symmetric

Survives elimination: A

Why: Convexity is equivalent to ∇²f ⪰ 0 everywhere — non-negative curvature in every direction at every point. Then every stationary point is a global minimum. For a constant Hessian (like diag(1,4)) one eigenvalue check settles it for all w.

94. Check yourself — the certificate for convexity

Check

Which condition proves convexity for a twice-differentiable f?

Check your understanding

A twice-differentiable f is convex if and only if its Hessian is…

  • A. positive semidefinite at every point in the domain (correct)
  • B. positive semidefinite at the minimum
  • C. invertible everywhere
  • D. symmetric

Answer: A

Why: Convexity is equivalent to ∇²f ⪰ 0 everywhere — non-negative curvature in every direction at every point. Then every stationary point is a global minimum. For a constant Hessian (like diag(1,4)) one eigenvalue check settles it for all w.

Why B tempts people
PSD only at the minimum certifies a LOCAL minimum. g(t)=t⁴−3t² is convex near each of its two minima yet non-convex overall — exactly the trap we dissected.
Why C tempts people
Invertibility (no zero eigenvalue) is about strict/positive-definite convexity, and isn't required — a convex function can have a singular PSD Hessian with a flat direction.
Why D tempts people
Hessians are ALWAYS symmetric (mixed partials commute), so symmetry says nothing about convexity — it can't distinguish a bowl from a saddle.

95. Answer it before you see the options: Check yourself — global minima

Prediction

Predict first

Gradient descent is run on a strictly convex loss from two different starting points (both with a stable learning rate). What happens?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Both converge to the same global minimum

Why: A strictly convex function has a single global minimum and no other stationary points, so any convergent run reaches the same point regardless of initialization — the ∇f(x)=0 ⇒ global proof guarantees it.

96. Check yourself — global minima

Check

What does convexity guarantee about where you land?

Check your understanding

Gradient descent is run on a strictly convex loss from two different starting points (both with a stable learning rate). What happens?

  • A. Both converge to the same global minimum (correct)
  • B. Each may land in a different local minimum
  • C. One will diverge, the other converge
  • D. They converge only if started close together

Answer: A

Why: A strictly convex function has a single global minimum and no other stationary points, so any convergent run reaches the same point regardless of initialization — the ∇f(x)=0 ⇒ global proof guarantees it.

Why B tempts people
Multiple local minima are a NON-convex phenomenon (like g(t)=t⁴−3t² or a neural net). A strictly convex function has exactly one — restarts are pointless.
Why C tempts people
Divergence is set by the learning rate (η vs 2/L), not the start point. With a stable η both converge; with η too large both diverge — the start doesn't decide it.
Why D tempts people
Convexity makes the result independent of where you start; proximity to the optimum is never required for convergence, only a stable step size.

97. Answer it before you see the options: Check yourself — learning-rate stability

Prediction

Predict first

For f(x) = x² (so f''=2, hence L=2), which learning rate makes gradient descent DIVERGE?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: η = 1.1

Why: The update multiplies x by (1−2η). Stability needs |1−2η| < 1, i.e. 0 < η < 1 = 2/L. η=1.1 gives factor −1.2, magnitude > 1, so |x| grows each step: 2 → −2.4 → 2.88 → … — divergence.

98. Check yourself — learning-rate stability

Check

Use the smoothness bound η < 2/L.

Check your understanding

For f(x) = x² (so f''=2, hence L=2), which learning rate makes gradient descent DIVERGE?

  • A. η = 1.1 (correct)
  • B. η = 0.5
  • C. η = 0.1
  • D. η = 0.9

Answer: A

Why: The update multiplies x by (1−2η). Stability needs |1−2η| < 1, i.e. 0 < η < 1 = 2/L. η=1.1 gives factor −1.2, magnitude > 1, so |x| grows each step: 2 → −2.4 → 2.88 → … — divergence.

Why B tempts people
η=0.5 gives factor 0 — it reaches the minimum in a single step. The fastest possible, the opposite of divergence.
Why C tempts people
η=0.1 gives factor +0.8 — slow but steadily converging, well inside the stable range 0 < η < 1.
Why D tempts people
η=0.9 gives factor |1−1.8| = 0.8 < 1 — it converges, just flipping sign each step. Under the η<1 threshold, so still stable.

99. Rule out three: Check yourself — conditioning and speed

Elimination

Eliminate the wrong options

A convex quadratic has Hessian eigenvalues μ=1 and L=100. Its learning rate must satisfy η<2/L=0.02, capped by the STEEP direction. What limits convergence speed?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The tiny μ direction: progress there scales with ημ ≤ 0.02, so κ=L/μ=100 makes it slow
  • B. Nothing — a convex problem always converges in one step
  • C. The steep L direction, which converges slowest
  • D. The learning rate is irrelevant to speed once the problem is convex

Survives elimination: A

Why: η is capped at ~1/L by the steep direction, but that same small η makes progress in the shallow μ direction glacial: its factor is 1−ημ ≈ 1−μ/L = 0.99. The condition number κ = L/μ = 100 sets the pace — O(κ log 1/ε) steps.

100. Check yourself — conditioning and speed

Check

Think about which eigenvalue caps the step and which one is slow.

Check your understanding

A convex quadratic has Hessian eigenvalues μ=1 and L=100. Its learning rate must satisfy η<2/L=0.02, capped by the STEEP direction. What limits convergence speed?

  • A. The tiny μ direction: progress there scales with ημ ≤ 0.02, so κ=L/μ=100 makes it slow (correct)
  • B. Nothing — a convex problem always converges in one step
  • C. The steep L direction, which converges slowest
  • D. The learning rate is irrelevant to speed once the problem is convex

Answer: A

Why: η is capped at ~1/L by the steep direction, but that same small η makes progress in the shallow μ direction glacial: its factor is 1−ημ ≈ 1−μ/L = 0.99. The condition number κ = L/μ = 100 sets the pace — O(κ log 1/ε) steps.

Why B tempts people
One-step convergence happens only when κ=1 (a round bowl). Here κ=100, so it takes on the order of 100·log(1/ε) steps.
Why C tempts people
The steep L direction actually converges FASTEST (its factor is smallest under η≈1/L). It's the shallow μ direction that lags — the common mix-up.
Why D tempts people
η is central to speed: too small crawls, too large (>2/L) diverges. Convexity guarantees you CAN converge, not that η stops mattering.

101. Your turn: GD convergence

Section

Part 6 of 6 — the project

102. Project: a learning-rate sweep

Concept

Run gradient descent on f(x)=x² for several learning rates, print each full trajectory, and classify every run as converging or diverging directly from its contraction factor |1−2η|. You derived every piece — now assemble it.

#requirementtool
1Define f and its gradientf = x², grad = 2x
2A GD loop that returns the trajectorya for-loop, append each x
3Sweep η and flag |1−2η| ≥ 1 as divergingcompare the factor

Build rules: type every line yourself, print the whole trajectory (don't just trust the final number), and watch the sign flip when η pushes past the stable range.

103. Break it if you can: Project: a learning-rate sweep

Counterexample

Discussion prompt

Build rules: type every line yourself, print the whole trajectory (don't just trust the final number), and watch the sign flip when η pushes past the stable range.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

104. Milestone 1 — function and gradient

Worked example

Your turn: write f and grad. Say aloud what grad(2) should be before you print it.

Hint: f = x**2, grad = 2*x — the same shape as Lesson 6's warm-up. The gradient at x=2 is 2·2 = 4.

def f(x):    return x**2
def grad(x): return 2*x
print(f(2), grad(2))
callvalue (verified)
f(2)4
grad(2)4
grad(0) (the minimum)0

105. Milestone 2 — the GD loop

Worked example

Your turn: from x=2, step six times at η=0.1 and collect the trajectory. Predict whether it reaches 0, and how fast.

Hint: x = x - lr*grad(x) inside the loop; append x each pass. The factor is 1−2·0.1 = 0.8, so each step keeps 80%.

def grad(x): return 2*x
x, lr = 2.0, 0.1
traj = [x]
for _ in range(6):
    x = x - lr*grad(x)
    traj.append(round(x, 4))
print(traj)
stepxf(x)=x²
02.00004.0000
11.60002.5600
21.28001.6384
31.02401.0486
60.52430.2749

106. Watch it run: Milestone 2 — the GD loop

Pattern

Step through it

Step through Milestone 2 — the GD loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: step is 0
  2. Step 2: step is 1
  3. Step 3: step is 2
  4. Step 4: step is 3
  5. Step 5: step is 6

107. Milestone 3 — sweep and classify

Worked example

Your turn: loop over η ∈ {0.1, 0.5, 1.1} and label each by its contraction factor |1−2η|. Predict which one diverges before you run it.

Hint: factor = abs(1 - 2*lr), then 'diverges' if factor >= 1 else 'converges'. Only η=1.1 should cross the line.

for lr in [0.1, 0.5, 1.1]:
    factor = abs(1 - 2*lr)
    tag = 'diverges' if factor >= 1 else 'converges'
    print(lr, round(factor, 2), tag)
η|1−2η|verdict (verified)
0.10.80converges
0.50.00converges (1 step)
1.11.20diverges

108. Milestone 4 — the full program

Worked example

Your turn: fold the loop and the classifier into one function run(lr) that returns both the trajectory and the factor, then sweep. This is the whole experiment in one screen.

Hint: return (traj, abs(1-2*lr)) from run, and tag each sweep line by the factor.

def f(x):    return x**2
def grad(x): return 2*x

def run(lr, x0=2.0, steps=6):
    x, traj = x0, [round(x0,4)]
    for _ in range(steps):
        x = x - lr*grad(x)
        traj.append(round(x, 4))
    return traj, abs(1 - 2*lr)

for lr in [0.1, 0.5, 1.1]:
    traj, factor = run(lr)
    tag = 'diverges' if factor >= 1 else 'converges'
    print(lr, tag, traj)
ηverdicttrajectory ends near
0.1converges0.5243 (heading to 0)
0.5converges0.0 (one step)
1.1diverges5.972 and growing

109. What each one costs: Milestone 4 — the full program

Trade off

Comparison matrix

From Milestone 4 — the full program: every row here is a choice with a cost. Fill the verdict column, then say which row you would actually pick and what you give up for it.

ηverdicttrajectory ends near
0.1converges0.5243 (heading to 0)
0.5converges0.0 (one step)
1.1diverges5.972 and growing

110. Stretch: watch the 2-D bound bite

Concept

Swap the scalar f for the 2-D bowl f(w)=½(w₁²+4w₂²) with grad = [w₁, 4w₂], then run η=0.6. Because 2/L = 0.5, the steep w₂ direction diverges while w₁ behaves — the single-bad-direction failure from Part 5.

import numpy as np
def grad2(w): return np.array([w[0], 4*w[1]])
w = np.array([4., 1.])
for t in range(6):
    w = w - 0.6*grad2(w)
    print(t, np.round(w, 3))
stepw = [w₁, w₂]note
1[1.600, −1.400]w₂ already overshot past 0
2[0.640, 1.960]w₂ growing, sign flipping
4[0.102, 3.842]w₂ diverging (factor −1.4)
6[0.016, 7.530]w₁ fine, w₂ blown up

111. Fill in: w = [w₁, w₂] for Stretch: watch the 2-D bound bite

Comparison

Comparison matrix

From Stretch: watch the 2-D bound bite: refill the w = [w₁, w₂] column from what you know. The rest of the table is as it appeared.

stepw = [w₁, w₂]note
1[1.600, −1.400]w₂ already overshot past 0
2[0.640, 1.960]w₂ growing, sign flipping
4[0.102, 3.842]w₂ diverging (factor −1.4)
6[0.016, 7.530]w₁ fine, w₂ blown up

112. Show it off

Concept

Slides closed, out loud: explain (1) why a convex function has no local-minima traps, using the tangent inequality; (2) what L is and where 2/L comes from; and (3) why the MSE loss is convex but a neural net's loss is not.

Stretch homework: prove the sum of two convex functions is convex (add their chord inequalities), and prove MSE is convex via the second-order condition (∇² = 2XᵀX ⪰ 0). Both feed straight into Lesson 13's PSD machinery and Week 6's SGD and momentum variants.

113. Connect it up: Lesson 11: Convex Optimization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — What makes a problem convex · When convexity fails · Why convex ⇒ only global minima · Gradient descent as a contraction · The step-size bound η < 2/L · Your turn: GD convergence. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

114. What you can do now

Recap

ideathe one thing to remember
convex testHessian PSD EVERYWHERE, not just at the min
minimaconvex ⇒ every stationary point is global
GD stepx_{t+1} = (1−2η)x_t on x²; factor |1−2η| decides all
step sizeη < 2/L, L = largest Hessian eigenvalue
speedκ = L/μ large ⇒ slow; μ>0 ⇒ geometric decay

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 11 (Week 4 — Convex Optimization) — Barron · USAAIO Round 2 Preparation, 2026
  2. Boyd & Vandenberghe, Convex Optimization, Ch. 3 (convex functions) & Ch. 9 (unconstrained minimization) — Cambridge University Press, 2004
  3. Nesterov, Introductory Lectures on Convex Optimization, §2.1 (smoothness, GD rates) — Springer, 2004
  4. NumPy linalg.eigvalsh
  5. Every trajectory, eigenvalue, and contraction factor produced by real execution — numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108