USAAIO Lesson 11, from Week 4 on optimization, fully worked. It proves the chord definition of convexity on a running quadratic, then derives the first-order tangent condition and the second-order condition that the Hessian be positive semi-definite, step by step, and dissects a non-convex counterexample. It shows why every stationary point of a convex function is global, writes gradient descent as a contraction with the update factor derived from scratch, proves the stable-step bound eta < 2/L, and covers the speed limit set by the condition number along with the classes of convergence rate. It ends with a from-scratch gradient-descent experiment verified against real execution. Every number was produced by running the code. The lesson runs to 61 slides.
Subject: Machine Learning · 114 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 11 · Week 4 (Optimization)
Why some loss surfaces are easy and others aren't. We prove convexity three equivalent ways, show why a convex bowl has no traps, then derive gradient descent as a contraction and the exact step-size bound η < 2/L — every algebra move shown, every number produced by running the code.
Objectives
g(t)=t⁴−3t² and locate exactly where the Hessian goes negativef, ∇f(x)=0 forces a global minimum — no local trapsη < 2/Lκ = L/μ to convergence speed and quote the O(1/t) / linear / O(1/t²) rate classesWarm-up
Discussion prompt
Before we open Lesson 11: Convex Optimization: without looking back, what was the main idea of Object Detection & YOLO, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
bounding-box parameterisation (corner vs centre-offset formats), IoU from scratch, non-maximum suppression (NMS), YOLO grid/anchor math (S=13, B=5, C=80 → 17 745 raw predictions), anchor-offset decoding with sigmoid/exp, Feature Pyramid Networks (FPN) for multi-scale detection, and mAP@0.5 calculation.
Section
Part 1 of 6 — the definition
Concept
We carry one function through the whole lesson so every slide compounds. It is a 2-variable quadratic in w = (w₁, w₂):
\[ f(w) = \tfrac{1}{2}\big(w_1^2 + 4\,w_2^2\big) \]
Its lowest point is clearly w = (0, 0), where f = 0. The 4 on w₂ makes the bowl steeper in the w₂ direction than the w₁ direction — that asymmetry becomes the whole story of Part 5.
For the pure learning-rate arithmetic we also keep a 1-D shadow, f(x) = x², because its gradient-descent step is a single clean multiplication.
Counterexample
Discussion prompt
We carry one function through the whole lesson so every slide compounds. It is a 2-variable quadratic in w = (w₁, w₂):
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Its lowest point is clearly w = (0, 0), where f = 0. The 4 on w₂ makes the bowl steeper in the w₂ direction than the w₁ direction — that asymmetry becomes the whole story of Part 5.
Intuition
Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.
If you stretch a string between any two points on the surface, the string never dips below the surface — it stays on or above it. That taut-string picture is exactly the formal definition we write next.
Why obsess over this? On a convex bowl, walking downhill always reaches the one true bottom. On a dented surface, downhill can dump you into the wrong valley — the difference between 'training just works' and 'training is an art'.
Analogy
Discussion prompt
Explain Convex = one bowl, no dents by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.
Concept
Take any two inputs x and y. The straight line between the two surface points (x, f(x)) and (y, f(y)) is the chord. Convexity says the function sits on or below its chords:
\[ f\big(\alpha x + (1-\alpha)y\big) \;\le\; \alpha f(x) + (1-\alpha)f(y), \qquad \alpha \in [0,1] \]
convex function — A function whose value at any weighted average of two points is at most the same weighted average of the two function values. Geometrically: the graph never rises above a chord. α sweeps the chord: α=1 is the x end, α=0 is the y end, α=½ is the midpoint.
Estimation
Predict first
Claim: f(w)=½(w₁²+4w₂²) is convex. Let's not take it on faith — pick x=(4,1) and y=(−2,3), sweep α, and check the chord inequality numerically. Runnable as written:
Commit before you compute: what does Verify the chord inequality on our bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: At every α, f(mix) ≤ chord — the inequality holds
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The surface height never exceeds the chord height.
Worked example
Claim: f(w)=½(w₁²+4w₂²) is convex. Let's not take it on faith — pick x=(4,1) and y=(−2,3), sweep α, and check the chord inequality numerically. Runnable as written:
import numpy as np
def f(w): return 0.5*(w[0]**2 + 4*w[1]**2)
x = np.array([4., 1.]); y = np.array([-2., 3.])
for a in [0., 0.25, 0.5, 0.75, 1.]:
mix = a*x + (1-a)*y # point on the segment
chord = a*f(x) + (1-a)*f(y) # height of the chord
print(a, round(f(mix),4), round(chord,4), f(mix) <= chord)At every α, f(mix) ≤ chord — the inequality holds
Why: The surface height never exceeds the chord height. A single violation would disprove convexity; none appears, consistent with (but not yet a proof of) convexity.
| α | f(mix) | chord height | f(mix) ≤ chord? |
|---|---|---|---|
| 0.00 | 20.0000 | 20.0000 | True (y endpoint) |
| 0.25 | 12.6250 | 17.5000 | True |
| 0.50 | 8.5000 | 15.0000 | True |
| 0.75 | 7.6250 | 12.5000 | True |
| 1.00 | 10.0000 | 10.0000 | True (x endpoint) |
Concept
Here is the chord condition drawn for a 1-D convex slice. The straight chord connects two surface points; the curve sags below it. That sag is convexity.
Figure (svg): A U-shaped convex curve with a straight chord drawn between two points on it; the curve lies entirely below the chord, and a vertical gap at the midpoint is labelled.
Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.
Explain it
Discussion prompt
Explain The picture: surface below its chord to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.
Intuition
The table checked one pair of points at five values of α. Convexity demands the inequality for every pair and every α — infinitely many checks. No finite table can settle it.
So we need a certificate: a condition we can check once, using derivatives, that guarantees the chord inequality for all pairs at once. That is what the first- and second-order conditions give us.
Concept
For a differentiable f, convexity is equivalent to: the tangent line (plane) at any point stays below the function everywhere. The linear approximation always underestimates.
\[ f(y) \;\ge\; f(x) + \nabla f(x)^\top (y - x) \qquad \forall\, x, y \]
Read it as: the true value f(y) is at least what the tangent at x predicts. This single inequality is the engine behind 'no local traps' — we cash it in shortly.
Missing information
Discussion prompt
Test the first-order condition on our bowl. Anchor the tangent at x=(2,1), where f(x)=4 and ∇f(x)=[2,4], then check f(y) ≥ tangent(y) at several y. Runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The tangent anchored at x underestimates f everywhere, touching only at y=x (both read 4). That is the first-order condition holding, and it is what forces a zero-gradient point to be global. Verified by execution.
Worked example
Test the first-order condition on our bowl. Anchor the tangent at x=(2,1), where f(x)=4 and ∇f(x)=[2,4], then check f(y) ≥ tangent(y) at several y. Runnable:
import numpy as np
def f(w): return 0.5*(w[0]**2 + 4*w[1]**2)
def grad(w): return np.array([w[0], 4*w[1]])
x = np.array([2., 1.])
for y in [[0.,0.], [3.,2.], [-1.,1.], [2.,1.]]:
y = np.array(y)
tang = f(x) + grad(x) @ (y - x) # tangent-plane height
print(y.tolist(), round(f(y),3), round(tang,3), f(y) >= tang - 1e-9)At every y, f(y) ≥ tangent — the plane stays under the surface
Why: The tangent anchored at x underestimates f everywhere, touching only at y=x (both read 4). That is the first-order condition holding, and it is what forces a zero-gradient point to be global. Verified by execution.
| y | f(y) | tangent at x=(2,1) | f(y) ≥ tangent? |
|---|---|---|---|
| [0, 0] | 0.0 | −4.0 | True |
| [3, 2] | 12.5 | 10.0 | True |
| [−1, 1] | 2.5 | −2.0 | True |
| [2, 1] | 4.0 | 4.0 | True (touches at y=x) |
Discrimination
Sort into buckets
Sort these by f(y) ≥ tangent?, from memory, without looking back at Verify the tangent lies below. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
For a twice-differentiable f, convexity is equivalent to the Hessian (the matrix of second partials) being positive semidefinite at every point of the domain:
\[ f \text{ convex} \iff \nabla^2 f(x) \succeq 0 \quad \forall\, x \]
positive semidefinite (PSD) — A symmetric matrix H with zᵀHz ≥ 0 for every vector z — equivalently, all eigenvalues ≥ 0. It means non-negative curvature in every direction: the surface never bends downward. This is the easiest of the three tests to check, since it is just an eigenvalue sign check.
Definition probe
Sort into buckets
Every line below is part of the definition of convex function or of positive semidefinite (PSD) — one or the other, never both. Put each where it belongs.
Ranking
Put in order
Put the moves of Compute the Hessian of our bowl by hand into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. ∂f/∂w₁ = ½·2w₁ = w₁ and ∂f/∂w₂ = ½·8w₂ = 4w₂.
Worked example
The Hessian collects all second partial derivatives. For f(w)=½(w₁²+4w₂²), take the partials one at a time — nothing skipped.
First partials
Why: ∂f/∂w₁ = ½·2w₁ = w₁ and ∂f/∂w₂ = ½·8w₂ = 4w₂. So ∇f = [w₁, 4w₂].
\[ \nabla f(w) = \begin{bmatrix} w_1 \\ 4 w_2 \end{bmatrix} \]
Second partials
Why: ∂²f/∂w₁² = 1, ∂²f/∂w₂² = 4, and the mixed partial ∂²f/∂w₁∂w₂ = 0 since w₁ and w₂ never multiply. The Hessian is constant — it does not depend on w.
\[ \nabla^2 f = \begin{bmatrix} 1 & 0 \\ 0 & 4 \end{bmatrix} \]
Verify: eigenvalues are 1 and 4, both ≥ 0
Why: A diagonal matrix's eigenvalues are its diagonal entries. Both positive ⇒ PSD (in fact positive definite) at every w ⇒ f is convex everywhere. Certificate obtained.
Notation
Annotate
From Compute the Hessian of our bowl by hand — read this one piece at a time. What is each part doing?
On: \( \nabla f(w) = \begin{bmatrix} w_1 \\ 4 w_2 \end{bmatrix} \)
Pattern
Predict first
The table runs: Hessian ∇²f | [[1, 0], [0, 4]] (constant in w) · eigenvalues | [1.0, 4.0]
In Confirm the Hessian test in code, given the rows so far: what is the next one — the row where quantity is all ≥ 0??
Correct: all ≥ 0? | True → PSD → convex everywhere
| quantity | value (verified) |
|---|---|
| Hessian ∇²f | [[1, 0], [0, 4]] (constant in w) |
| eigenvalues | [1.0, 4.0] |
| all ≥ 0? | True → PSD → convex everywhere |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The Hessian is constant, so this one eigenvalue check certifies convexity over the ENTIRE domain — no need to re-check at other points.
Worked example
Ask NumPy for the eigenvalues of the Hessian and read off the sign. eigvalsh is the symmetric-matrix eigenvalue routine — the right tool for a Hessian.
import numpy as np
H = np.array([[1., 0.],
[0., 4.]]) # Hessian of 0.5(w1^2 + 4 w2^2)
eig = np.linalg.eigvalsh(H)
print(eig) # [1. 4.]
print('convex everywhere?', np.all(eig >= 0))eig(H) = [1, 4], all ≥ 0 → PSD → convex
Why: The Hessian is constant, so this one eigenvalue check certifies convexity over the ENTIRE domain — no need to re-check at other points. Verified by execution.
| quantity | value (verified) |
|---|---|
| Hessian ∇²f | [[1, 0], [0, 4]] (constant in w) |
| eigenvalues | [1.0, 4.0] |
| all ≥ 0? | True → PSD → convex everywhere |
Comparison
Comparison matrix
From Confirm the Hessian test in code: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| quantity | value (verified) |
|---|---|
| Hessian ∇²f | [[1, 0], [0, 4]] (constant in w) |
| eigenvalues | [1.0, 4.0] |
| all ≥ 0? | True → PSD → convex everywhere |
Anomaly
Predict first
A student writes this, and it looks reasonable:
To test convexity, find the critical point and check the Hessian there. If ∇²f ⪰ 0 at that point, call f convex.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: A PSD Hessian at ONE point only certifies a LOCAL minimum.
Convexity is a global property. The Hessian must be PSD at every point in the domain, not just at the minimum.
Why: A PSD Hessian at ONE point only certifies a LOCAL minimum. The function is free to curve downward somewhere else and still pass this local test — so you can wrongly certify a non-convex function.
Trap
To test convexity, find the critical point and check the Hessian there. If ∇²f ⪰ 0 at that point, call f convex.
Test ∇²f ⪰ 0 only at the stationary point
Why: A PSD Hessian at ONE point only certifies a LOCAL minimum. The function is free to curve downward somewhere else and still pass this local test — so you can wrongly certify a non-convex function.
Convexity is a global property. The Hessian must be PSD at every point in the domain, not just at the minimum.
Verify ∇²f ⪰ 0 for ALL w
Why: For our bowl the Hessian is CONSTANT (= diag(1,4)), so one check covers all w. When the Hessian depends on w you must show it stays PSD throughout — as the next slide's counterexample makes painfully clear.
Section
Part 2 of 6 — a counterexample
Concept
To feel what convexity buys us, break it. Consider the 1-D function
\[ g(t) = t^4 - 3t^2 \]
It has a hump in the middle and two dips on the sides — a classic double well. Downhill from the left leads to one valley, downhill from the right to another. That is exactly the trap a convex function forbids.
Step zero
Discussion prompt
Its second derivative changes sign — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: First derivative
Answer:
Worked example
Differentiate g(t)=t⁴−3t² twice, one move per line:
First derivative
Why: g'(t) = 4t³ − 6t by the power rule, term by term.
\[ g'(t) = 4t^3 - 6t \]
Second derivative
Why: g''(t) = 12t² − 6, again by the power rule. This is the 1-D Hessian.
\[ g''(t) = 12t^2 - 6 \]
Find where g'' < 0
Why: 12t² − 6 < 0 ⇔ t² < ½ ⇔ |t| < 1/√2 ≈ 0.707. On that whole interval the curvature is NEGATIVE — the graph bends downward — so g is not convex.
\[ g''(t) < 0 \iff -\tfrac{1}{\sqrt2} < t < \tfrac{1}{\sqrt2} \]
Blank canvas
Draw it
Draw what Its second derivative changes sign just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Worked example
Tabulate g'' across the domain, then test the chord between t=−1.2 and t=+1.2 at the midpoint t=0. Runnable:
import numpy as np
def g(t): return t**4 - 3*t**2
def gpp(t): return 12*t**2 - 6 # second derivative
for t in [-1.5, -1., -0.5, 0., 0.5, 1., 1.5]:
print(t, gpp(t), 'convex' if gpp(t) >= 0 else 'CONCAVE')
mid = g(0.0); chord = 0.5*(g(-1.2) + g(1.2))
print('g(0) =', mid, ' chord =', round(chord,4), ' convex?', mid <= chord)g''< 0 on the middle rows; the chord test fails at t=0
Why: g(0) = 0 but the chord between t=±1.2 sits at −2.2464, which is BELOW 0. The surface rises above its chord — the definitional violation, matching the negative-curvature interval exactly.
| t | g''(t) = 12t²−6 | curvature | verdict |
|---|---|---|---|
| −1.5 | 21.0 | positive | convex here |
| −1.0 | 6.0 | positive | convex here |
| −0.5 | −3.0 | negative | CONCAVE |
| 0.0 | −6.0 | negative | CONCAVE |
| 0.5 | −3.0 | negative | CONCAVE |
| 1.0 | 6.0 | positive | convex here |
| 1.5 | 21.0 | positive | convex here |
Discrimination
Sort into buckets
Sort these by curvature, from memory, without looking back at Trace the curvature and catch the chord violation. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Set g'(t)=4t³−6t = 2t(2t²−3) = 0. The stationary points are t=0 and t=±√(3/2) ≈ ±1.2247.
| stationary t | g(t) | g''(t) | type |
|---|---|---|---|
| −1.2247 | −2.25 | 12 > 0 | local (global) minimum |
| 0 | 0 | −6 < 0 | local MAXIMUM (the hump) |
| +1.2247 | −2.25 | 12 > 0 | local (global) minimum |
Two separate minima. Gradient descent started at t=−0.3 slides left; started at t=+0.3 it slides right. Where you start decides where you land — the pathology convexity rules out.
Trade off
Comparison matrix
From Two minima — the trap made concrete: every row here is a choice with a cost. Fill the type column, then say which row you would actually pick and what you give up for it.
| stationary t | g(t) | g''(t) | type |
|---|---|---|---|
| −1.2247 | −2.25 | 12 > 0 | local (global) minimum |
| 0 | 0 | −6 < 0 | local MAXIMUM (the hump) |
| +1.2247 | −2.25 | 12 > 0 | local (global) minimum |
Section
Part 3 of 6 — the payoff
Intuition
Everything so far — chords, tangents, Hessians — was setup. Here is the payoff that makes convex optimization the well-behaved corner of machine learning:
On a convex function, a flat point is the best point. Find any w where the gradient is zero and you are done — it is the global minimum, not merely a local one. No restarts, no luck, no wondering if a better valley exists.
The next slide proves it in two lines using only the first-order (tangent-below) condition — the whole reason we bothered to state it.
Step zero
Discussion prompt
Prove: ∇f(x)=0 forces a global minimum — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Start from the first-order condition
Answer:
Worked example
This is the theorem that makes convex problems easy. Suppose f is convex and differentiable and x is a stationary point, ∇f(x)=0. We show f(x) is the global minimum in two lines.
Start from the first-order condition
Why: Convexity gives the tangent-below inequality at x, valid for EVERY y.
\[ f(y) \ge f(x) + \nabla f(x)^\top (y - x) \]
Substitute ∇f(x) = 0
Why: The stationary condition kills the entire linear term, since 0ᵀ(y−x) = 0.
\[ f(y) \ge f(x) + 0 = f(x) \]
Verify the conclusion holds for every y
Why: f(y) ≥ f(x) for ALL y means x is a GLOBAL minimum. No 'other valley' can be lower — the tangent inequality forbids it. This is why a convex problem has no local traps.
\[ \boxed{\,f(y) \ge f(x)\ \ \forall y \;\Longrightarrow\; x \text{ is a global minimizer}\,} \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Verify the conclusion holds for every y
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
This is the theorem that makes convex problems easy. Suppose f is convex and differentiable and x is a stationary point, ∇f(x)=0. We show f(x) is the global minimum in two lines.
Intuition
For a convex loss, any point where the gradient vanishes is already the best possible fit. You never have to worry 'is this just a local minimum?' — there is only one.
Linear regression, ridge, logistic regression, SVMs — all convex. That is why we can trust gradient descent to find the answer, and why Lesson 7's normal equations gave the unique global optimum.
Neural nets are not convex. Their loss is a landscape of many valleys, so initialization, momentum, and luck all matter. Convexity is the dividing line between 'solved' and 'an art'.
Sorting
Sort into buckets
These are the pieces of Lesson 11: Convex Optimization, out of order. Put each one back under the part of the lesson it belongs to.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This treats a convex bowl like a neural-net landscape.
A convex function has no local minimum other than the global one. Convergence is about speed, not about getting stuck.
Why: This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.
Trap
Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.
Restart MSE from many seeds to escape local minima
Why: This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.
A convex function has no local minimum other than the global one. Convergence is about speed, not about getting stuck.
Every stationary point of a convex f is global
Why: We just proved it: ∇f(x)=0 ⇒ f(y) ≥ f(x) for all y. So the only thing left to tune is the learning rate — which controls how FAST you reach the one minimum, the subject of Parts 4–5.
Break the constraint
Discussion prompt
The rule this trap just fixed:
We just proved it: ∇f(x)=0 ⇒ f(y) ≥ f(x) for all y. So the only thing left to tune is the learning rate — which controls how FAST you reach the one minimum, the subject of Parts 4–5.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This treats a convex bowl like a neural-net landscape. For a strictly convex loss there is EXACTLY one valley — restarts all converge to the same point. The extra runs waste compute and learn nothing.
Fill the middle
Fill in the blanks
From The MSE loss is convex — tying back to Lesson 7 — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1.,2.,3.,4.,5.]); y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
H = **2 * X.T @ X** # Hessian of ||Xw - y||^2
print(H) # [[ 10. 30.] [ 30. 110.]]
print(np.linalg.eigvalsh(H)) # [ 1.690481 118.309519]
Why: H is what everything below it consumes, so the wrong expression here fails later and somewhere else. zᵀ(2XᵀX)z = 2‖Xz‖² ≥ 0 for any z, so 2XᵀX is PSD by construction — and it is CONSTANT in w.
Worked example
The least-squares loss L(w)=‖Xw−y‖² from Lesson 7 is convex, and now we can prove it in one line: its Hessian is 2XᵀX, which is PSD. On the 5-student data:
import numpy as np
x = np.array([1.,2.,3.,4.,5.]); y = np.array([2.,4.,5.,4.,5.])
X = np.c_[np.ones(5), x]
H = 2 * X.T @ X # Hessian of ||Xw - y||^2
print(H) # [[ 10. 30.] [ 30. 110.]]
print(np.linalg.eigvalsh(H)) # [ 1.690481 118.309519]eig(2XᵀX) = [1.69, 118.31], both ≥ 0 → convex everywhere
Why: zᵀ(2XᵀX)z = 2‖Xz‖² ≥ 0 for any z, so 2XᵀX is PSD by construction — and it is CONSTANT in w. Convex everywhere ⇒ the normal-equation solution is the unique GLOBAL optimum. Verified by execution.
| quantity | value (verified) |
|---|---|
| Hessian | 2XᵀX = [[10, 30], [30, 110]] (constant) |
| eigenvalues | [1.6905, 118.3095] |
| convex? | yes — both ≥ 0, PSD everywhere |
Concept
Convex (∇²f ⪰ 0, eigenvalues ≥ 0) guarantees any minimum is global, but there could be a flat valley of equally-good minima. Strict convexity (∇²f ≻ 0, eigenvalues > 0) rules that out.
\[ \nabla^2 f \succ 0 \;\Longrightarrow\; \text{the global minimizer is unique} \]
Our bowl (eigenvalues 1, 4 > 0) and the full-rank MSE Hessian (1.69, 118.3 > 0) are both strictly convex — one point, not a valley. A rank-deficient X gives a zero eigenvalue: still convex, but with infinitely many optimal w — exactly the collinearity case ridge fixed in Lesson 7.
Ranking
Put in order
Put the moves of Prove 2XᵀX is PSD without eigenvalues into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Pull the constant 2 out front and group the middle: zᵀXᵀXz = (Xz)ᵀ(Xz).
Worked example
The eigenvalue check is numeric; here is the airtight algebraic reason 2XᵀX ⪰ 0 for any X. Test the PSD definition zᵀHz ≥ 0 directly.
Write out zᵀ(2XᵀX)z
Why: Pull the constant 2 out front and group the middle: zᵀXᵀXz = (Xz)ᵀ(Xz).
\[ z^\top (2 X^\top X)\, z = 2\,(Xz)^\top (Xz) \]
Recognize a squared norm
Why: (Xz)ᵀ(Xz) is the dot product of the vector Xz with itself — that is ‖Xz‖², which can never be negative.
\[ = 2\,\lVert Xz \rVert^2 \;\ge\; 0 \quad \forall z \]
Verify against the numbers
Why: zᵀHz ≥ 0 for all z is precisely the PSD definition — so 2XᵀX is convex-certifying for any data X, matching the eigenvalues [1.69, 118.31] ≥ 0 we computed. No noise, no exceptions.
| step | expression | sign |
|---|---|---|
| group | zᵀ(2XᵀX)z = 2(Xz)ᵀ(Xz) | regrouped |
| norm | = 2‖Xz‖² | square |
| conclude | ≥ 0 for every z | PSD ✓ |
Translation
\( = 2\,\lVert Xz \rVert^2 \;\ge\; 0 \quad \forall z \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Section
Part 4 of 6 — the dynamics
Intuition
For our quadratic we can solve ∇f=0 in closed form — that is exactly the normal equations of Lesson 7. So why iterate at all?
Because for most losses ∇f(w)=0 is a giant nonlinear system with no closed form. Logistic regression, neural nets — you cannot write the answer down. The only tool that scales is: compute the gradient, take a step downhill, repeat.
Convexity is what makes that humble loop trustworthy: on a convex loss, following the gradient down always arrives at the one global minimum. We study the loop on the quadratic because there we know the true answer and can check the iterates against it.
Concept
Gradient descent takes repeated downhill steps. From the current point wₜ, step against the gradient by a learning rate η > 0:
\[ w_{t+1} = w_t - \eta\,\nabla f(w_t) \]
The gradient points uphill (steepest ascent), so the minus sign heads downhill. η sets the stride. The entire question of Parts 4–5 is: how big can η be before the strides overshoot and blow up?
Concept
Why step against ∇f specifically? Among all unit directions u, the directional derivative ∇f(w)ᵀu — the instantaneous slope — is largest when u points along ∇f, and most negative when u = −∇f.
\[ \min_{\lVert u\rVert = 1} \nabla f(w)^\top u = -\lVert \nabla f(w)\rVert \;\text{ at }\; u = -\frac{\nabla f(w)}{\lVert \nabla f(w)\rVert} \]
So −∇f is the locally fastest way down — the greediest possible descent direction. Gradient descent is just 'walk the steepest way down, one stride at a time'.
Step zero
Discussion prompt
The 1-D shadow: derive the update factor — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Plug f'(x)=2x into the update
Answer:
Worked example
Use the clean shadow f(x)=x². Its gradient is f'(x)=2x. Substitute into the update and simplify — this reveals GD as pure multiplication.
Plug f'(x)=2x into the update
Why: The general rule x_{t+1} = x_t − η f'(x_t) becomes x_{t+1} = x_t − η·2x_t.
\[ x_{t+1} = x_t - \eta\,(2 x_t) \]
Factor out xₜ
Why: Both terms carry x_t, so pull it out. Each step just MULTIPLIES the current point by the fixed number (1 − 2η).
\[ x_{t+1} = (1 - 2\eta)\, x_t \]
Unroll to a closed form
Why: Applying the same factor t times gives a geometric sequence. Whether it shrinks or grows is decided entirely by |1 − 2η|.
\[ x_t = (1 - 2\eta)^t\, x_0 \]
Blank canvas
Draw it
Draw what The 1-D shadow: derive the update factor just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
xₜ = (1−2η)ᵗ x₀ is a geometric sequence. Its fate hinges on one number, the contraction factor |1 − 2η|:
|1 − 2η| < 1 → each step shrinks |x| → converges to 0|1 − 2η| = 1 → |x| never changes → stalls (or oscillates forever)|1 − 2η| > 1 → each step grows |x| → divergesSolve |1 − 2η| < 1: it means −1 < 1 − 2η < 1, i.e. 0 < η < 1. So for f(x)=x², gradient descent converges exactly when 0 < η < 1. Keep that interval — the next slide watches it break.
Estimation
Predict first
Start at x₀=2 and run six steps for η ∈ {0.1, 0.5, 0.9, 1.1}. Watch the contraction factor decide each fate. Runnable:
Commit before you compute: what does Four learning rates on f(x)=x² come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: η=0.5 lands at 0 in ONE step; η=1.1 blows up
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. η=0.5 gives factor 0 — the perfect step for this f.
Worked example
Start at x₀=2 and run six steps for η ∈ {0.1, 0.5, 0.9, 1.1}. Watch the contraction factor decide each fate. Runnable:
for lr in [0.1, 0.5, 0.9, 1.1]:
x = 2.0; traj = [x]
for _ in range(6):
x = x - lr*(2*x) # gradient step, grad = 2x
traj.append(round(x, 4))
print(lr, 'factor', round(1-2*lr,2), traj)η=0.5 lands at 0 in ONE step; η=1.1 blows up
Why: η=0.5 gives factor 0 — the perfect step for this f. η=0.9 gives factor −0.8, so it converges but flips sign each step. η=1.1 gives factor −1.2 (>1 in magnitude), so |x| grows every step: divergence.
| η | factor 1−2η | x trajectory from 2.0 | fate |
|---|---|---|---|
| 0.1 | +0.80 | 2 → 1.6 → 1.28 → 1.024 → 0.819 → 0.655 → 0.524 | converges (slow) |
| 0.5 | 0.00 | 2 → 0 → 0 → 0 → 0 → 0 → 0 | converges (1 step) |
| 0.9 | −0.80 | 2 → −1.6 → 1.28 → −1.024 → 0.819 → −0.655 → 0.524 | converges (oscillating) |
| 1.1 | −1.20 | 2 → −2.4 → 2.88 → −3.456 → 4.147 → −4.977 → 5.972 | DIVERGES |
Fill the middle
Fill in the blanks
From Full trace: loss and gradient at η = 0.1 — one line has had its right-hand side removed. Put it back.
x, lr = 2.0, 0.1
for t in range(6):
print(t, round(x,5), round(xx,6), round(2x,5)) # t, x, f=x^2, grad=2x
x = *x - lr(2x)* # step downhill
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Near the minimum the gradient fades, so each step gets smaller on its own — GD naturally decelerates as it homes in.
Worked example
Zoom in on the safe run η=0.1 from x₀=2 and log every quantity — position, loss, and gradient — pass by pass. This is the trace to internalize.
x, lr = 2.0, 0.1
for t in range(6):
print(t, round(x,5), round(x*x,6), round(2*x,5)) # t, x, f=x^2, grad=2x
x = x - lr*(2*x) # step downhillAs x → 0, both the loss x² and the gradient 2x shrink toward 0
Why: Near the minimum the gradient fades, so each step gets smaller on its own — GD naturally decelerates as it homes in. The loss falls by the factor (1−2η)² = 0.64 per step, from 4.0 down.
| step t | x | f(x) = x² | gradient 2x |
|---|---|---|---|
| 0 | 2.00000 | 4.000000 | 4.00000 |
| 1 | 1.60000 | 2.560000 | 3.20000 |
| 2 | 1.28000 | 1.638400 | 2.56000 |
| 3 | 1.02400 | 1.048576 | 2.04800 |
| 4 | 0.81920 | 0.671089 | 1.63840 |
| 5 | 0.65536 | 0.429497 | 1.31072 |
Pattern
Step through it
Step through Full trace: loss and gradient at η = 0.1 one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
Bigger steps cover more ground, so cranking η up must reach the minimum sooner. Use η = 1.1 on f(x)=x².
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound.
There is a ceiling. Stay inside 0 < η < 1 (here), and the fastest step is the one that makes the factor smallest — η = 0.5 (factor 0).
Why: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound. Bigger is not faster — bigger is broken.
Trap
Bigger steps cover more ground, so cranking η up must reach the minimum sooner. Use η = 1.1 on f(x)=x².
η = 1.1 → factor |1−2η| = 1.2 > 1
Why: Above the stability threshold each step OVERSHOOTS the minimum by more than it started: 2 → −2.4 → 2.88 → −3.456 → … The iterate grows without bound. Bigger is not faster — bigger is broken.
There is a ceiling. Stay inside 0 < η < 1 (here), and the fastest step is the one that makes the factor smallest — η = 0.5 (factor 0).
η = 0.5 → factor 0 → one-step convergence
Why: Within the stable range the error contracts every step; right at η = 1/L the factor hits its minimum. The 'sweet spot' is a balance the smoothness constant L pins down exactly — derived next.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
w = (w₁, w₂):; Picture the graph of a function as a landscape. Convex means it is a single smooth bowl: no separate valleys, no dents, no hidden pockets to fall into.; Flip the sag upward and you get the non-convex double well of Part 2, where a chord can dip under the curve. One picture, two fates.∇²f ⪰ 0 at that point, call f convex.; Like any function, a convex one might have local minima that trap gradient descent, so we still have to worry about initialization on linear regression.Section
Part 5 of 6 — smoothness & conditioning
Concept
L-smoothness — A function is L-smooth if its gradient changes no faster than L: ‖∇f(a) − ∇f(b)‖ ≤ L‖a − b‖. For a twice-differentiable f this L is the LARGEST eigenvalue of the Hessian — the steepest curvature anywhere.
L is the tightest speed limit on the gradient. Along the steepest-curvature direction, the same step size travels the most — so L is what caps how big a step you can safely take.
\[ L = \lambda_{\max}\!\big(\nabla^2 f\big) \]
Explain it
Discussion prompt
Explain Smoothness L bounds the curvature to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
L is the tightest speed limit on the gradient. Along the steepest-curvature direction, the same step size travels the most — so L is what caps how big a step you can safely take.
Ranking
Put in order
Put the moves of Derive the stable-step bound η < 2/L into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Convergence in a direction with curvature λ needs |1 − ηλ| < 1, exactly the 1-D condition applied per eigen-direction.
Worked example
On a quadratic with Hessian eigenvalue λ in some direction, GD multiplies the error in that direction by (1 − ηλ) each step — the vector version of the (1−2η) we found. Stability needs that factor below 1 in magnitude.
Require the contraction factor < 1 in magnitude
Why: Convergence in a direction with curvature λ needs |1 − ηλ| < 1, exactly the 1-D condition applied per eigen-direction.
\[ |1 - \eta \lambda| < 1 \]
Unfold the absolute value
Why: |1 − ηλ| < 1 means −1 < 1 − ηλ < 1. The right inequality gives ηλ > 0 (always true for η,λ > 0); the left gives ηλ < 2.
\[ 0 < \eta \lambda < 2 \;\Longrightarrow\; \eta < \frac{2}{\lambda} \]
Verify: the tightest direction is λ = L
Why: The bound must hold for EVERY eigen-direction at once, so the smallest ceiling wins — the one from the largest λ, namely L. Hence the single stable-step rule below.
\[ \boxed{\,0 < \eta < \frac{2}{L}, \qquad L = \lambda_{\max}(\nabla^2 f)\,} \]
Notation
Annotate
From Derive the stable-step bound η < 2/L — read this one piece at a time. What is each part doing?
On: \( \boxed{\,0 < \eta < \frac{2}{L}, \qquad L = \lambda_{\max}(\nabla^2 f)\,} \)
Pattern
Predict first
The table runs: 0.2 | [1.049, 0.000] | +0.20 | converges · 0.5 | [0.062, 1.000] | −1.00 | w₂ stalls (|factor|=1)
In The bound on our 2-D bowl, given the rows so far: what is the next one — the row where η is 0.6?
Correct: 0.6 | [0.016, 7.530] | −1.40 | DIVERGES in w₂
| η | w after 6 steps | w₂ factor 1−4η | fate |
|---|---|---|---|
| 0.2 | [1.049, 0.000] | +0.20 | converges |
| 0.5 | [0.062, 1.000] | −1.00 | w₂ stalls (|factor|=1) |
| 0.6 | [0.016, 7.530] | −1.40 | DIVERGES in w₂ |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The w₂ curvature is L=4, so its factor is 1−0.6·4 = −1.4, magnitude > 1: w₂ explodes (1 → −1.4 → 1.96 → …) even though w₁'s factor 1−0.6·1 = 0.4 is fine.
Worked example
Back to f(w)=½(w₁²+4w₂²), Hessian diag(1,4), so L=4, μ=1, and the bound is η < 2/4 = 0.5. Run η=0.2 (safe), η=0.5 (edge), η=0.6 (over). Runnable:
import numpy as np
def grad(w): return np.array([w[0], 4*w[1]]) # Hessian diag(1,4)
for eta in [0.2, 0.5, 0.6]:
w = np.array([4., 1.])
for _ in range(6):
w = w - eta*grad(w)
print(eta, np.round(w, 3), 'w2 factor', round(1-eta*4, 2))η=0.6 > 2/L=0.5 → the w₂ direction diverges
Why: The w₂ curvature is L=4, so its factor is 1−0.6·4 = −1.4, magnitude > 1: w₂ explodes (1 → −1.4 → 1.96 → …) even though w₁'s factor 1−0.6·1 = 0.4 is fine. ONE bad direction dooms the whole run.
| η | w after 6 steps | w₂ factor 1−4η | fate |
|---|---|---|---|
| 0.2 | [1.049, 0.000] | +0.20 | converges |
| 0.5 | [0.062, 1.000] | −1.00 | w₂ stalls (|factor|=1) |
| 0.6 | [0.016, 7.530] | −1.40 | DIVERGES in w₂ |
Concept
The stable step is capped by the largest curvature L, but progress along the smallest-curvature direction μ = λ_min is what's slow. The gap between them is the trouble.
condition number κ — κ = L/μ = λ_max / λ_min of the Hessian. κ=1 is a perfectly round bowl (GD converges in ~1 step); large κ is a long narrow ravine where GD zig-zags. Convex GD needs about O(κ log 1/ε) steps.
Our bowl has κ = 4/1 = 4 — mild. The Lesson 7 MSE Hessian had κ ≈ 118.3/1.69 ≈ 70 — the SAME number that warned us off inverting XᵀX. Big κ is slow GD and fragile linear algebra; rescaling features fixes both.
Analogy
Discussion prompt
Explain Conditioning κ = L/μ sets the pace by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The stable step is capped by the largest curvature L, but progress along the smallest-curvature direction μ = λ_min is what's slow. The gap between them is the trouble.
Missing information
Discussion prompt
When μ > 0 (a curvature floor — 'strongly convex'), the loss gap shrinks by a constant ratio every step. On our bowl at η = 1/L = 0.25, the per-step factor settles to (1 − μ/L)² = (1 − 0.25)² = 0.5625. Runnable:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Geometric (a.k.a. LINEAR) convergence: the gap is multiplied by a constant < 1 each step, so it reaches ε in O(log 1/ε) steps. The w₂ term dies instantly (factor 0), leaving the slow w₁ direction to set the 0.5625 rate. Verified by execution.
Worked example
When μ > 0 (a curvature floor — 'strongly convex'), the loss gap shrinks by a constant ratio every step. On our bowl at η = 1/L = 0.25, the per-step factor settles to (1 − μ/L)² = (1 − 0.25)² = 0.5625. Runnable:
import numpy as np
def f(w): return 0.5*(w[0]**2 + 4*w[1]**2)
def grad(w): return np.array([w[0], 4*w[1]])
w = np.array([4., 1.]); eta = 0.25; prev = None
for t in range(6):
if prev: print(t, round(f(w),6), 'ratio', round(f(w)/prev,4))
prev = f(w); w = w - eta*grad(w)The loss ratio locks onto 0.5625 = (1−μ/L)²
Why: Geometric (a.k.a. LINEAR) convergence: the gap is multiplied by a constant < 1 each step, so it reaches ε in O(log 1/ε) steps. The w₂ term dies instantly (factor 0), leaving the slow w₁ direction to set the 0.5625 rate. Verified by execution.
| step t | f(wₜ) | ratio f(wₜ)/f(wₜ₋₁) |
|---|---|---|
| 1 | 4.500000 | 0.4500 |
| 2 | 2.531250 | 0.5625 |
| 3 | 1.423828 | 0.5625 |
| 4 | 0.800903 | 0.5625 |
| 5 | 0.450508 | 0.5625 |
Pattern
Step through it
Step through Strong convexity gives geometric decay one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
The rate classes sound abstract until you price them. O(1/t) (sublinear) means to halve the error you roughly double the steps — get to 1e-3, and 1e-6 costs a thousand times more work. Painful for high precision.
O(cᵗ) (linear/geometric) means each fixed block of steps multiplies the error by a constant — so reaching 1e-6 costs only about twice the steps of reaching 1e-3. That is the prize strong convexity (μ>0) hands you, and why our 0.5625-ratio bowl felt fast.
Concept
How fast GD closes the gap to the minimum depends on the curvature class. Memorize this table — it is exam bread and butter:
| objective class | GD error after t steps | meaning |
|---|---|---|
| convex + L-smooth | O(1/t) | sublinear — halving error costs ~2× steps |
| strongly convex (μ>0) | O(cᵗ), c<1 | linear/geometric — like our 0.5625 ratio |
| convex + Nesterov momentum | O(1/t²) | accelerated — the optimal first-order rate |
A curvature floor μ>0 upgrades O(1/t) to exponential decay. Momentum buys the 1/t² rate without needing that floor — it's the free lunch of Week 6's optimizers.
Comparison
Comparison matrix
From The convergence-rate classes: refill the meaning column from what you know. The rest of the table is as it appeared.
| objective class | GD error after t steps | meaning |
|---|---|---|
| convex + L-smooth | O(1/t) | sublinear — halving error costs ~2× steps |
| strongly convex (μ>0) | O(cᵗ), c<1 | linear/geometric — like our 0.5625 ratio |
| convex + Nesterov momentum | O(1/t²) | accelerated — the optimal first-order rate |
Intuition
Picture a long narrow valley — steep walls (large L), gentle floor (small μ). Plain GD bounces across the steep walls while barely inching along the floor toward the minimum. That is the κ=L/μ slowdown made visual.
Momentum fixes it by accumulating velocity: v ← βv − η∇f, then w ← w + v. The back-and-forth wall-bounces cancel while the steady floor-ward motion builds up — so you glide down the valley instead of rattling across it.
Estimation
Predict first
Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:
Commit before you compute: what does Momentum beats plain GD on our bowl come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Momentum reaches 1e-8 in 12 steps vs plain GD's 21
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Heavy-ball with η=0.4444, β=0.1111 nearly halves the step count on this mild κ=4 bowl.
Worked example
Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:
import numpy as np
f = lambda w: 0.5*(w[0]**2 + 4*w[1]**2)
grad = lambda w: np.array([w[0], 4.*w[1]])
L, mu = 4., 1.
def plain():
w = np.array([4.,1.]); eta = 2/(L+mu)
k = 0
while f(w) >= 1e-8: w = w - eta*grad(w); k += 1
return k
def heavy_ball():
w = np.array([4.,1.]); v = np.zeros(2)
eta = 4/((np.sqrt(L)+np.sqrt(mu))**2) # 0.4444
beta = ((np.sqrt(L/mu)-1)/(np.sqrt(L/mu)+1))**2 # 0.1111
k = 0
while f(w) >= 1e-8: v = beta*v - eta*grad(w); w = w + v; k += 1
return k
print('plain GD steps:', plain(), ' heavy-ball steps:', heavy_ball())Momentum reaches 1e-8 in 12 steps vs plain GD's 21
Why: Heavy-ball with η=0.4444, β=0.1111 nearly halves the step count on this mild κ=4 bowl. The bigger κ is, the larger the gap — momentum's O(√κ) beats plain GD's O(κ). Verified by execution.
| method | params | steps to loss < 1e-8 |
|---|---|---|
| plain GD | η = 2/(L+μ) = 0.4 | 21 |
| heavy-ball momentum | η = 0.4444, β = 0.1111 | 12 |
| momentum trace f(w) | 10 → 3.68 → 0.87 → 0.17 → … | geometric, faster ratio |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Momentum reaches 1e-8 in 12 steps vs plain GD's 21
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Compare plain GD against heavy-ball momentum on f(w)=½(w₁²+4w₂²), both from w=(4,1), counting steps to drive the loss below 1e-8. Runnable:
Constraint
Discussion prompt
Run Diagnosing any optimization problem with this step confiscated:
Conditioning κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.Pattern
L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.Edge cases
Discussion prompt
Diagnosing any optimization problem works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
L: the largest Hessian eigenvalue → keep η < 2/L; the sweet spot is near 1/L.κ = L/μ: large κ ⇒ slow zig-zagging → rescale features or precondition to shrink it.η is above 2/L, lower it. Crawling? η too small or κ too large → raise η or add momentum.Elimination
Eliminate the wrong options
A twice-differentiable f is convex if and only if its Hessian is…
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Convexity is equivalent to ∇²f ⪰ 0 everywhere — non-negative curvature in every direction at every point. Then every stationary point is a global minimum. For a constant Hessian (like diag(1,4)) one eigenvalue check settles it for all w.
Check
Which condition proves convexity for a twice-differentiable f?
Check your understanding
A twice-differentiable f is convex if and only if its Hessian is…
Answer: A
Why: Convexity is equivalent to ∇²f ⪰ 0 everywhere — non-negative curvature in every direction at every point. Then every stationary point is a global minimum. For a constant Hessian (like diag(1,4)) one eigenvalue check settles it for all w.
Prediction
Predict first
Gradient descent is run on a strictly convex loss from two different starting points (both with a stable learning rate). What happens?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Both converge to the same global minimum
Why: A strictly convex function has a single global minimum and no other stationary points, so any convergent run reaches the same point regardless of initialization — the ∇f(x)=0 ⇒ global proof guarantees it.
Check
What does convexity guarantee about where you land?
Check your understanding
Gradient descent is run on a strictly convex loss from two different starting points (both with a stable learning rate). What happens?
Answer: A
Why: A strictly convex function has a single global minimum and no other stationary points, so any convergent run reaches the same point regardless of initialization — the ∇f(x)=0 ⇒ global proof guarantees it.
Prediction
Predict first
For f(x) = x² (so f''=2, hence L=2), which learning rate makes gradient descent DIVERGE?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: η = 1.1
Why: The update multiplies x by (1−2η). Stability needs |1−2η| < 1, i.e. 0 < η < 1 = 2/L. η=1.1 gives factor −1.2, magnitude > 1, so |x| grows each step: 2 → −2.4 → 2.88 → … — divergence.
Check
Use the smoothness bound η < 2/L.
Check your understanding
For f(x) = x² (so f''=2, hence L=2), which learning rate makes gradient descent DIVERGE?
Answer: A
Why: The update multiplies x by (1−2η). Stability needs |1−2η| < 1, i.e. 0 < η < 1 = 2/L. η=1.1 gives factor −1.2, magnitude > 1, so |x| grows each step: 2 → −2.4 → 2.88 → … — divergence.
Elimination
Eliminate the wrong options
A convex quadratic has Hessian eigenvalues μ=1 and L=100. Its learning rate must satisfy η<2/L=0.02, capped by the STEEP direction. What limits convergence speed?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: η is capped at ~1/L by the steep direction, but that same small η makes progress in the shallow μ direction glacial: its factor is 1−ημ ≈ 1−μ/L = 0.99. The condition number κ = L/μ = 100 sets the pace — O(κ log 1/ε) steps.
Check
Think about which eigenvalue caps the step and which one is slow.
Check your understanding
A convex quadratic has Hessian eigenvalues μ=1 and L=100. Its learning rate must satisfy η<2/L=0.02, capped by the STEEP direction. What limits convergence speed?
Answer: A
Why: η is capped at ~1/L by the steep direction, but that same small η makes progress in the shallow μ direction glacial: its factor is 1−ημ ≈ 1−μ/L = 0.99. The condition number κ = L/μ = 100 sets the pace — O(κ log 1/ε) steps.
Section
Part 6 of 6 — the project
Concept
Run gradient descent on f(x)=x² for several learning rates, print each full trajectory, and classify every run as converging or diverging directly from its contraction factor |1−2η|. You derived every piece — now assemble it.
| # | requirement | tool |
|---|---|---|
| 1 | Define f and its gradient | f = x², grad = 2x |
| 2 | A GD loop that returns the trajectory | a for-loop, append each x |
| 3 | Sweep η and flag |1−2η| ≥ 1 as diverging | compare the factor |
Build rules: type every line yourself, print the whole trajectory (don't just trust the final number), and watch the sign flip when η pushes past the stable range.
Counterexample
Discussion prompt
Build rules: type every line yourself, print the whole trajectory (don't just trust the final number), and watch the sign flip when η pushes past the stable range.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: write f and grad. Say aloud what grad(2) should be before you print it.
Hint: f = x**2, grad = 2*x — the same shape as Lesson 6's warm-up. The gradient at x=2 is 2·2 = 4.
def f(x): return x**2
def grad(x): return 2*x
print(f(2), grad(2))| call | value (verified) |
|---|---|
| f(2) | 4 |
| grad(2) | 4 |
| grad(0) (the minimum) | 0 |
Worked example
Your turn: from x=2, step six times at η=0.1 and collect the trajectory. Predict whether it reaches 0, and how fast.
Hint: x = x - lr*grad(x) inside the loop; append x each pass. The factor is 1−2·0.1 = 0.8, so each step keeps 80%.
def grad(x): return 2*x
x, lr = 2.0, 0.1
traj = [x]
for _ in range(6):
x = x - lr*grad(x)
traj.append(round(x, 4))
print(traj)| step | x | f(x)=x² |
|---|---|---|
| 0 | 2.0000 | 4.0000 |
| 1 | 1.6000 | 2.5600 |
| 2 | 1.2800 | 1.6384 |
| 3 | 1.0240 | 1.0486 |
| 6 | 0.5243 | 0.2749 |
Pattern
Step through it
Step through Milestone 2 — the GD loop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: loop over η ∈ {0.1, 0.5, 1.1} and label each by its contraction factor |1−2η|. Predict which one diverges before you run it.
Hint: factor = abs(1 - 2*lr), then 'diverges' if factor >= 1 else 'converges'. Only η=1.1 should cross the line.
for lr in [0.1, 0.5, 1.1]:
factor = abs(1 - 2*lr)
tag = 'diverges' if factor >= 1 else 'converges'
print(lr, round(factor, 2), tag)| η | |1−2η| | verdict (verified) |
|---|---|---|
| 0.1 | 0.80 | converges |
| 0.5 | 0.00 | converges (1 step) |
| 1.1 | 1.20 | diverges |
Worked example
Your turn: fold the loop and the classifier into one function run(lr) that returns both the trajectory and the factor, then sweep. This is the whole experiment in one screen.
Hint: return (traj, abs(1-2*lr)) from run, and tag each sweep line by the factor.
def f(x): return x**2
def grad(x): return 2*x
def run(lr, x0=2.0, steps=6):
x, traj = x0, [round(x0,4)]
for _ in range(steps):
x = x - lr*grad(x)
traj.append(round(x, 4))
return traj, abs(1 - 2*lr)
for lr in [0.1, 0.5, 1.1]:
traj, factor = run(lr)
tag = 'diverges' if factor >= 1 else 'converges'
print(lr, tag, traj)| η | verdict | trajectory ends near |
|---|---|---|
| 0.1 | converges | 0.5243 (heading to 0) |
| 0.5 | converges | 0.0 (one step) |
| 1.1 | diverges | 5.972 and growing |
Trade off
Comparison matrix
From Milestone 4 — the full program: every row here is a choice with a cost. Fill the verdict column, then say which row you would actually pick and what you give up for it.
| η | verdict | trajectory ends near |
|---|---|---|
| 0.1 | converges | 0.5243 (heading to 0) |
| 0.5 | converges | 0.0 (one step) |
| 1.1 | diverges | 5.972 and growing |
Concept
Swap the scalar f for the 2-D bowl f(w)=½(w₁²+4w₂²) with grad = [w₁, 4w₂], then run η=0.6. Because 2/L = 0.5, the steep w₂ direction diverges while w₁ behaves — the single-bad-direction failure from Part 5.
import numpy as np
def grad2(w): return np.array([w[0], 4*w[1]])
w = np.array([4., 1.])
for t in range(6):
w = w - 0.6*grad2(w)
print(t, np.round(w, 3))| step | w = [w₁, w₂] | note |
|---|---|---|
| 1 | [1.600, −1.400] | w₂ already overshot past 0 |
| 2 | [0.640, 1.960] | w₂ growing, sign flipping |
| 4 | [0.102, 3.842] | w₂ diverging (factor −1.4) |
| 6 | [0.016, 7.530] | w₁ fine, w₂ blown up |
Comparison
Comparison matrix
From Stretch: watch the 2-D bound bite: refill the w = [w₁, w₂] column from what you know. The rest of the table is as it appeared.
| step | w = [w₁, w₂] | note |
|---|---|---|
| 1 | [1.600, −1.400] | w₂ already overshot past 0 |
| 2 | [0.640, 1.960] | w₂ growing, sign flipping |
| 4 | [0.102, 3.842] | w₂ diverging (factor −1.4) |
| 6 | [0.016, 7.530] | w₁ fine, w₂ blown up |
Concept
Slides closed, out loud: explain (1) why a convex function has no local-minima traps, using the tangent inequality; (2) what L is and where 2/L comes from; and (3) why the MSE loss is convex but a neural net's loss is not.
Stretch homework: prove the sum of two convex functions is convex (add their chord inequalities), and prove MSE is convex via the second-order condition (∇² = 2XᵀX ⪰ 0). Both feed straight into Lesson 13's PSD machinery and Week 6's SGD and momentum variants.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What makes a problem convex · When convexity fails · Why convex ⇒ only global minima · Gradient descent as a contraction · The step-size bound η < 2/L · Your turn: GD convergence. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
g(t)=t⁴−3t² has g''<0 on |t|<1/√2, two minima at ±√1.5, and a chord violation at t=0∇f(x)=0 ⇒ f(y)≥f(x) everywhere for convex f — every stationary point is global, no trapsx_{t}=(1−2η)^t x₀ and read the stable-step bound η < 2/L, L = λ_max(∇²f)κ=L/μ, and quote the O(1/t) / linear / O(1/t²) rate classes| idea | the one thing to remember |
|---|---|
| convex test | Hessian PSD EVERYWHERE, not just at the min |
| minima | convex ⇒ every stationary point is global |
| GD step | x_{t+1} = (1−2η)x_t on x²; factor |1−2η| decides all |
| step size | η < 2/L, L = largest Hessian eigenvalue |
| speed | κ = L/μ large ⇒ slow; μ>0 ⇒ geometric decay |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.