USAAIO Lesson 37, from Week 13, a fully worked Phase 1 consolidation built around ONE running dataset, the five students from Lesson 7. OLS is derived and verified FOUR equivalent ways - by the normal equations, by Gaussian MLE, by the pseudoinverse and SVD, and by hat-matrix projection - with every step shown. Each of the four pillars is then re-derived on that same data: eigenvalues and PSD matrices, the SVD singular values, softmax with cross-entropy giving dL/dz = p - y, KL divergence and cross-entropy, bias and variance, ROC against PR-AUC under imbalance, gradient descent and Newton's method both converging to the closed form, and ridge regression lifting the zero eigenvalue. Every snippet runs self-contained, and every number was produced by real execution. The lesson runs to 62 slides.
Subject: Machine Learning · 112 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 37 · Week 13
Twelve weeks, one coherent foundation. We prove it by fitting one line to five students four different ways — normal equations, MLE, pseudoinverse, projection — then re-derive every pillar on that same data, all verified against real execution.
Objectives
w = [2.2, 0.6]w, and know when GD divergesWarm-up
Discussion prompt
Before we open Lesson 37: Phase 1 Review & Mock Prep: without looking back, what was the main idea of Initialization & Gradient Clipping, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
exploding gradients and norm clipping, Xavier and He initialization derived from variance propagation, why He init keeps ReLU activations stable through deep networks, and why zero initialization fails. Run a 50-layer signal-propagation experiment and compare initializations.
Section
Part 1 of 5 — the synthesis capstone
Concept
Everything today runs on Lesson 7's five students: hours studied x and score y. We fit one line ŷ = w₀ + w₁·x and reach the SAME w from four Phase-1 directions.
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Design matrix X = [1, x] (a bias column then the feature), target y. Keep these five numbers in your head — every slide reuses them.
Comparison
Comparison matrix
From The running dataset: refill the hours x column from what you know. The rest of the table is as it appeared.
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Intuition
Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.
Geometry says project y onto the column space. Probability says maximize the Gaussian likelihood. Linear algebra says apply the pseudoinverse. Optimization says set the gradient to zero. They must agree — and today we watch them agree to the fourth decimal.
If you can explain the equivalence out loud, you understand Phase 1. That is the mock-exam bar.
Counterexample
Discussion prompt
Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
If you can explain the equivalence out loud, you understand Phase 1. That is the mock-exam bar.
Ranking
Put in order
Put the moves of Route 1 — the normal equations into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Five equations, two unknowns — no exact solution.
Worked example
Start from the overdetermined system Xw = y
Why: Five equations, two unknowns — no exact solution. Least squares finds the closest w.
\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix},\quad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \]
Left-multiply both sides by Xᵀ
Why: Xᵀ is (2×5), so XᵀX is a square (2×2) system — the normal equations from Lesson 7.
\[ \boxed{\,X^\top X\, w = X^\top y\,} \]
Assemble the entries by hand
Why: XᵀX top-left = Σ1 = 5, off-diagonal = Σx = 15, bottom-right = Σx² = 55; Xᵀy = [Σy, Σxy] = [20, 66].
\[ \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} w = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]
Verify: solve gives w = [2.2, 0.6]
Why: det = 5·55 − 15² = 50 ≠ 0, so the (2×2) is invertible and w is unique. Real execution below confirms it.
Notation
Annotate
From Route 1 — the normal equations — read this one piece at a time. What is each part doing?
On: \( X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix},\quad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \)
Step zero
Discussion prompt
Route 1 — XᵀX and Xᵀy, entry by entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Top-left = ones·ones = Σ1 = 5
Answer:
Worked example
Before trusting the solver, build the two objects by hand. (XᵀX)ⱼₖ is column j dotted with column k; column 0 is ones, column 1 is x.
Top-left = ones·ones = Σ1 = 5
Why: Five 1s dotted with five 1s. This entry is just n, the number of data points.
Off-diagonal = ones·x = Σx = 1+2+3+4+5 = 15
Why: The ones vector dotted with x sums x; symmetry puts 15 in both off-diagonal slots.
Bottom-right = x·x = Σx² = 1+4+9+16+25 = 55; and Xᵀy = [Σy, Σxy] = [20, 66]
Why: Σxy = 1·2 + 2·4 + 3·5 + 4·4 + 5·5 = 2+8+15+16+25 = 66. Verified by execution below.
\[ X^\top X = \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix},\quad X^\top y = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Bottom-right = x·x = Σx² = 1+4+9+16+25 = 55; and Xᵀy = [Σy, Σxy] = [20, 66]
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Before trusting the solver, build the two objects by hand. (XᵀX)ⱼₖ is column j dotted with column k; column 0 is ones, column 1 is x.
Estimation
Predict first
Self-contained: re-import NumPy, re-define x, y, X, then solve XᵀX w = Xᵀy. Never invert — solve factors the system.
Commit before you compute: what does Route 1 in code — solve the normal equations come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: solve returns [2.2, 0.6]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Line 5 builds XᵀX and Xᵀy and factors the system in one call.
Worked example
Self-contained: re-import NumPy, re-define x, y, X, then solve XᵀX w = Xᵀy. Never invert — solve factors the system.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x] # bias column + feature
w = np.linalg.solve(X.T @ X, X.T @ y)
print(X.T @ X) # [[ 5 15] [15 55]]
print(X.T @ y) # [20 66]
print(w) # [2.2 0.6]solve returns [2.2, 0.6]
Why: Line 5 builds XᵀX and Xᵀy and factors the system in one call. Verified: intercept 2.2, slope 0.6.
| object | value (verified) |
|---|---|
| XᵀX | [[5, 15], [15, 55]] |
| Xᵀy | [20, 66] |
| w = [w₀, w₁] | [2.2, 0.6] |
Trade off
Comparison matrix
From Route 1 in code — solve the normal equations: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| object | value (verified) |
|---|---|
| XᵀX | [[5, 15], [15, 55]] |
| Xᵀy | [20, 66] |
| w = [w₀, w₁] | [2.2, 0.6] |
Intuition
Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.
Assume each score is the true line plus Gaussian noise. The Gaussian density has exp(−(error)²) in it, so its log has −(error)². Maximizing the log-likelihood is minimizing the sum of squared errors — exactly OLS.
So 'least squares' is quietly saying 'I believe the noise is Gaussian.' Change the noise model and the loss changes: Laplace noise → absolute error (L1).
Analogy
Discussion prompt
Explain Why squared error is a probability statement by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.
Ranking
Put in order
Put the moves of Route 2 — OLS is Gaussian MLE into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Lesson 8: assume each residual is independent Gaussian noise around the linear mean.
Worked example
Model the noise as Gaussian: y = Xw + ε, ε ~ N(0, σ²)
Why: Lesson 8: assume each residual is independent Gaussian noise around the linear mean.
Write the negative log-likelihood
Why: Product of Gaussian densities → sum of log-densities. Constants aside, the −log-likelihood is proportional to the sum of squared residuals.
\[ -\log \mathcal{L}(w) = \tfrac{n}{2}\log(2\pi\sigma^2) + \frac{1}{2\sigma^2}\sum_i (y_i - x_i^\top w)^2 \]
Minimize over w
Why: Only the SSE term depends on w, and 1/(2σ²) is a positive constant, so argmin of the NLL = argmin of the SSE = the OLS solution.
\[ \arg\min_w \big[-\log\mathcal{L}(w)\big] = \arg\min_w \lVert Xw - y\rVert^2 \]
Verify: the NLL is smallest exactly at w = [2.2, 0.6]
Why: Perturbing w by [0.1, 0] raises the NLL from 5.7947 to 5.8197 — the OLS point is the maximum-likelihood point. Real execution below.
Blank canvas
Draw it
Draw what Route 2 — OLS is Gaussian MLE just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Missing information
Discussion prompt
Compute the OLS w, then show the Gaussian negative log-likelihood is minimized there — any nudge makes it larger.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Nudging w to [2.3, 0.6] raises the NLL to 5.8197. The likelihood peaks at the OLS coefficients — MLE and OLS are the same estimator under Gaussian noise.
Worked example
Compute the OLS w, then show the Gaussian negative log-likelihood is minimized there — any nudge makes it larger.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
def nll(w, sigma=1.0):
r = y - X @ w
n = len(y)
return 0.5*n*np.log(2*np.pi*sigma**2) + (r @ r)/(2*sigma**2)
print(w_ols) # [2.2 0.6]
print(round(nll(w_ols), 4)) # 5.7947
print(round(nll(w_ols + np.array([0.1, 0.0])), 4)) # 5.8197nll(w_ols) = 5.7947 is the minimum
Why: Nudging w to [2.3, 0.6] raises the NLL to 5.8197. The likelihood peaks at the OLS coefficients — MLE and OLS are the same estimator under Gaussian noise.
| w passed to nll | negative log-likelihood |
|---|---|
| [2.2, 0.6] (OLS) | 5.7947 (minimum) |
| [2.3, 0.6] (nudged) | 5.8197 (larger) |
| conclusion | MLE = OLS under Gaussian noise |
Step zero
Discussion prompt
Route 3 — the pseudoinverse (SVD) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Recall the SVD: X = UΣVᵀ (Lesson 16)
Answer:
Worked example
Recall the SVD: X = UΣVᵀ (Lesson 16)
Why: Every matrix factors this way; Σ holds the singular values σᵢ = √λᵢ(XᵀX) on its diagonal.
Define the Moore–Penrose pseudoinverse X⁺ = VΣ⁺Uᵀ
Why: Σ⁺ inverts each nonzero singular value. When X has full column rank, X⁺ = (XᵀX)⁻¹Xᵀ exactly.
\[ X^{+} = (X^\top X)^{-1} X^\top \quad\text{(full column rank)} \]
The least-squares solution is w = X⁺y
Why: Applying X⁺ to y projects onto the row space and returns the minimum-norm least-squares w — identical to the normal-equation answer, but computed via QR/SVD so it never squares the condition number.
\[ w = X^{+} y = \begin{bmatrix} 2.2 \\ 0.6 \end{bmatrix} \]
Verify: np.linalg.pinv(X) @ y = [2.2, 0.6]
Why: Confirmed by real execution below — the pseudoinverse route lands on the same coefficients as routes 1 and 2.
Translation
\( X^{+} = (X^\top X)^{-1} X^\top \quad\text{(full column rank)} \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Pattern
Predict first
The table runs: pseudoinverse X⁺y | L16 | w = [2.2, 0.6] · lstsq (QR/SVD) | L16 | w = [2.2, 0.6]
In Route 3 & 4 in code — pinv, lstsq, and the hat matrix, given the rows so far: what is the next one — the row where route is hat matrix Hy?
Correct: hat matrix Hy | L25 | ŷ = [2.8, 3.4, 4.0, 4.6, 5.2]
| route | lesson | output (verified) |
|---|---|---|
| pseudoinverse X⁺y | L16 | w = [2.2, 0.6] |
| lstsq (QR/SVD) | L16 | w = [2.2, 0.6] |
| hat matrix Hy | L25 | ŷ = [2.8, 3.4, 4.0, 4.6, 5.2] |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Hy is the projection of y onto the column space — the same predictions X@w produces.
Worked example
Three more routes in one runnable block: pseudoinverse, lstsq, and the projection ŷ = Hy with H = X(XᵀX)⁻¹Xᵀ. All must agree.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_pinv = np.linalg.pinv(X) @ y # Route 3: SVD
w_lstsq = np.linalg.lstsq(X, y, rcond=None)[0]
H = X @ np.linalg.inv(X.T @ X) @ X.T # Route 4: hat matrix
print(w_pinv.round(4)) # [2.2 0.6]
print(w_lstsq.round(4)) # [2.2 0.6]
print((H @ y).round(4)) # [2.8 3.4 4. 4.6 5.2]pinv and lstsq both give [2.2, 0.6]; Hy gives the fitted ŷ
Why: Hy is the projection of y onto the column space — the same predictions X@w produces. H is the linear operator that maps targets to fits.
| route | lesson | output (verified) |
|---|---|---|
| pseudoinverse X⁺y | L16 | w = [2.2, 0.6] |
| lstsq (QR/SVD) | L16 | w = [2.2, 0.6] |
| hat matrix Hy | L25 | ŷ = [2.8, 3.4, 4.0, 4.6, 5.2] |
Intuition
Picture y as a point floating above a flat plane — the column space of X, every prediction the model can reach. Because no line fits all five points, y is off the plane.
The nearest reachable point is the perpendicular shadow of y on the plane. Any other point on the plane is farther, because the straight-down drop is the shortest path.
That shadow is Hy, and the leftover arrow y − Hy = r points straight out — perpendicular to every column, which is exactly Xᵀr = 0. Geometry and the normal equations are the same statement.
Explain it
Discussion prompt
Explain Why projection is 'closest' to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The nearest reachable point is the perpendicular shadow of y on the plane. Any other point on the plane is farther, because the straight-down drop is the shortest path.
Concept
H = X(XᵀX)⁻¹Xᵀ is not just a formula — it is an orthogonal projector onto the column space. Two properties prove it: H is symmetric and idempotent (H² = H), and trace(H) equals the number of parameters.
\[ H^2 = H,\quad H^\top = H,\quad \operatorname{trace}(H) = 2,\quad (I - H)y = r \]
Applying H twice does nothing new — once you are on the plane, projecting again keeps you there. trace(H) = 2 counts the effective degrees of freedom (bias + slope).
Fill the middle
Fill in the blanks
From Verify the projection properties — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
H = X @ np.linalg.inv(X.T @ X) @ X.T
print(np.allclose(H, H @ H)) # True (idempotent)
print(np.allclose(H, H.T)) # True (symmetric)
print(round(np.trace(H), 4)) # 2.0
print((y - H @ y).round(4)) # residuals
Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. All four checks pass. The residual (I−H)y sticks straight out of the column space — the geometric signature of the least-squares fit.
Worked example
Confirm idempotence, symmetry, trace = 2, and that (I − H)y reproduces the residuals we know are [−0.8, 0.6, 1.0, −0.6, −0.2].
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
H = X @ np.linalg.inv(X.T @ X) @ X.T
print(np.allclose(H, H @ H)) # True (idempotent)
print(np.allclose(H, H.T)) # True (symmetric)
print(round(np.trace(H), 4)) # 2.0
print((y - H @ y).round(4)) # residualsH² = H, Hᵀ = H, trace = 2.0, (I−H)y = residuals
Why: All four checks pass. The residual (I−H)y sticks straight out of the column space — the geometric signature of the least-squares fit.
| property | value (verified) |
|---|---|
| H == H@H (idempotent) | True |
| H == Hᵀ (symmetric) | True |
| trace(H) | 2.0 |
| (I − H)y | [−0.8, 0.6, 1.0, −0.6, −0.2] |
Discrimination
Sort into buckets
Sort these by value (verified), from memory, without looking back at Verify the projection properties. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
With w = [2.2, 0.6], predict each student, subtract from the true score, and read off the residuals and R² — the numbers the exam asks you to produce.
| x | y | ŷ = 2.2 + 0.6x | r = y − ŷ | r² |
|---|---|---|---|---|
| 1 | 2 | 2.8 | −0.8 | 0.64 |
| 2 | 4 | 3.4 | +0.6 | 0.36 |
| 3 | 5 | 4.0 | +1.0 | 1.00 |
| 4 | 4 | 4.6 | −0.6 | 0.36 |
| 5 | 5 | 5.2 | −0.2 | 0.04 |
SSres = Σr² = 2.4, SStot = Σ(y−4)² = 6.0, R² = 1 − 2.4/6.0 = 0.6
Why: ȳ = 20/5 = 4; deviations [−2,0,1,0,1] squared and summed give 6.0. The line removes 60% of that baseline error. Also note Σr = 0 — the bias column forces zero mean residual.
Xᵀr = [0, 0] — the perpendicularity check
Why: Row 0 (ones·r): −0.8+0.6+1.0−0.6−0.2 = 0. Row 1 (x·r): −0.8+1.2+3.0−2.4−1.0 = 0. The residual sticks straight out of col(X), confirming w is the projection.
Pattern
Step through it
Step through The fit in numbers — residuals and R² one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Four Phase-1 lessons, four derivations, one answer. This equivalence IS the synthesis the mock exam probes.
| route | lesson | the claim | w |
|---|---|---|---|
| normal equations XᵀXw = Xᵀy | L7 | residual ⟂ columns | [2.2, 0.6] |
| Gaussian MLE | L8 | −log L ∝ SSE | [2.2, 0.6] |
| pseudoinverse X⁺y | L16 | SVD inverts Σ | [2.2, 0.6] |
| hat-matrix projection Hy | L25 | H projects onto col(X) | [2.2, 0.6] |
Say it aloud: geometry, probability, and optimization all point at the same [2.2, 0.6]. If you can explain why, Phase 1 is consolidated.
Section
Part 2 of 5
Concept
| result | the one line |
|---|---|
| eigen (L10) | Av = λv; symmetric ⇒ real λ, orthogonal vectors |
| SVD (L16) | A = UΣVᵀ for ANY matrix; σ = √eig(AᵀA) |
| PSD/PD (L13) | xᵀAx ≥ 0; XᵀX always PSD; PD ⟺ full rank |
| projection (L25) | ŷ = Hy, H = X(XᵀX)⁻¹Xᵀ, idempotent |
We just met H and PSD via OLS. Next we pin down eigenvalues, PD, and singular values with reproducible numbers.
Ranking
Put in order
Put the moves of Eigenvalues by hand, then PD into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Eigenvalues are the λ that make A − λI singular.
Worked example
Take A = [[2, 1], [1, 2]] — a symmetric matrix that shows up as a mini covariance. Find its eigenvalues from the characteristic polynomial, no solver.
Write det(A − λI) = 0
Why: Eigenvalues are the λ that make A − λI singular.
\[ \det\!\begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} = (2-\lambda)^2 - 1 = 0 \]
Expand to λ² − 4λ + 3 = 0
Why: (2−λ)² − 1 = 4 − 4λ + λ² − 1 = λ² − 4λ + 3. Note trace = 4 = sum of roots, det = 3 = product.
\[ \lambda^2 - 4\lambda + 3 = (\lambda - 1)(\lambda - 3) = 0 \;\Rightarrow\; \lambda = 1,\, 3 \]
Verify: both λ > 0, so A is positive definite
Why: np.roots([1,−4,3]) returns [3, 1] and eigh returns [1, 3] — matching the hand result. All-positive eigenvalues ⇒ PD (L13).
Notation
Annotate
From Eigenvalues by hand, then PD — read this one piece at a time. What is each part doing?
On: \( \lambda^2 - 4\lambda + 3 = (\lambda - 1)(\lambda - 3) = 0 \;\Rightarrow\; \lambda = 1,\, 3 \)
Worked example
Exam pace: one problem touching three lessons. For A = [[2,1],[1,2]], give eigenvalues, decide PD, and name the dominant eigenvector (what power iteration converges to).
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
vals, vecs = np.linalg.eigh(A) # ascending eigenvalues
print(vals.round(4)) # [1. 3.]
print(bool(vals.min() > 0)) # True (PD)
print(vecs[:, -1].round(4)) # dominant eigenvector
print((A @ vecs[:, -1]).round(4)) # = lambda * vλ = [1, 3], PD = True, dominant eigenvector = [0.7071, 0.7071]
Why: The largest eigenvalue is 3; its eigenvector [0.7071, 0.7071] is the direction power iteration finds. Av = [2.1213, 2.1213] = 3·v confirms it.
| question | answer (verified) | lesson |
|---|---|---|
| eigenvalues | [1, 3] | L10 |
| positive definite? | yes (all λ > 0) | L13 |
| dominant eigenvector | [0.7071, 0.7071] | L10 |
| Av vs λv check | [2.1213, 2.1213] both | L10 |
Intuition
Eigenvalues only exist for square matrices. But a 3×2 data matrix isn't square — so what measures its 'stretch'?
The SVD answers: A = UΣVᵀ factors any matrix into a rotation, a stretch by the singular values σ, and another rotation. The σ are the amounts A stretches its input directions.
And AᵀA is square and PSD, so its eigenvalues exist and are ≥ 0. The singular values are their square roots — the bridge from the non-square world back to eigenvalues.
Worked example
The exam loves σ = √eig(AᵀA). Test it on a 3×2 matrix that has no eigenvalues of its own — only singular values exist.
import numpy as np
A = np.array([[3., 0.], [0., 2.], [1., 1.]]) # 3x2, non-square
U, S, Vt = np.linalg.svd(A, full_matrices=False)
eigAtA = np.linalg.eigvalsh(A.T @ A) # ascending
print(S.round(4)) # [3.1926 2.1926]
print(np.sqrt(eigAtA[::-1]).round(4)) # [3.1926 2.1926]
print(np.allclose(A, U @ np.diag(S) @ Vt)) # TrueS = [3.1926, 2.1926] = √eig(AᵀA), and UΣVᵀ rebuilds A
Why: AᵀA is (2×2) and PSD, so its eigenvalues are ≥ 0 and their roots are the singular values — even though the non-square A has no eigenvalues. Reconstruction is exact.
| quantity | value (verified) |
|---|---|
| singular values S | [3.1926, 2.1926] |
| √eig(AᵀA) (desc) | [3.1926, 2.1926] |
| UΣVᵀ == A | True |
Concept
| result | the one line |
|---|---|
| MLE (L8) | argmax Σ log p; losses are negative log-likelihoods |
| MAP (L12) | MLE + log-prior; Gaussian→L2, Laplace→L1 |
| MVN (L23) | N(μ,Σ), sample x = μ + Lz (Cholesky) |
| info theory (L14/29) | KL ≥ 0; cross-entropy = H + KL; MI = 0 ⟺ independent |
We already used MLE for Route 2. Now nail the information-theory identity CE = H + KL with reproducible numbers.
Intuition
Entropy H(p) is the irreducible average surprise of the true distribution — the best any model could hope for. It doesn't depend on your model at all.
KL divergence D(p‖q) is the extra surprise from using the wrong distribution q instead of p. It is never negative — a wrong model can only cost you.
Cross-entropy is their sum. Since H(p) is fixed, minimizing cross-entropy loss over q is minimizing KL — driving your model toward the data. That's why classifiers train on cross-entropy.
Faded example
Fill in the blanks
KL ≥ 0 and cross-entropy = H + KL, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
p = np.array([0.7, 0.2, 0.1]) # true distribution
q = np.array([0.5, 0.3, 0.2]) # model distribution
H = -np.sum(p * np.log(p)) # entropy of p
KL = np.sum(p * np.log(p / q)) # KL(p || q) >= 0
CE = **-np.sum(p * np.log(q))** # cross-entropy H(p,q)
print(round(H, 4), round(KL, 4), round(CE, 4))
print(np.isclose(CE, H + KL)) # True
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Minimizing cross-entropy over q minimizes KL (H(p) is fixed), so CE-loss drives the model toward the data distribution.
Worked example
Take a true distribution p and a model q. Verify the identity that every training loss rests on: cross-entropy splits into the data's own entropy plus the KL gap.
\[ \underbrace{H(p,q)}_{\text{cross-entropy}} = \underbrace{H(p)}_{\text{entropy}} + \underbrace{D_{KL}(p\,\Vert\, q)}_{\ge\, 0} \]
import numpy as np
p = np.array([0.7, 0.2, 0.1]) # true distribution
q = np.array([0.5, 0.3, 0.2]) # model distribution
H = -np.sum(p * np.log(p)) # entropy of p
KL = np.sum(p * np.log(p / q)) # KL(p || q) >= 0
CE = -np.sum(p * np.log(q)) # cross-entropy H(p,q)
print(round(H, 4), round(KL, 4), round(CE, 4))
print(np.isclose(CE, H + KL)) # TrueH = 0.8018, KL = 0.0851 ≥ 0, CE = 0.8869 = H + KL
Why: Minimizing cross-entropy over q minimizes KL (H(p) is fixed), so CE-loss drives the model toward the data distribution. KL is never negative — that is why CE ≥ H always.
| quantity | value (verified) |
|---|---|
| H(p) entropy | 0.8018 |
| KL(p‖q) | 0.0851 (≥ 0) |
| CE(p,q) | 0.8869 |
| H + KL | 0.8869 (= CE) |
Pattern
Step through it
Step through KL ≥ 0 and cross-entropy = H + KL one row at a time. What is driving the change, and what would the row after the last one be?
Concept
| result | the one line |
|---|---|
| gradient (L6/9) | ∇ points uphill; backprop = chain rule of Jacobians |
| convexity (L11) | Hessian PSD everywhere ⇒ only global minima |
| optimizers (L15) | Adam = momentum + RMSprop + bias correction |
| Newton (L21) | θ −= H⁻¹∇f; quadratic convergence, O(d³) |
OLS is a convex quadratic, so both gradient descent and Newton must converge to the closed-form [2.2, 0.6]. We watch each do it — and watch GD blow up when the step is too big.
Concept
The SSE loss ‖Xw − y‖² has Hessian ∇²L = 2XᵀX, which is PSD everywhere (zᵀXᵀXz = ‖Xz‖² ≥ 0). A PSD-everywhere Hessian means the loss is convex — one global bowl, no bad local minima.
\[ \nabla^2 L = 2X^\top X \succeq 0 \;\Rightarrow\; L \text{ convex} \;\Rightarrow\; \nabla L = 0 \text{ is the global min} \]
So any descent method that keeps moving downhill must reach the same w★. Gradient descent and Newton below are two such methods — and they land on the identical [2.2, 0.6].
Estimation
Predict first
The SSE gradient is ∇L = 2Xᵀ(Xw − y). Step downhill from zero; with a small enough learning rate it lands on the closed-form solution.
Commit before you compute: what does Gradient descent converges to the OLS w come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: After 20000 steps at lr = 0.001, GD w = [2.2, 0.6]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Convexity guarantees a single global minimum; GD slides into it and matches the closed form.
Worked example
The SSE gradient is ∇L = 2Xᵀ(Xw − y). Step downhill from zero; with a small enough learning rate it lands on the closed-form solution.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.zeros(2)
lr = 0.001
for t in range(20000):
grad = 2 * X.T @ (X @ w - y) # gradient of SSE
w = w - lr * grad
print(w.round(4)) # [2.2 0.6]
print(np.linalg.solve(X.T @ X, X.T @ y).round(4)) # [2.2 0.6]After 20000 steps at lr = 0.001, GD w = [2.2, 0.6]
Why: Convexity guarantees a single global minimum; GD slides into it and matches the closed form. The gradient at the optimum is zero, so the update stops moving.
| method | w (verified) |
|---|---|
| gradient descent (lr 0.001) | [2.2, 0.6] |
| closed-form solve | [2.2, 0.6] |
| difference | < 1e-3 |
Comparison
Comparison matrix
From Gradient descent converges to the OLS w: refill the w (verified) column from what you know. The rest of the table is as it appeared.
| method | w (verified) |
|---|---|
| gradient descent (lr 0.001) | [2.2, 0.6] |
| closed-form solve | [2.2, 0.6] |
| difference | < 1e-3 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Bigger steps should learn faster, so crank the learning rate to lr = 0.02 for the same SSE gradient descent.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118.
Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.
Why: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.
Trap
Bigger steps should learn faster, so crank the learning rate to lr = 0.02 for the same SSE gradient descent.
lr = 0.02 → w explodes past 1e300 to inf, NaN territory
Why: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.
Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.
lr = 0.001 < 2/L → w converges to [2.2, 0.6]
Why: Below the 2/L threshold every step contracts toward the minimum. The convex bowl has one bottom; a safe step size reaches it. Always bound the step by the curvature.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.
Fill the middle
Fill in the blanks
From Newton solves OLS in ONE step — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.array([0.0, 0.0]) # any start
grad = 2 * X.T @ (X @ w - y) # gradient at w
Hess = 2 * X.T @ X # Hessian (constant for OLS)
w_new = w - np.linalg.solve(Hess, grad) # one Newton step
print(w_new.round(4)) # [2.2 0.6]
Why: w is what everything below it consumes, so the wrong expression here fails later and somewhere else. For a quadratic the Hessian is constant and the second-order model is exact, so a single Newton step is the closed-form solution.
Worked example
OLS is exactly quadratic, so Newton's method — w ← w − H⁻¹∇ — jumps to the exact minimum in a single step from any start, because the quadratic model is the function itself.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.array([0.0, 0.0]) # any start
grad = 2 * X.T @ (X @ w - y) # gradient at w
Hess = 2 * X.T @ X # Hessian (constant for OLS)
w_new = w - np.linalg.solve(Hess, grad) # one Newton step
print(w_new.round(4)) # [2.2 0.6]One Newton step from [0, 0] lands on [2.2, 0.6]
Why: For a quadratic the Hessian is constant and the second-order model is exact, so a single Newton step is the closed-form solution. This is why Newton has quadratic convergence — but each step costs O(d³) to solve.
| step | w (verified) |
|---|---|
| start | [0, 0] |
| after 1 Newton step | [2.2, 0.6] |
| closed form | [2.2, 0.6] |
Concept
| result | the one line |
|---|---|
| losses (L17) | softmax+CE ⇒ dL/dz = p − y |
| bias-variance (L32) | MSE = Bias² + Variance + Noise (U-curve) |
| metrics (L26/35) | imbalance ⇒ PR-AUC; AUC = P(pos > neg) |
| SVM/kernels (L28/30) | dual in α; k = φ·φ; Mercer ⇒ PSD Gram |
The single most exam-tested gradient is softmax + cross-entropy. We derive dL/dz = p − y and confirm it against a numerical gradient.
Worked example
The clean gradient that makes classification trainable: the derivative of cross-entropy with respect to the logits is just predicted probability minus the one-hot label.
\[ L = -\sum_k y_k \log p_k,\quad p = \operatorname{softmax}(z) \;\Rightarrow\; \frac{\partial L}{\partial z} = p - y \]
import numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z = np.array([2.0, 1.0, 0.1]) # logits
y = np.array([1.0, 0.0, 0.0]) # one-hot label
p = softmax(z)
print(p.round(4)) # [0.659 0.2424 0.0986]
print((p - y).round(4)) # analytic gradient
eps = 1e-6
g = np.array([( -np.sum(y*np.log(softmax(z+eps*np.eye(3)[i])))
+np.sum(y*np.log(softmax(z-eps*np.eye(3)[i]))) )/(2*eps)
for i in range(3)])
print(g.round(4)) # numerical gradient (matches)p − y = [−0.341, 0.2424, 0.0986] matches the numerical gradient
Why: The messy softmax Jacobian and the cross-entropy derivative cancel into p − y. The finite-difference check agrees to 1e-5 — the analytic shortcut is correct, which is why frameworks fuse softmax+CE.
| quantity | value (verified) |
|---|---|
| p = softmax(z) | [0.6590, 0.2424, 0.0986] |
| p − y (analytic) | [−0.3410, 0.2424, 0.0986] |
| numerical ∂L/∂z | [−0.3410, 0.2424, 0.0986] |
| match | True (atol 1e-5) |
Discrimination
Sort into buckets
Sort these by value (verified), from memory, without looking back at softmax + cross-entropy ⇒ dL/dz = p − y. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Intuition
Test error splits into three sources. Bias is being systematically off — a too-simple model that can't capture the shape. Variance is being jumpy — a too-flexible model that chases the particular training noise.
Noise is the irreducible floor: even the perfect model can't predict the random part of y. No amount of cleverness removes it.
The trade: making a model more flexible lowers bias but raises variance. The sweet spot minimizes their sum — the U-shaped test-error curve. The simulation next reads off all three.
Worked example
MSE = Bias² + Variance + Noise. Estimate a constant model's prediction at a point over many noisy training sets and read off the three pieces.
\[ \mathbb{E}\big[(y - \hat f)^2\big] = \underbrace{(\mathbb{E}\hat f - f)^2}_{\text{Bias}^2} + \underbrace{\operatorname{Var}(\hat f)}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise}} \]
import numpy as np
np.random.seed(0)
f_true, sigma = 2.0, 1.0
preds = []
for _ in range(20000):
sample = f_true + np.random.randn(5)*sigma # 5 noisy points
preds.append(sample.mean()) # constant model
preds = np.array(preds)
bias2 = (preds.mean() - f_true)**2
var = preds.var()
print(round(bias2, 4), round(var, 4), round(sigma**2, 4))
print(round(bias2 + var + sigma**2, 4)) # total MSEBias² ≈ 0.0, Variance ≈ 0.1984, Noise = 1.0 → MSE ≈ 1.1984
Why: The mean estimator is unbiased (Bias² ≈ 0) with variance σ²/5 = 0.2, plus irreducible noise σ² = 1. seed(0) makes these reproducible. Reducing model bias trades against rising variance — the U-curve.
| component | value (verified, seed 0) |
|---|---|
| Bias² | 0.0 (unbiased mean) |
| Variance | 0.1984 (≈ σ²/5) |
| Noise σ² | 1.0 |
| MSE = B² + V + N | 1.1984 |
Concept
ROC-AUC has a clean probabilistic meaning: it is the chance that a random positive gets a higher score than a random negative.
\[ \text{ROC-AUC} = \mathbb{P}\big(\text{score(pos)} > \text{score(neg)}\big) \]
The catch: its false-positive rate divides by the number of negatives. Under 2% positives there are ~49× more negatives, so a flood of false positives barely moves the rate — ROC-AUC stays flattering. PR-AUC, which uses precision, does not forgive that.
Faded example
Fill in the blanks
ROC-AUC vs PR-AUC under 2% imbalance, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
np.random.seed(0)
n = 5000
y = (np.random.rand(n) < 0.02).astype(int) # ~2% positives
scores = *np.random.rand(n) + 0.15y** # weak signal
print(int(y.sum()), 'of', n) # 105 of 5000
print(round(roc_auc_score(y, scores), 4)) # 0.6369
print(round(average_precision_score(y, scores), 4)) # 0.1771
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. ROC-AUC's false-positive rate has a huge true-negative denominator, so it stays flattering.
Worked example
On a rare-positive problem, ROC-AUC looks optimistic while PR-AUC tells the truth about the minority class. Simulate 2% positives with a weak scorer.
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
np.random.seed(0)
n = 5000
y = (np.random.rand(n) < 0.02).astype(int) # ~2% positives
scores = np.random.rand(n) + 0.15*y # weak signal
print(int(y.sum()), 'of', n) # 105 of 5000
print(round(roc_auc_score(y, scores), 4)) # 0.6369
print(round(average_precision_score(y, scores), 4)) # 0.1771ROC-AUC = 0.6369 but PR-AUC = 0.1771 on the same scores
Why: ROC-AUC's false-positive rate has a huge true-negative denominator, so it stays flattering. PR-AUC watches precision on the rare positives and exposes the weakness. Under imbalance, report PR-AUC.
| metric | value (verified, seed 0) |
|---|---|
| positives | 105 of 5000 (2.1%) |
| ROC-AUC | 0.6369 (optimistic) |
| PR-AUC | 0.1771 (honest) |
| base rate | 0.021 |
Section
Part 3 of 5
Concept
The threads weave together: eigenvalues → SVD → PCA; MLE → cross-entropy → the training loop; gradients → backprop → optimizers; PSD → kernels → SVM.
Nothing in Phase 1 stands alone — every result is a prerequisite for Phase 2 (classical ML) and beyond (transformers, generative models). OLS-four-ways is the smallest example of the whole web.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.
You've beaten all three this phase — keep them straight.
Why: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.
Trap
Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.
Carry these three misconceptions into the exam
Why: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.
You've beaten all three this phase — keep them straight.
p = P(data or more extreme | H₀) (L20); PD needs ALL leading minors > 0 (L13); use cross-entropy for classification (L17)
Why: The p-value conditions on H₀, not the reverse. det>0 is necessary but not sufficient for PD. Cross-entropy gives the clean p − y gradient; MSE saturates. Name the correct statement for each cold.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
X = [1, x] (a bias column then the feature), target y. Keep these five numbers in your head — every slide reuses them.; Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.; Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.lr = 0.02 for the same SSE gradient descent.; Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.Pattern
Predict first
The table runs: det(A) | 1.0 > 0 | misleadingly 'passes' · eigenvalues | [−1, −1] | both < 0 → NOT PD
In Trap in code — det > 0 is NOT positive definite, given the rows so far: what is the next one — the row where test is xᵀAx at [1,1]?
Correct: xᵀAx at [1,1] | −2.0 | confirms not PD
| test | result | verdict |
|---|---|---|
| det(A) | 1.0 > 0 | misleadingly 'passes' |
| eigenvalues | [−1, −1] | both < 0 → NOT PD |
| xᵀAx at [1,1] | −2.0 | confirms not PD |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Two negatives multiply to a positive determinant, so det>0 is fooled.
Worked example
A concrete counterexample: −I has determinant +1 (two negative eigenvalues multiply to positive) but is negative definite. Determinant alone cannot certify PD.
import numpy as np
A = np.array([[-1., 0.], [0., -1.]])
print(round(np.linalg.det(A), 4)) # 1.0 (positive!)
print(np.linalg.eigvalsh(A)) # [-1. -1.]
x = np.array([1., 1.])
print(x @ A @ x) # -2.0 (< 0, so NOT PD)det(A) = 1.0 > 0 yet eigenvalues are [−1, −1] and xᵀAx = −2
Why: Two negatives multiply to a positive determinant, so det>0 is fooled. The honest PD test is 'all eigenvalues > 0' or 'all leading minors > 0'. Here xᵀAx < 0 proves A is not PD.
| test | result | verdict |
|---|---|---|
| det(A) | 1.0 > 0 | misleadingly 'passes' |
| eigenvalues | [−1, −1] | both < 0 → NOT PD |
| xᵀAx at [1,1] | −2.0 | confirms not PD |
Ranking
Put in order
These are the steps of The exam-pace retrieval protocol, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Apply this to the five drills next. Struggle, then check — the drills after this point are your cold attempts.
Sorting
Sort into buckets
These are the pieces of Lesson 37: Phase 1 Review & Mock Prep, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 5 — slides closed
Elimination
Eliminate the wrong options
Why do the normal equations, Gaussian MLE, the pseudoinverse, and the hat-matrix projection all return the same w = [2.2, 0.6]?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Least squares has one minimizer. The perpendicular-residual condition (projection), ∇SSE = 0 (normal equations), X⁺y (SVD), and the Gaussian MLE all characterize that same point, so they must coincide exactly when X has full column rank.
Check
Cold attempt, then click. This is the capstone idea of the whole session.
Check your understanding
Why do the normal equations, Gaussian MLE, the pseudoinverse, and the hat-matrix projection all return the same w = [2.2, 0.6]?
Answer: A
Why: Least squares has one minimizer. The perpendicular-residual condition (projection), ∇SSE = 0 (normal equations), X⁺y (SVD), and the Gaussian MLE all characterize that same point, so they must coincide exactly when X has full column rank.
Check
Cold attempt, then click.
Check your understanding
The singular values of any matrix A are:
Answer: A
Why: σᵢ = √λᵢ(AᵀA). AᵀA is PSD so the σ are real and ≥ 0, and the SVD exists for ANY matrix (L16). We verified [3.1926, 2.1926] = √eig(AᵀA) for a 3×2 A.
Prediction
Predict first
Minimizing mean squared error is maximum likelihood under which assumption?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Gaussian noise with constant variance
Why: With y = f(x) + ε, ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² — exactly MSE (L8). We saw the NLL bottom out at the OLS w = [2.2, 0.6].
Check
Slides closed first.
Check your understanding
Minimizing mean squared error is maximum likelihood under which assumption?
Answer: A
Why: With y = f(x) + ε, ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² — exactly MSE (L8). We saw the NLL bottom out at the OLS w = [2.2, 0.6].
Prediction
Predict first
You run gradient descent on the OLS loss with learning rate 0.02 and the weights explode to infinity. The largest eigenvalue of the Hessian 2XᵀX is about 118. What is the fix?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Lower the learning rate below 2/L ≈ 0.017
Why: Gradient descent on a quadratic is stable only when lr < 2/L, where L is the largest Hessian eigenvalue. Here 2/L ≈ 0.017, so lr = 0.02 overshoots every step and diverges; lr = 0.001 converges to [2.2, 0.6].
Check
Attempt before clicking.
Check your understanding
You run gradient descent on the OLS loss with learning rate 0.02 and the weights explode to infinity. The largest eigenvalue of the Hessian 2XᵀX is about 118. What is the fix?
Answer: A
Why: Gradient descent on a quadratic is stable only when lr < 2/L, where L is the largest Hessian eigenvalue. Here 2/L ≈ 0.017, so lr = 0.02 overshoots every step and diverges; lr = 0.001 converges to [2.2, 0.6].
Elimination
Eliminate the wrong options
On a 2%-positive dataset your model gets ROC-AUC 0.64 but PR-AUC 0.18. Which metric better reflects minority-class performance, and why?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Under heavy imbalance ROC-AUC stays optimistic because the false-positive rate is diluted by many true negatives. PR-AUC watches precision on the positives and drops to 0.18, honestly reflecting the minority performance (verified on 105/5000 positives).
Check
Last one — then we find your gaps.
Check your understanding
On a 2%-positive dataset your model gets ROC-AUC 0.64 but PR-AUC 0.18. Which metric better reflects minority-class performance, and why?
Answer: A
Why: Under heavy imbalance ROC-AUC stays optimistic because the false-positive rate is diluted by many true negatives. PR-AUC watches precision on the positives and drops to 0.18, honestly reflecting the minority performance (verified on 105/5000 positives).
Section
Part 5 of 5
Concept
Turn this session into a plan: for every topic, rate your confidence 1–5 and log anything that needed a hint.
| confidence | action |
|---|---|
| 5 — cold, fast | retire from active rotation |
| 3–4 — slow / shaky | re-quiz in 3 days, then 1 week |
| 1–2 — needed a hint | re-derive now, re-quiz next session |
The prioritized 1–2 list IS your mock-exam study plan. A concept is 'done' only when you reproduce it cold on a later day.
Concept
The capstone build: fit the five-student line and prove three independent routes agree. You derived every piece — now assemble it yourself, typing each line and predicting each output before you run it.
| # | milestone | route |
|---|---|---|
| 1 | Design matrix + normal equations | np.linalg.solve |
| 2 | Pseudoinverse route | np.linalg.pinv |
| 3 | Score R² and match sklearn | LinearRegression |
Build rules: type every line yourself, run after each line, and when something errors, read the array shapes first — don't delete the error.
Worked example
Your turn: build X with a bias column and solve the normal equations. Say the shape (5, 2) and the answer [2.2, 0.6] out loud before printing.
Hint: np.c_[np.ones_like(x), x] stacks the bias column, then np.linalg.solve(X.T @ X, X.T @ y). Never use inv.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
XtX = X.T @ X
Xty = X.T @ y
w = np.linalg.solve(XtX, Xty)
print(X.shape) # (5, 2)
print(XtX) # [[5 15] [15 55]]
print(w) # [2.2 0.6]| check | value (verified) |
|---|---|
| X.shape | (5, 2) |
| XᵀX | [[5, 15], [15, 55]] |
| Xᵀy | [20, 66] |
| w | [2.2, 0.6] |
Worked example
Your turn: get the same w a second, independent way — the SVD-based pseudoinverse. Predict it will match milestone 1 before printing.
Hint: np.linalg.pinv(X) @ y computes X⁺y directly via the SVD, no XᵀX formed.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w_normal = np.linalg.solve(X.T @ X, X.T @ y)
w_pinv = np.linalg.pinv(X) @ y
print(w_normal.round(4)) # [2.2 0.6]
print(w_pinv.round(4)) # [2.2 0.6]
print(np.allclose(w_normal, w_pinv)) # True| route | w (verified) |
|---|---|
| normal equations | [2.2, 0.6] |
| pseudoinverse X⁺y | [2.2, 0.6] |
| allclose | True |
Fill the middle
Fill in the blanks
From Milestone 3 — score R² and match sklearn — one line has had its right-hand side removed. Put it back.
import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
yhat = X @ w
r2 = 1 - np.sum((y - yhat)2) / np.sum((y - y.mean())2)
print('w =', w.round(4), ' R2 =', round(r2, 4)) # [2.2 0.6] 0.6
m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))
Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Your from-scratch line explains 60% of the score variance, and sklearn's fitted intercept and slope match your hand-built coefficients exactly.
Worked example
Your turn: compute R² from scratch, then fit sklearn and confirm intercept, slope, and score all agree. Predict whether they match exactly.
Hint: ŷ = X @ w; R² = 1 − Σ(y−ŷ)² / Σ(y−ȳ)². Fit sklearn on x.reshape(-1, 1) and compare intercept_, coef_.
import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print('w =', w.round(4), ' R2 =', round(r2, 4)) # [2.2 0.6] 0.6
m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))R² = 0.6, and sklearn prints 2.2 0.6 — everything agrees
Why: Your from-scratch line explains 60% of the score variance, and sklearn's fitted intercept and slope match your hand-built coefficients exactly. You reproduced linear regression from the math up.
| source | intercept | slope | R² |
|---|---|---|---|
| your OLS | 2.2 | 0.6 | 0.6 |
| sklearn | 2.2 | 0.6 | 0.6 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
R² = 0.6, and sklearn prints 2.2 0.6 — everything agrees
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Your turn: compute R² from scratch, then fit sklearn and confirm intercept, slope, and score all agree. Predict whether they match exactly.
Intuition
Suppose one feature is a copy of another — income and income in cents. The second column points the same direction as the first; it adds no new axis to the column space.
Now infinitely many weight vectors give the identical prediction — shift weight from one twin to the other and nothing changes. There is no unique w, so XᵀX is singular and solve can't pick an answer.
Ridge breaks the tie: by charging for weight size, it prefers the smallest ‖w‖ among all the equivalent fits — a unique answer, always. That's the fallback we verify next.
Concept
When a feature is collinear, XᵀX has a zero eigenvalue and solve fails. Ridge adds λI, lifting every eigenvalue by λ so the system is always solvable.
\[ (X^\top X + \lambda I)\, w = X^\top y \]
Add a redundant column 2x to our data: XᵀX becomes rank-2 of 3 with eigenvalues {0, 0.90, 279.1}. That 0 is what kills solve. Ridge with λ = 1 shifts it to 1.
Explain it
Discussion prompt
Explain Ridge — the always-invertible fallback to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
When a feature is collinear, XᵀX has a zero eigenvalue and solve fails. Ridge adds λI, lifting every eigenvalue by λ so the system is always solvable.
Missing information
Discussion prompt
Same collinear Xc that makes solve raise. Show the singular eigenvalue, add λI, and watch it become solvable — runnable as written.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
rank 2 of 3 gives a zero eigenvalue, so plain solve raises LinAlgError. Adding λI = np.eye(3) lifts every eigenvalue by 1, making the matrix non-singular. Same solve now returns [1.0734, 0.1808, 0.3616].
Worked example
Same collinear Xc that makes solve raise. Show the singular eigenvalue, add λI, and watch it become solvable — runnable as written.
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x] # col 3 = 2 * col 2
G = Xc.T @ Xc
print(np.linalg.matrix_rank(Xc)) # 2 (not 3)
print(np.linalg.eigvalsh(G).round(4)) # [-0. 0.8957 279.1043]
wr = np.linalg.solve(G + 1.0*np.eye(3), Xc.T @ y)
print(np.linalg.eigvalsh(G + np.eye(3)).round(4)) # [1. 1.8957 280.1043]
print(wr.round(4)) # [1.0734 0.1808 0.3616]The 0 eigenvalue becomes 1 after adding λI, and solve returns weights
Why: rank 2 of 3 gives a zero eigenvalue, so plain solve raises LinAlgError. Adding λI = np.eye(3) lifts every eigenvalue by 1, making the matrix non-singular. Same solve now returns [1.0734, 0.1808, 0.3616].
| matrix | eigenvalues (verified) | solvable? |
|---|---|---|
| XcᵀXc (collinear) | [0, 0.8957, 279.1] | no — singular |
| XcᵀXc + I (ridge) | [1, 1.8957, 280.1] | yes |
| ridge w | [1.0734, 0.1808, 0.3616] | returned |
Trade off
Comparison matrix
From Ridge lifts the zero eigenvalue: every row here is a choice with a cost. Fill the eigenvalues (verified) column, then say which row you would actually pick and what you give up for it.
| matrix | eigenvalues (verified) | solvable? |
|---|---|---|
| XcᵀXc (collinear) | [0, 0.8957, 279.1] | no — singular |
| XcᵀXc + I (ridge) | [1, 1.8957, 280.1] | yes |
| ridge w | [1.0734, 0.1808, 0.3616] | returned |
Estimation
Predict first
One more Phase-1 bridge: adding λ‖w‖² is the same as putting a Gaussian prior w ~ N(0, τ²I) on the weights and doing MAP (L12). The penalty is the log-prior. Ridge shrinks ‖w‖; verify it.
Commit before you compute: what does Ridge is MAP with a Gaussian prior come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: λ = 1 shrinks w from [2.2, 0.6] to [1.1712, 0.8649], and ‖w‖ from 2.2804 to 1.4559
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The Gaussian prior pulls weights toward 0, trading a little training fit for stability — the MAP estimate under a zero-mean weight prior.
Worked example
One more Phase-1 bridge: adding λ‖w‖² is the same as putting a Gaussian prior w ~ N(0, τ²I) on the weights and doing MAP (L12). The penalty is the log-prior. Ridge shrinks ‖w‖; verify it.
\[ \underbrace{\lVert Xw - y\rVert^2}_{\text{Gaussian likelihood}} + \underbrace{\lambda \lVert w\rVert^2}_{\text{Gaussian log-prior}} \;\Rightarrow\; (X^\top X + \lambda I)w = X^\top y \]
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
w_ridge = np.linalg.solve(X.T @ X + 1.0*np.eye(2), X.T @ y)
print(w_ols.round(4)) # [2.2 0.6]
print(w_ridge.round(4)) # [1.1712 0.8649]
print(round(np.linalg.norm(w_ols), 4),
round(np.linalg.norm(w_ridge), 4)) # 2.2804 1.4559λ = 1 shrinks w from [2.2, 0.6] to [1.1712, 0.8649], and ‖w‖ from 2.2804 to 1.4559
Why: The Gaussian prior pulls weights toward 0, trading a little training fit for stability — the MAP estimate under a zero-mean weight prior. As λ → 0 ridge recovers OLS; as λ → ∞ it forces w → 0.
| λ | w = [w₀, w₁] (verified) | ‖w‖ |
|---|---|---|
| 0 (OLS / MLE) | [2.2, 0.6] | 2.2804 |
| 1 (ridge / MAP) | [1.1712, 0.8649] | 1.4559 |
Comparison
Comparison matrix
From Ridge is MAP with a Gaussian prior: refill the w = [w₀, w₁] (verified) column from what you know. The rest of the table is as it appeared.
| λ | w = [w₀, w₁] (verified) | ‖w‖ |
|---|---|---|
| 0 (OLS / MLE) | [2.2, 0.6] | 2.2804 |
| 1 (ridge / MAP) | [1.1712, 0.8649] | 1.4559 |
Concept
The mock tests theory AND code at pace. Before it: a clean environment ready, the 1–2 confidence list reviewed, and the retrieval log current.
On the exam: attempt cold first, show every derivation step, and watch the three traps (p-value conditioning, det>0 ⇏ PD, MSE-for-classification). You built the whole foundation — now prove it sticks.
Analogy
Discussion prompt
Explain Mock-exam readiness by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The mock tests theory AND code at pace. Before it: a clean environment ready, the 1–2 confidence list reviewed, and the retrieval log current.
Concept
Slides closed, out loud: explain (1) why the four OLS routes must agree, (2) why softmax+CE gives p − y, and (3) what λI does to the eigenvalues of XᵀX.
Stretch: run lr = 0.02 gradient descent and watch it diverge, then drop to 0.001 and watch it converge to [2.2, 0.6] — the exact stability boundary the exam likes to probe.
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why the four OLS routes must agree, (2) why softmax+CE gives p − y, and (3) what λI does to the eigenvalues of XᵀX.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Stretch: run lr = 0.02 gradient descent and watch it diverge, then drop to 0.001 and watch it converge to [2.2, 0.6] — the exact stability boundary the exam likes to probe.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — One dataset, four routes to OLS · The four pillars, re-derived · Traps & synthesis · Retrieval drills · Find the gaps & build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
[2.2, 0.6], verified[1, 3], SVD σ = √eig(AᵀA), CE = H + KL, softmax ∂L/∂z = p − ylr = 2/L| pillar | the one keystone (verified) |
|---|---|
| synthesis | four routes → one w = [2.2, 0.6] |
| linear algebra | SVD: A = UΣVᵀ, σ = √eig(AᵀA) |
| probability | losses are negative log-likelihoods; CE = H + KL |
| optimization | convex ⇒ global min; lr < 2/L for stability |
| ML foundations | softmax+CE ⇒ ∂L/∂z = p − y; imbalance ⇒ PR-AUC |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.