Lesson 37: Phase 1 Review & Mock Prep

USAAIO Lesson 37, from Week 13, a fully worked Phase 1 consolidation built around ONE running dataset, the five students from Lesson 7. OLS is derived and verified FOUR equivalent ways - by the normal equations, by Gaussian MLE, by the pseudoinverse and SVD, and by hat-matrix projection - with every step shown. Each of the four pillars is then re-derived on that same data: eigenvalues and PSD matrices, the SVD singular values, softmax with cross-entropy giving dL/dz = p - y, KL divergence and cross-entropy, bias and variance, ROC against PR-AUC under imbalance, gradient descent and Newton's method both converging to the closed form, and ridge regression lifting the zero eigenvalue. Every snippet runs self-contained, and every number was produced by real execution. The lesson runs to 62 slides.

Subject: Machine Learning · 112 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Phase 1 Review & Mock Prep

Title

USAAIO · Lesson 37 · Week 13

Twelve weeks, one coherent foundation. We prove it by fitting one line to five students four different ways — normal equations, MLE, pseudoinverse, projection — then re-derive every pillar on that same data, all verified against real execution.

2. By the end of this session you can

Objectives

  1. Fit OLS on the running dataset four equivalent ways and prove all four give w = [2.2, 0.6]
  2. Show why normal equations, Gaussian MLE, the pseudoinverse, and the hat-matrix projection are one object
  3. Re-derive each Phase-1 pillar — eigen/PSD, SVD, softmax+CE, KL, bias-variance, AUC — with numbers you can reproduce
  4. Watch gradient descent and Newton converge to the closed-form w, and know when GD diverges
  5. Beat the three high-yield exam traps and build a prioritized gap list for the mock

3. What survived from Initialization & Gradient Clipping?

Warm-up

Discussion prompt

Before we open Lesson 37: Phase 1 Review & Mock Prep: without looking back, what was the main idea of Initialization & Gradient Clipping, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

exploding gradients and norm clipping, Xavier and He initialization derived from variance propagation, why He init keeps ReLU activations stable through deep networks, and why zero initialization fails. Run a 50-layer signal-propagation experiment and compare initializations.

4. One dataset, four routes to OLS

Section

Part 1 of 5 — the synthesis capstone

5. The running dataset

Concept

Everything today runs on Lesson 7's five students: hours studied x and score y. We fit one line ŷ = w₀ + w₁·x and reach the SAME w from four Phase-1 directions.

studenthours xscore y
112
224
335
444
555

Design matrix X = [1, x] (a bias column then the feature), target y. Keep these five numbers in your head — every slide reuses them.

6. Fill in: hours x for The running dataset

Comparison

Comparison matrix

From The running dataset: refill the hours x column from what you know. The rest of the table is as it appeared.

studenthours xscore y
112
224
335
444
555

7. Why 'four ways' is the whole point

Intuition

Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.

Geometry says project y onto the column space. Probability says maximize the Gaussian likelihood. Linear algebra says apply the pseudoinverse. Optimization says set the gradient to zero. They must agree — and today we watch them agree to the fourth decimal.

If you can explain the equivalence out loud, you understand Phase 1. That is the mock-exam bar.

8. Break it if you can: Why 'four ways' is the whole point

Counterexample

Discussion prompt

Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

If you can explain the equivalence out loud, you understand Phase 1. That is the mock-exam bar.

9. What has to happen first: Route 1 — the normal equations

Ranking

Put in order

Put the moves of Route 1 — the normal equations into the order they have to happen.

  1. Start from the overdetermined system Xw = y
  2. Left-multiply both sides by Xᵀ
  3. Assemble the entries by hand
  4. Verify: solve gives w = [2.2, 0.6]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Five equations, two unknowns — no exact solution.

10. Route 1 — the normal equations

Worked example

Start from the overdetermined system Xw = y

Why: Five equations, two unknowns — no exact solution. Least squares finds the closest w.

\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix},\quad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \]

Left-multiply both sides by Xᵀ

Why: Xᵀ is (2×5), so XᵀX is a square (2×2) system — the normal equations from Lesson 7.

\[ \boxed{\,X^\top X\, w = X^\top y\,} \]

Assemble the entries by hand

Why: XᵀX top-left = Σ1 = 5, off-diagonal = Σx = 15, bottom-right = Σx² = 55; Xᵀy = [Σy, Σxy] = [20, 66].

\[ \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} w = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]

Verify: solve gives w = [2.2, 0.6]

Why: det = 5·55 − 15² = 50 ≠ 0, so the (2×2) is invertible and w is unique. Real execution below confirms it.

11. Decode the notation: Route 1 — the normal equations

Notation

Annotate

From Route 1 — the normal equations — read this one piece at a time. What is each part doing?

On: \( X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix},\quad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \)

  • Five equations, two unknowns — no exact solution. Least squares finds the closest w.
  • Xᵀ is (2×5), so XᵀX is a square (2×2) system — the normal equations from Lesson 7.
  • XᵀX top-left = Σ1 = 5, off-diagonal = Σx = 15, bottom-right = Σx² = 55; Xᵀy = [Σy, Σxy] = [20, 66].

12. Plan first: Route 1 — XᵀX and Xᵀy, entry by entry

Step zero

Discussion prompt

Route 1 — XᵀX and Xᵀy, entry by entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Top-left = ones·ones = Σ1 = 5

Answer:

  1. Top-left = ones·ones = Σ1 = 5
  2. Off-diagonal = ones·x = Σx = 1+2+3+4+5 = 15
  3. Bottom-right = x·x = Σx² = 1+4+9+16+25 = 55; and Xᵀy = [Σy, Σxy] = [20, 66]

13. Route 1 — XᵀX and Xᵀy, entry by entry

Worked example

Before trusting the solver, build the two objects by hand. (XᵀX)ⱼₖ is column j dotted with column k; column 0 is ones, column 1 is x.

Top-left = ones·ones = Σ1 = 5

Why: Five 1s dotted with five 1s. This entry is just n, the number of data points.

Off-diagonal = ones·x = Σx = 1+2+3+4+5 = 15

Why: The ones vector dotted with x sums x; symmetry puts 15 in both off-diagonal slots.

Bottom-right = x·x = Σx² = 1+4+9+16+25 = 55; and Xᵀy = [Σy, Σxy] = [20, 66]

Why: Σxy = 1·2 + 2·4 + 3·5 + 4·4 + 5·5 = 2+8+15+16+25 = 66. Verified by execution below.

\[ X^\top X = \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix},\quad X^\top y = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]

14. Work backwards from the answer: Route 1 — XᵀX and Xᵀy, entry by entry

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Bottom-right = x·x = Σx² = 1+4+9+16+25 = 55; and Xᵀy = [Σy, Σxy] = [20, 66]

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Before trusting the solver, build the two objects by hand. (XᵀX)ⱼₖ is column j dotted with column k; column 0 is ones, column 1 is x.

15. Guess the shape of the answer: Route 1 in code — solve the normal equations

Estimation

Predict first

Self-contained: re-import NumPy, re-define x, y, X, then solve XᵀX w = Xᵀy. Never invert — solve factors the system.

Commit before you compute: what does Route 1 in code — solve the normal equations come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: solve returns [2.2, 0.6]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Line 5 builds XᵀX and Xᵀy and factors the system in one call.

16. Route 1 in code — solve the normal equations

Worked example

Self-contained: re-import NumPy, re-define x, y, X, then solve XᵀX w = Xᵀy. Never invert — solve factors the system.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]          # bias column + feature
w = np.linalg.solve(X.T @ X, X.T @ y)
print(X.T @ X)                    # [[ 5 15] [15 55]]
print(X.T @ y)                    # [20 66]
print(w)                          # [2.2 0.6]

solve returns [2.2, 0.6]

Why: Line 5 builds XᵀX and Xᵀy and factors the system in one call. Verified: intercept 2.2, slope 0.6.

objectvalue (verified)
XᵀX[[5, 15], [15, 55]]
Xᵀy[20, 66]
w = [w₀, w₁][2.2, 0.6]

17. What each one costs: Route 1 in code — solve the normal equations

Trade off

Comparison matrix

From Route 1 in code — solve the normal equations: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

objectvalue (verified)
XᵀX[[5, 15], [15, 55]]
Xᵀy[20, 66]
w = [w₀, w₁][2.2, 0.6]

18. Why squared error is a probability statement

Intuition

Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.

Assume each score is the true line plus Gaussian noise. The Gaussian density has exp(−(error)²) in it, so its log has −(error)². Maximizing the log-likelihood is minimizing the sum of squared errors — exactly OLS.

So 'least squares' is quietly saying 'I believe the noise is Gaussian.' Change the noise model and the loss changes: Laplace noise → absolute error (L1).

19. By analogy: Why squared error is a probability statement

Analogy

Discussion prompt

Explain Why squared error is a probability statement by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.

20. What has to happen first: Route 2 — OLS is Gaussian MLE

Ranking

Put in order

Put the moves of Route 2 — OLS is Gaussian MLE into the order they have to happen.

  1. Model the noise as Gaussian: y = Xw + ε, ε ~ N(0, σ²)
  2. Write the negative log-likelihood
  3. Minimize over w
  4. Verify: the NLL is smallest exactly at w = [2.2, 0.6]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Lesson 8: assume each residual is independent Gaussian noise around the linear mean.

21. Route 2 — OLS is Gaussian MLE

Worked example

Model the noise as Gaussian: y = Xw + ε, ε ~ N(0, σ²)

Why: Lesson 8: assume each residual is independent Gaussian noise around the linear mean.

Write the negative log-likelihood

Why: Product of Gaussian densities → sum of log-densities. Constants aside, the −log-likelihood is proportional to the sum of squared residuals.

\[ -\log \mathcal{L}(w) = \tfrac{n}{2}\log(2\pi\sigma^2) + \frac{1}{2\sigma^2}\sum_i (y_i - x_i^\top w)^2 \]

Minimize over w

Why: Only the SSE term depends on w, and 1/(2σ²) is a positive constant, so argmin of the NLL = argmin of the SSE = the OLS solution.

\[ \arg\min_w \big[-\log\mathcal{L}(w)\big] = \arg\min_w \lVert Xw - y\rVert^2 \]

Verify: the NLL is smallest exactly at w = [2.2, 0.6]

Why: Perturbing w by [0.1, 0] raises the NLL from 5.7947 to 5.8197 — the OLS point is the maximum-likelihood point. Real execution below.

22. Draw the shape of it: Route 2 — OLS is Gaussian MLE

Blank canvas

Draw it

Draw what Route 2 — OLS is Gaussian MLE just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

23. What has to be given first: Route 2 in code — the MLE point is the OLS…

Missing information

Discussion prompt

Compute the OLS w, then show the Gaussian negative log-likelihood is minimized there — any nudge makes it larger.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Nudging w to [2.3, 0.6] raises the NLL to 5.8197. The likelihood peaks at the OLS coefficients — MLE and OLS are the same estimator under Gaussian noise.

24. Route 2 in code — the MLE point is the OLS point

Worked example

Compute the OLS w, then show the Gaussian negative log-likelihood is minimized there — any nudge makes it larger.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols = np.linalg.solve(X.T @ X, X.T @ y)
def nll(w, sigma=1.0):
    r = y - X @ w
    n = len(y)
    return 0.5*n*np.log(2*np.pi*sigma**2) + (r @ r)/(2*sigma**2)
print(w_ols)              # [2.2 0.6]
print(round(nll(w_ols), 4))                    # 5.7947
print(round(nll(w_ols + np.array([0.1, 0.0])), 4))  # 5.8197

nll(w_ols) = 5.7947 is the minimum

Why: Nudging w to [2.3, 0.6] raises the NLL to 5.8197. The likelihood peaks at the OLS coefficients — MLE and OLS are the same estimator under Gaussian noise.

w passed to nllnegative log-likelihood
[2.2, 0.6] (OLS)5.7947 (minimum)
[2.3, 0.6] (nudged)5.8197 (larger)
conclusionMLE = OLS under Gaussian noise

25. Plan first: Route 3 — the pseudoinverse (SVD)

Step zero

Discussion prompt

Route 3 — the pseudoinverse (SVD) — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Recall the SVD: X = UΣVᵀ (Lesson 16)

Answer:

  1. Recall the SVD: X = UΣVᵀ (Lesson 16)
  2. Define the Moore–Penrose pseudoinverse X⁺ = VΣ⁺Uᵀ
  3. The least-squares solution is w = X⁺y
  4. Verify: np.linalg.pinv(X) @ y = [2.2, 0.6]

26. Route 3 — the pseudoinverse (SVD)

Worked example

Recall the SVD: X = UΣVᵀ (Lesson 16)

Why: Every matrix factors this way; Σ holds the singular values σᵢ = √λᵢ(XᵀX) on its diagonal.

Define the Moore–Penrose pseudoinverse X⁺ = VΣ⁺Uᵀ

Why: Σ⁺ inverts each nonzero singular value. When X has full column rank, X⁺ = (XᵀX)⁻¹Xᵀ exactly.

\[ X^{+} = (X^\top X)^{-1} X^\top \quad\text{(full column rank)} \]

The least-squares solution is w = X⁺y

Why: Applying X⁺ to y projects onto the row space and returns the minimum-norm least-squares w — identical to the normal-equation answer, but computed via QR/SVD so it never squares the condition number.

\[ w = X^{+} y = \begin{bmatrix} 2.2 \\ 0.6 \end{bmatrix} \]

Verify: np.linalg.pinv(X) @ y = [2.2, 0.6]

Why: Confirmed by real execution below — the pseudoinverse route lands on the same coefficients as routes 1 and 2.

27. Say it in words: Route 3 — the pseudoinverse (SVD)

Translation

\( X^{+} = (X^\top X)^{-1} X^\top \quad\text{(full column rank)} \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

28. Predict the next row: Route 3 & 4 in code — pinv, lstsq, and the hat…

Pattern

Predict first

The table runs: pseudoinverse X⁺y | L16 | w = [2.2, 0.6] · lstsq (QR/SVD) | L16 | w = [2.2, 0.6]

In Route 3 & 4 in code — pinv, lstsq, and the hat matrix, given the rows so far: what is the next one — the row where route is hat matrix Hy?

Correct: hat matrix Hy | L25 | ŷ = [2.8, 3.4, 4.0, 4.6, 5.2]

routelessonoutput (verified)
pseudoinverse X⁺yL16w = [2.2, 0.6]
lstsq (QR/SVD)L16w = [2.2, 0.6]
hat matrix HyL25ŷ = [2.8, 3.4, 4.0, 4.6, 5.2]

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Hy is the projection of y onto the column space — the same predictions X@w produces.

29. Route 3 & 4 in code — pinv, lstsq, and the hat matrix

Worked example

Three more routes in one runnable block: pseudoinverse, lstsq, and the projection ŷ = Hy with H = X(XᵀX)⁻¹Xᵀ. All must agree.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_pinv  = np.linalg.pinv(X) @ y            # Route 3: SVD
w_lstsq = np.linalg.lstsq(X, y, rcond=None)[0]
H = X @ np.linalg.inv(X.T @ X) @ X.T       # Route 4: hat matrix
print(w_pinv.round(4))                      # [2.2 0.6]
print(w_lstsq.round(4))                     # [2.2 0.6]
print((H @ y).round(4))                     # [2.8 3.4 4. 4.6 5.2]

pinv and lstsq both give [2.2, 0.6]; Hy gives the fitted ŷ

Why: Hy is the projection of y onto the column space — the same predictions X@w produces. H is the linear operator that maps targets to fits.

routelessonoutput (verified)
pseudoinverse X⁺yL16w = [2.2, 0.6]
lstsq (QR/SVD)L16w = [2.2, 0.6]
hat matrix HyL25ŷ = [2.8, 3.4, 4.0, 4.6, 5.2]

30. Why projection is 'closest'

Intuition

Picture y as a point floating above a flat plane — the column space of X, every prediction the model can reach. Because no line fits all five points, y is off the plane.

The nearest reachable point is the perpendicular shadow of y on the plane. Any other point on the plane is farther, because the straight-down drop is the shortest path.

That shadow is Hy, and the leftover arrow y − Hy = r points straight out — perpendicular to every column, which is exactly Xᵀr = 0. Geometry and the normal equations are the same statement.

31. Teach it back: Why projection is 'closest'

Explain it

Discussion prompt

Explain Why projection is 'closest' to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The nearest reachable point is the perpendicular shadow of y on the plane. Any other point on the plane is farther, because the straight-down drop is the shortest path.

32. The hat matrix is a projection

Concept

H = X(XᵀX)⁻¹Xᵀ is not just a formula — it is an orthogonal projector onto the column space. Two properties prove it: H is symmetric and idempotent (H² = H), and trace(H) equals the number of parameters.

\[ H^2 = H,\quad H^\top = H,\quad \operatorname{trace}(H) = 2,\quad (I - H)y = r \]

Applying H twice does nothing new — once you are on the plane, projecting again keeps you there. trace(H) = 2 counts the effective degrees of freedom (bias + slope).

33. Restore the missing line: Verify the projection properties

Fill the middle

Fill in the blanks

From Verify the projection properties — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
H = X @ np.linalg.inv(X.T @ X) @ X.T
print(np.allclose(H, H @ H)) # True (idempotent)
print(np.allclose(H, H.T)) # True (symmetric)
print(round(np.trace(H), 4)) # 2.0
print((y - H @ y).round(4)) # residuals

Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. All four checks pass. The residual (I−H)y sticks straight out of the column space — the geometric signature of the least-squares fit.

34. Verify the projection properties

Worked example

Confirm idempotence, symmetry, trace = 2, and that (I − H)y reproduces the residuals we know are [−0.8, 0.6, 1.0, −0.6, −0.2].

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
H = X @ np.linalg.inv(X.T @ X) @ X.T
print(np.allclose(H, H @ H))      # True  (idempotent)
print(np.allclose(H, H.T))        # True  (symmetric)
print(round(np.trace(H), 4))      # 2.0
print((y - H @ y).round(4))       # residuals

H² = H, Hᵀ = H, trace = 2.0, (I−H)y = residuals

Why: All four checks pass. The residual (I−H)y sticks straight out of the column space — the geometric signature of the least-squares fit.

propertyvalue (verified)
H == H@H (idempotent)True
H == Hᵀ (symmetric)True
trace(H)2.0
(I − H)y[−0.8, 0.6, 1.0, −0.6, −0.2]

35. Which is which, by value (verified)

Discrimination

Sort into buckets

Sort these by value (verified), from memory, without looking back at Verify the projection properties. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

True
H == H@H (idempotent); H == Hᵀ (symmetric)
2.0
trace(H)
[−0.8, 0.6, 1.0, −0.6, −0.2]
(I − H)y
g1
value (verified) is "True" for H == H@H (idempotent), H == Hᵀ (symmetric) — that is what the table on "Verify the projection properties" records, and it is the single property separating this group from the rest.
g2
value (verified) is "2.0" for trace(H) — that is what the table on "Verify the projection properties" records, and it is the single property separating this group from the rest.
g3
value (verified) is "[−0.8, 0.6, 1.0, −0.6, −0.2]" for (I − H)y — that is what the table on "Verify the projection properties" records, and it is the single property separating this group from the rest.

36. The fit in numbers — residuals and R²

Worked example

With w = [2.2, 0.6], predict each student, subtract from the true score, and read off the residuals and R² — the numbers the exam asks you to produce.

xyŷ = 2.2 + 0.6xr = y − ŷr²
122.8−0.80.64
243.4+0.60.36
354.0+1.01.00
444.6−0.60.36
555.2−0.20.04

SSres = Σr² = 2.4, SStot = Σ(y−4)² = 6.0, R² = 1 − 2.4/6.0 = 0.6

Why: ȳ = 20/5 = 4; deviations [−2,0,1,0,1] squared and summed give 6.0. The line removes 60% of that baseline error. Also note Σr = 0 — the bias column forces zero mean residual.

Xᵀr = [0, 0] — the perpendicularity check

Why: Row 0 (ones·r): −0.8+0.6+1.0−0.6−0.2 = 0. Row 1 (x·r): −0.8+1.2+3.0−2.4−1.0 = 0. The residual sticks straight out of col(X), confirming w is the projection.

37. Watch it run: The fit in numbers — residuals and R²

Pattern

Step through it

Step through The fit in numbers — residuals and R² one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: x is 1
  2. Step 2: x is 2
  3. Step 3: x is 3
  4. Step 4: x is 4
  5. Step 5: x is 5

38. The four routes, one truth

Concept

Four Phase-1 lessons, four derivations, one answer. This equivalence IS the synthesis the mock exam probes.

routelessonthe claimw
normal equations XᵀXw = XᵀyL7residual ⟂ columns[2.2, 0.6]
Gaussian MLEL8−log L ∝ SSE[2.2, 0.6]
pseudoinverse X⁺yL16SVD inverts Σ[2.2, 0.6]
hat-matrix projection HyL25H projects onto col(X)[2.2, 0.6]

Say it aloud: geometry, probability, and optimization all point at the same [2.2, 0.6]. If you can explain why, Phase 1 is consolidated.

39. The four pillars, re-derived

Section

Part 2 of 5

40. Pillar 1 — Linear Algebra

Concept

resultthe one line
eigen (L10)Av = λv; symmetric ⇒ real λ, orthogonal vectors
SVD (L16)A = UΣVᵀ for ANY matrix; σ = √eig(AᵀA)
PSD/PD (L13)xᵀAx ≥ 0; XᵀX always PSD; PD ⟺ full rank
projection (L25)ŷ = Hy, H = X(XᵀX)⁻¹Xᵀ, idempotent

We just met H and PSD via OLS. Next we pin down eigenvalues, PD, and singular values with reproducible numbers.

41. What has to happen first: Eigenvalues by hand, then PD

Ranking

Put in order

Put the moves of Eigenvalues by hand, then PD into the order they have to happen.

  1. Write det(A − λI) = 0
  2. Expand to λ² − 4λ + 3 = 0
  3. Verify: both λ > 0, so A is positive definite

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Eigenvalues are the λ that make A − λI singular.

42. Eigenvalues by hand, then PD

Worked example

Take A = [[2, 1], [1, 2]] — a symmetric matrix that shows up as a mini covariance. Find its eigenvalues from the characteristic polynomial, no solver.

Write det(A − λI) = 0

Why: Eigenvalues are the λ that make A − λI singular.

\[ \det\!\begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} = (2-\lambda)^2 - 1 = 0 \]

Expand to λ² − 4λ + 3 = 0

Why: (2−λ)² − 1 = 4 − 4λ + λ² − 1 = λ² − 4λ + 3. Note trace = 4 = sum of roots, det = 3 = product.

\[ \lambda^2 - 4\lambda + 3 = (\lambda - 1)(\lambda - 3) = 0 \;\Rightarrow\; \lambda = 1,\, 3 \]

Verify: both λ > 0, so A is positive definite

Why: np.roots([1,−4,3]) returns [3, 1] and eigh returns [1, 3] — matching the hand result. All-positive eigenvalues ⇒ PD (L13).

43. Decode the notation: Eigenvalues by hand, then PD

Notation

Annotate

From Eigenvalues by hand, then PD — read this one piece at a time. What is each part doing?

On: \( \lambda^2 - 4\lambda + 3 = (\lambda - 1)(\lambda - 3) = 0 \;\Rightarrow\; \lambda = 1,\, 3 \)

  • Eigenvalues are the λ that make A − λI singular.
  • (2−λ)² − 1 = 4 − 4λ + λ² − 1 = λ² − 4λ + 3. Note trace = 4 = sum of roots, det = 3 = product.
  • np.roots([1,−4,3]) returns [3, 1] and eigh returns [1, 3] — matching the hand result. All-positive eigenvalues ⇒ PD (L13).

44. Eigen + PSD + dominant vector — the 60-second drill

Worked example

Exam pace: one problem touching three lessons. For A = [[2,1],[1,2]], give eigenvalues, decide PD, and name the dominant eigenvector (what power iteration converges to).

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
vals, vecs = np.linalg.eigh(A)     # ascending eigenvalues
print(vals.round(4))               # [1. 3.]
print(bool(vals.min() > 0))        # True  (PD)
print(vecs[:, -1].round(4))        # dominant eigenvector
print((A @ vecs[:, -1]).round(4))  # = lambda * v

λ = [1, 3], PD = True, dominant eigenvector = [0.7071, 0.7071]

Why: The largest eigenvalue is 3; its eigenvector [0.7071, 0.7071] is the direction power iteration finds. Av = [2.1213, 2.1213] = 3·v confirms it.

questionanswer (verified)lesson
eigenvalues[1, 3]L10
positive definite?yes (all λ > 0)L13
dominant eigenvector[0.7071, 0.7071]L10
Av vs λv check[2.1213, 2.1213] bothL10

45. Singular values, even without eigenvalues

Intuition

Eigenvalues only exist for square matrices. But a 3×2 data matrix isn't square — so what measures its 'stretch'?

The SVD answers: A = UΣVᵀ factors any matrix into a rotation, a stretch by the singular values σ, and another rotation. The σ are the amounts A stretches its input directions.

And AᵀA is square and PSD, so its eigenvalues exist and are ≥ 0. The singular values are their square roots — the bridge from the non-square world back to eigenvalues.

46. SVD — singular values of a non-square matrix

Worked example

The exam loves σ = √eig(AᵀA). Test it on a 3×2 matrix that has no eigenvalues of its own — only singular values exist.

import numpy as np
A = np.array([[3., 0.], [0., 2.], [1., 1.]])   # 3x2, non-square
U, S, Vt = np.linalg.svd(A, full_matrices=False)
eigAtA = np.linalg.eigvalsh(A.T @ A)            # ascending
print(S.round(4))                               # [3.1926 2.1926]
print(np.sqrt(eigAtA[::-1]).round(4))           # [3.1926 2.1926]
print(np.allclose(A, U @ np.diag(S) @ Vt))      # True

S = [3.1926, 2.1926] = √eig(AᵀA), and UΣVᵀ rebuilds A

Why: AᵀA is (2×2) and PSD, so its eigenvalues are ≥ 0 and their roots are the singular values — even though the non-square A has no eigenvalues. Reconstruction is exact.

quantityvalue (verified)
singular values S[3.1926, 2.1926]
√eig(AᵀA) (desc)[3.1926, 2.1926]
UΣVᵀ == ATrue

47. Pillar 2 — Probability & Information

Concept

resultthe one line
MLE (L8)argmax Σ log p; losses are negative log-likelihoods
MAP (L12)MLE + log-prior; Gaussian→L2, Laplace→L1
MVN (L23)N(μ,Σ), sample x = μ + Lz (Cholesky)
info theory (L14/29)KL ≥ 0; cross-entropy = H + KL; MI = 0 ⟺ independent

We already used MLE for Route 2. Now nail the information-theory identity CE = H + KL with reproducible numbers.

48. Cross-entropy = surprise you can't avoid + surprise you added

Intuition

Entropy H(p) is the irreducible average surprise of the true distribution — the best any model could hope for. It doesn't depend on your model at all.

KL divergence D(p‖q) is the extra surprise from using the wrong distribution q instead of p. It is never negative — a wrong model can only cost you.

Cross-entropy is their sum. Since H(p) is fixed, minimizing cross-entropy loss over q is minimizing KL — driving your model toward the data. That's why classifiers train on cross-entropy.

49. Finish it with less help: KL ≥ 0 and cross-entropy = H + KL

Faded example

Fill in the blanks

KL ≥ 0 and cross-entropy = H + KL, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
p = np.array([0.7, 0.2, 0.1]) # true distribution
q = np.array([0.5, 0.3, 0.2]) # model distribution
H = -np.sum(p * np.log(p)) # entropy of p
KL = np.sum(p * np.log(p / q)) # KL(p || q) >= 0
CE = **-np.sum(p * np.log(q))** # cross-entropy H(p,q)
print(round(H, 4), round(KL, 4), round(CE, 4))
print(np.isclose(CE, H + KL)) # True

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Minimizing cross-entropy over q minimizes KL (H(p) is fixed), so CE-loss drives the model toward the data distribution.

50. KL ≥ 0 and cross-entropy = H + KL

Worked example

Take a true distribution p and a model q. Verify the identity that every training loss rests on: cross-entropy splits into the data's own entropy plus the KL gap.

\[ \underbrace{H(p,q)}_{\text{cross-entropy}} = \underbrace{H(p)}_{\text{entropy}} + \underbrace{D_{KL}(p\,\Vert\, q)}_{\ge\, 0} \]

import numpy as np
p = np.array([0.7, 0.2, 0.1])          # true distribution
q = np.array([0.5, 0.3, 0.2])          # model distribution
H  = -np.sum(p * np.log(p))            # entropy of p
KL =  np.sum(p * np.log(p / q))        # KL(p || q) >= 0
CE = -np.sum(p * np.log(q))            # cross-entropy H(p,q)
print(round(H, 4), round(KL, 4), round(CE, 4))
print(np.isclose(CE, H + KL))          # True

H = 0.8018, KL = 0.0851 ≥ 0, CE = 0.8869 = H + KL

Why: Minimizing cross-entropy over q minimizes KL (H(p) is fixed), so CE-loss drives the model toward the data distribution. KL is never negative — that is why CE ≥ H always.

quantityvalue (verified)
H(p) entropy0.8018
KL(p‖q)0.0851 (≥ 0)
CE(p,q)0.8869
H + KL0.8869 (= CE)

51. Watch it run: KL ≥ 0 and cross-entropy = H + KL

Pattern

Step through it

Step through KL ≥ 0 and cross-entropy = H + KL one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is H(p) entropy
  2. Step 2: quantity is KL(p‖q)
  3. Step 3: quantity is CE(p,q)
  4. Step 4: quantity is H + KL

52. Pillar 3 — Calculus & Optimization

Concept

resultthe one line
gradient (L6/9)∇ points uphill; backprop = chain rule of Jacobians
convexity (L11)Hessian PSD everywhere ⇒ only global minima
optimizers (L15)Adam = momentum + RMSprop + bias correction
Newton (L21)θ −= H⁻¹∇f; quadratic convergence, O(d³)

OLS is a convex quadratic, so both gradient descent and Newton must converge to the closed-form [2.2, 0.6]. We watch each do it — and watch GD blow up when the step is too big.

53. Why OLS is a friendly landscape

Concept

The SSE loss ‖Xw − y‖² has Hessian ∇²L = 2XᵀX, which is PSD everywhere (zᵀXᵀXz = ‖Xz‖² ≥ 0). A PSD-everywhere Hessian means the loss is convex — one global bowl, no bad local minima.

\[ \nabla^2 L = 2X^\top X \succeq 0 \;\Rightarrow\; L \text{ convex} \;\Rightarrow\; \nabla L = 0 \text{ is the global min} \]

So any descent method that keeps moving downhill must reach the same w★. Gradient descent and Newton below are two such methods — and they land on the identical [2.2, 0.6].

54. Guess the shape of the answer: Gradient descent converges to the OLS w

Estimation

Predict first

The SSE gradient is ∇L = 2Xᵀ(Xw − y). Step downhill from zero; with a small enough learning rate it lands on the closed-form solution.

Commit before you compute: what does Gradient descent converges to the OLS w come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: After 20000 steps at lr = 0.001, GD w = [2.2, 0.6]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Convexity guarantees a single global minimum; GD slides into it and matches the closed form.

55. Gradient descent converges to the OLS w

Worked example

The SSE gradient is ∇L = 2Xᵀ(Xw − y). Step downhill from zero; with a small enough learning rate it lands on the closed-form solution.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.zeros(2)
lr = 0.001
for t in range(20000):
    grad = 2 * X.T @ (X @ w - y)     # gradient of SSE
    w = w - lr * grad
print(w.round(4))                    # [2.2 0.6]
print(np.linalg.solve(X.T @ X, X.T @ y).round(4))  # [2.2 0.6]

After 20000 steps at lr = 0.001, GD w = [2.2, 0.6]

Why: Convexity guarantees a single global minimum; GD slides into it and matches the closed form. The gradient at the optimum is zero, so the update stops moving.

methodw (verified)
gradient descent (lr 0.001)[2.2, 0.6]
closed-form solve[2.2, 0.6]
difference< 1e-3

56. Fill in: w (verified) for Gradient descent converges to the OLS w

Comparison

Comparison matrix

From Gradient descent converges to the OLS w: refill the w (verified) column from what you know. The rest of the table is as it appeared.

methodw (verified)
gradient descent (lr 0.001)[2.2, 0.6]
closed-form solve[2.2, 0.6]
difference< 1e-3

57. Something is wrong here: a learning rate above 2/L diverges

Anomaly

Predict first

A student writes this, and it looks reasonable:

Bigger steps should learn faster, so crank the learning rate to lr = 0.02 for the same SSE gradient descent.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118.

Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.

Why: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.

58. Trap: a learning rate above 2/L diverges

Trap

The trap

Bigger steps should learn faster, so crank the learning rate to lr = 0.02 for the same SSE gradient descent.

lr = 0.02 → w explodes past 1e300 to inf, NaN territory

Why: The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.

The fix

Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.

lr = 0.001 < 2/L → w converges to [2.2, 0.6]

Why: Below the 2/L threshold every step contracts toward the minimum. The convex bowl has one bottom; a safe step size reaches it. Always bound the step by the curvature.

59. Break it on purpose: a learning rate above 2/L diverges

Break the constraint

Discussion prompt

The rule this trap just fixed:

Pick lr < 2/L where L is the largest eigenvalue of the Hessian 2XᵀX.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

The Hessian is 2XᵀX with largest eigenvalue L ≈ 118. Stability needs lr < 2/L ≈ 0.017. At 0.02 each step overshoots and the iterate grows geometrically without bound — verified: after enough steps w overflows to inf.

60. Restore the missing line: Newton solves OLS in ONE step

Fill the middle

Fill in the blanks

From Newton solves OLS in ONE step — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.array([0.0, 0.0]) # any start
grad = 2 * X.T @ (X @ w - y) # gradient at w
Hess = 2 * X.T @ X # Hessian (constant for OLS)
w_new = w - np.linalg.solve(Hess, grad) # one Newton step
print(w_new.round(4)) # [2.2 0.6]

Why: w is what everything below it consumes, so the wrong expression here fails later and somewhere else. For a quadratic the Hessian is constant and the second-order model is exact, so a single Newton step is the closed-form solution.

61. Newton solves OLS in ONE step

Worked example

OLS is exactly quadratic, so Newton's method — w ← w − H⁻¹∇ — jumps to the exact minimum in a single step from any start, because the quadratic model is the function itself.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.array([0.0, 0.0])           # any start
grad = 2 * X.T @ (X @ w - y)       # gradient at w
Hess = 2 * X.T @ X                 # Hessian (constant for OLS)
w_new = w - np.linalg.solve(Hess, grad)   # one Newton step
print(w_new.round(4))              # [2.2 0.6]

One Newton step from [0, 0] lands on [2.2, 0.6]

Why: For a quadratic the Hessian is constant and the second-order model is exact, so a single Newton step is the closed-form solution. This is why Newton has quadratic convergence — but each step costs O(d³) to solve.

stepw (verified)
start[0, 0]
after 1 Newton step[2.2, 0.6]
closed form[2.2, 0.6]

62. Pillar 4 — ML Foundations

Concept

resultthe one line
losses (L17)softmax+CE ⇒ dL/dz = p − y
bias-variance (L32)MSE = Bias² + Variance + Noise (U-curve)
metrics (L26/35)imbalance ⇒ PR-AUC; AUC = P(pos > neg)
SVM/kernels (L28/30)dual in α; k = φ·φ; Mercer ⇒ PSD Gram

The single most exam-tested gradient is softmax + cross-entropy. We derive dL/dz = p − y and confirm it against a numerical gradient.

63. softmax + cross-entropy ⇒ dL/dz = p − y

Worked example

The clean gradient that makes classification trainable: the derivative of cross-entropy with respect to the logits is just predicted probability minus the one-hot label.

\[ L = -\sum_k y_k \log p_k,\quad p = \operatorname{softmax}(z) \;\Rightarrow\; \frac{\partial L}{\partial z} = p - y \]

import numpy as np
def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()
z = np.array([2.0, 1.0, 0.1])      # logits
y = np.array([1.0, 0.0, 0.0])      # one-hot label
p = softmax(z)
print(p.round(4))                  # [0.659 0.2424 0.0986]
print((p - y).round(4))            # analytic gradient
eps = 1e-6
g = np.array([( -np.sum(y*np.log(softmax(z+eps*np.eye(3)[i])))
               +np.sum(y*np.log(softmax(z-eps*np.eye(3)[i]))) )/(2*eps)
              for i in range(3)])
print(g.round(4))                  # numerical gradient (matches)

p − y = [−0.341, 0.2424, 0.0986] matches the numerical gradient

Why: The messy softmax Jacobian and the cross-entropy derivative cancel into p − y. The finite-difference check agrees to 1e-5 — the analytic shortcut is correct, which is why frameworks fuse softmax+CE.

quantityvalue (verified)
p = softmax(z)[0.6590, 0.2424, 0.0986]
p − y (analytic)[−0.3410, 0.2424, 0.0986]
numerical ∂L/∂z[−0.3410, 0.2424, 0.0986]
matchTrue (atol 1e-5)

64. Which is which, by value (verified)

Discrimination

Sort into buckets

Sort these by value (verified), from memory, without looking back at softmax + cross-entropy ⇒ dL/dz = p − y. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

[0.6590, 0.2424, 0.0986]
p = softmax(z)
[−0.3410, 0.2424, 0.0986]
p − y (analytic); numerical ∂L/∂z
True (atol 1e-5)
match
g1
value (verified) is "[0.6590, 0.2424, 0.0986]" for p = softmax(z) — that is what the table on "softmax + cross-entropy ⇒ dL/dz = p − y" records, and it is the single property separating this group from the rest.
g2
value (verified) is "[−0.3410, 0.2424, 0.0986]" for p − y (analytic), numerical ∂L/∂z — that is what the table on "softmax + cross-entropy ⇒ dL/dz = p − y" records, and it is the single property separating this group from the rest.
g3
value (verified) is "True (atol 1e-5)" for match — that is what the table on "softmax + cross-entropy ⇒ dL/dz = p − y" records, and it is the single property separating this group from the rest.

65. Three ways to be wrong

Intuition

Test error splits into three sources. Bias is being systematically off — a too-simple model that can't capture the shape. Variance is being jumpy — a too-flexible model that chases the particular training noise.

Noise is the irreducible floor: even the perfect model can't predict the random part of y. No amount of cleverness removes it.

The trade: making a model more flexible lowers bias but raises variance. The sweet spot minimizes their sum — the U-shaped test-error curve. The simulation next reads off all three.

66. Bias-variance, simulated

Worked example

MSE = Bias² + Variance + Noise. Estimate a constant model's prediction at a point over many noisy training sets and read off the three pieces.

\[ \mathbb{E}\big[(y - \hat f)^2\big] = \underbrace{(\mathbb{E}\hat f - f)^2}_{\text{Bias}^2} + \underbrace{\operatorname{Var}(\hat f)}_{\text{Variance}} + \underbrace{\sigma^2}_{\text{Noise}} \]

import numpy as np
np.random.seed(0)
f_true, sigma = 2.0, 1.0
preds = []
for _ in range(20000):
    sample = f_true + np.random.randn(5)*sigma   # 5 noisy points
    preds.append(sample.mean())                  # constant model
preds = np.array(preds)
bias2 = (preds.mean() - f_true)**2
var   = preds.var()
print(round(bias2, 4), round(var, 4), round(sigma**2, 4))
print(round(bias2 + var + sigma**2, 4))          # total MSE

Bias² ≈ 0.0, Variance ≈ 0.1984, Noise = 1.0 → MSE ≈ 1.1984

Why: The mean estimator is unbiased (Bias² ≈ 0) with variance σ²/5 = 0.2, plus irreducible noise σ² = 1. seed(0) makes these reproducible. Reducing model bias trades against rising variance — the U-curve.

componentvalue (verified, seed 0)
Bias²0.0 (unbiased mean)
Variance0.1984 (≈ σ²/5)
Noise σ²1.0
MSE = B² + V + N1.1984

67. What ROC-AUC actually is

Concept

ROC-AUC has a clean probabilistic meaning: it is the chance that a random positive gets a higher score than a random negative.

\[ \text{ROC-AUC} = \mathbb{P}\big(\text{score(pos)} > \text{score(neg)}\big) \]

The catch: its false-positive rate divides by the number of negatives. Under 2% positives there are ~49× more negatives, so a flood of false positives barely moves the rate — ROC-AUC stays flattering. PR-AUC, which uses precision, does not forgive that.

68. Finish it with less help: ROC-AUC vs PR-AUC under 2% imbalance

Faded example

Fill in the blanks

ROC-AUC vs PR-AUC under 2% imbalance, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
np.random.seed(0)
n = 5000
y = (np.random.rand(n) < 0.02).astype(int) # ~2% positives
scores = *np.random.rand(n) + 0.15y** # weak signal
print(int(y.sum()), 'of', n) # 105 of 5000
print(round(roc_auc_score(y, scores), 4)) # 0.6369
print(round(average_precision_score(y, scores), 4)) # 0.1771

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. ROC-AUC's false-positive rate has a huge true-negative denominator, so it stays flattering.

69. ROC-AUC vs PR-AUC under 2% imbalance

Worked example

On a rare-positive problem, ROC-AUC looks optimistic while PR-AUC tells the truth about the minority class. Simulate 2% positives with a weak scorer.

import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
np.random.seed(0)
n = 5000
y = (np.random.rand(n) < 0.02).astype(int)   # ~2% positives
scores = np.random.rand(n) + 0.15*y           # weak signal
print(int(y.sum()), 'of', n)                  # 105 of 5000
print(round(roc_auc_score(y, scores), 4))     # 0.6369
print(round(average_precision_score(y, scores), 4))  # 0.1771

ROC-AUC = 0.6369 but PR-AUC = 0.1771 on the same scores

Why: ROC-AUC's false-positive rate has a huge true-negative denominator, so it stays flattering. PR-AUC watches precision on the rare positives and exposes the weakness. Under imbalance, report PR-AUC.

metricvalue (verified, seed 0)
positives105 of 5000 (2.1%)
ROC-AUC0.6369 (optimistic)
PR-AUC0.1771 (honest)
base rate0.021

70. Traps & synthesis

Section

Part 3 of 5

71. The dependency map

Concept

The threads weave together: eigenvalues → SVD → PCA; MLE → cross-entropy → the training loop; gradients → backprop → optimizers; PSD → kernels → SVM.

Nothing in Phase 1 stands alone — every result is a prerequisite for Phase 2 (classical ML) and beyond (transformers, generative models). OLS-four-ways is the smallest example of the whole web.

72. Something is wrong here: the high-yield exam misreads

Anomaly

Predict first

A student writes this, and it looks reasonable:

Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.

You've beaten all three this phase — keep them straight.

Why: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.

73. Trap: the high-yield exam misreads

Trap

The trap

Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.

Carry these three misconceptions into the exam

Why: All three are the classic traps — the p-value reverses the conditioning, det>0 allows two negative eigenvalues (their product is positive), and MSE's gradient vanishes on confident wrong predictions.

The fix

You've beaten all three this phase — keep them straight.

p = P(data or more extreme | H₀) (L20); PD needs ALL leading minors > 0 (L13); use cross-entropy for classification (L17)

Why: The p-value conditions on H₀, not the reverse. det>0 is necessary but not sufficient for PD. Cross-entropy gives the clean p − y gradient; MSE saturates. Name the correct statement for each cold.

74. Which of these survive contact with Lesson 37: Phase 1 Review & Mock Prep?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Design matrix X = [1, x] (a bias column then the feature), target y. Keep these five numbers in your head — every slide reuses them.; Phase 1 taught linear algebra, probability, and optimization as separate weeks. The deep truth is that they are one idea wearing four costumes.; Route 1 just chose to square the residuals. Route 2 explains why squaring is the right penalty — it falls out of a noise model, not a whim.
Breaks
Bigger steps should learn faster, so crank the learning rate to lr = 0.02 for the same SSE gradient descent.; Quick recall: a p-value of 0.03 means a 3% chance the null is true; det > 0 means positive definite; and MSE works fine for classification.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 37: Phase 1 Review & Mock Prep puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

75. Predict the next row: Trap in code — det > 0 is NOT positive definite

Pattern

Predict first

The table runs: det(A) | 1.0 > 0 | misleadingly 'passes' · eigenvalues | [−1, −1] | both < 0 → NOT PD

In Trap in code — det > 0 is NOT positive definite, given the rows so far: what is the next one — the row where test is xᵀAx at [1,1]?

Correct: xᵀAx at [1,1] | −2.0 | confirms not PD

testresultverdict
det(A)1.0 > 0misleadingly 'passes'
eigenvalues[−1, −1]both < 0 → NOT PD
xᵀAx at [1,1]−2.0confirms not PD

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Two negatives multiply to a positive determinant, so det>0 is fooled.

76. Trap in code — det > 0 is NOT positive definite

Worked example

A concrete counterexample: −I has determinant +1 (two negative eigenvalues multiply to positive) but is negative definite. Determinant alone cannot certify PD.

import numpy as np
A = np.array([[-1., 0.], [0., -1.]])
print(round(np.linalg.det(A), 4))   # 1.0  (positive!)
print(np.linalg.eigvalsh(A))        # [-1. -1.]
x = np.array([1., 1.])
print(x @ A @ x)                    # -2.0  (< 0, so NOT PD)

det(A) = 1.0 > 0 yet eigenvalues are [−1, −1] and xᵀAx = −2

Why: Two negatives multiply to a positive determinant, so det>0 is fooled. The honest PD test is 'all eigenvalues > 0' or 'all leading minors > 0'. Here xᵀAx < 0 proves A is not PD.

testresultverdict
det(A)1.0 > 0misleadingly 'passes'
eigenvalues[−1, −1]both < 0 → NOT PD
xᵀAx at [1,1]−2.0confirms not PD

77. Rebuild the recipe: The exam-pace retrieval protocol

Ranking

Put in order

These are the steps of The exam-pace retrieval protocol, scrambled. Put them back in order before the next slide shows you.

  1. Cold attempt first — slides closed, paper out (the struggle is the encoding)
  2. Then check — reveal only after a genuine attempt
  3. Reconstruct aloud — re-derive the result with the answer hidden again
  4. Log every gap — any hint or miss goes on the retrieval list
  5. Re-quiz on a schedule — 3 days, 1 week, 2 weeks (spaced retrieval)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

78. The exam-pace retrieval protocol

Pattern

  1. Cold attempt first — slides closed, paper out (the struggle is the encoding)
  2. Then check — reveal only after a genuine attempt
  3. Reconstruct aloud — re-derive the result with the answer hidden again
  4. Log every gap — any hint or miss goes on the retrieval list
  5. Re-quiz on a schedule — 3 days, 1 week, 2 weeks (spaced retrieval)

Apply this to the five drills next. Struggle, then check — the drills after this point are your cold attempts.

79. Where does each piece belong: Lesson 37: Phase 1 Review & Mock Prep

Sorting

Sort into buckets

These are the pieces of Lesson 37: Phase 1 Review & Mock Prep, out of order. Put each one back under the part of the lesson it belongs to.

One dataset, four routes to OLS
The running dataset; Why 'four ways' is the whole point; Route 1 — the normal equations
The four pillars, re-derived
Pillar 1 — Linear Algebra; Eigenvalues by hand, then PD; Eigen + PSD + dominant vector — the 60-second drill
Traps & synthesis
The dependency map; Trap in code — det > 0 is NOT positive definite; The exam-pace retrieval protocol
s1
One dataset, four routes to OLS is where Lesson 37: Phase 1 Review & Mock Prep puts The running dataset, Why 'four ways' is the whole point, Route 1 — the normal equations. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The four pillars, re-derived is where Lesson 37: Phase 1 Review & Mock Prep puts Pillar 1 — Linear Algebra, Eigenvalues by hand, then PD, Eigen + PSD + dominant vector — the 60-second drill. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Traps & synthesis is where Lesson 37: Phase 1 Review & Mock Prep puts The dependency map, Trap in code — det > 0 is NOT positive definite, The exam-pace retrieval protocol. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

80. Retrieval drills

Section

Part 4 of 5 — slides closed

81. Rule out three: Drill — the four routes

Elimination

Eliminate the wrong options

Why do the normal equations, Gaussian MLE, the pseudoinverse, and the hat-matrix projection all return the same w = [2.2, 0.6]?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. They are four descriptions of the same optimum: minimizing squared error = projecting y onto col(X) = maximizing the Gaussian likelihood
  • B. It is a numerical coincidence specific to this 5-point dataset
  • C. Only the normal equations are exact; the other three are approximations that happened to round to the same value
  • D. They agree only because the data is perfectly linear

Survives elimination: A

Why: Least squares has one minimizer. The perpendicular-residual condition (projection), ∇SSE = 0 (normal equations), X⁺y (SVD), and the Gaussian MLE all characterize that same point, so they must coincide exactly when X has full column rank.

82. Drill — the four routes

Check

Cold attempt, then click. This is the capstone idea of the whole session.

Check your understanding

Why do the normal equations, Gaussian MLE, the pseudoinverse, and the hat-matrix projection all return the same w = [2.2, 0.6]?

  • A. They are four descriptions of the same optimum: minimizing squared error = projecting y onto col(X) = maximizing the Gaussian likelihood (correct)
  • B. It is a numerical coincidence specific to this 5-point dataset
  • C. Only the normal equations are exact; the other three are approximations that happened to round to the same value
  • D. They agree only because the data is perfectly linear

Answer: A

Why: Least squares has one minimizer. The perpendicular-residual condition (projection), ∇SSE = 0 (normal equations), X⁺y (SVD), and the Gaussian MLE all characterize that same point, so they must coincide exactly when X has full column rank.

Why B tempts people
Not a coincidence — the equivalence is a theorem that holds for any full-column-rank X, not just these five points.
Why C tempts people
All four are exact (in exact arithmetic). pinv/lstsq differ only in floating-point stability, not in the answer they target.
Why D tempts people
The data is NOT perfectly linear here — the minimum error is 2.4, not 0. The routes agree precisely in the noisy, no-exact-fit case that least squares is built for.

83. Drill — Linear Algebra

Check

Cold attempt, then click.

Check your understanding

The singular values of any matrix A are:

  • A. the non-negative square roots of the eigenvalues of AᵀA (correct)
  • B. the eigenvalues of A
  • C. the diagonal entries of A
  • D. always equal to the eigenvalues in magnitude

Answer: A

Why: σᵢ = √λᵢ(AᵀA). AᵀA is PSD so the σ are real and ≥ 0, and the SVD exists for ANY matrix (L16). We verified [3.1926, 2.1926] = √eig(AᵀA) for a 3×2 A.

Why B tempts people
Non-square matrices have no eigenvalues yet always have singular values; even for square A the two generally differ.
Why C tempts people
Only true for an already-diagonal matrix; in general the diagonal is unrelated to the σ.
Why D tempts people
|λ| = σ only for normal matrices (e.g. symmetric); generally they differ.

84. Answer it before you see the options: Drill — Probability

Prediction

Predict first

Minimizing mean squared error is maximum likelihood under which assumption?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Gaussian noise with constant variance

Why: With y = f(x) + ε, ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² — exactly MSE (L8). We saw the NLL bottom out at the OLS w = [2.2, 0.6].

85. Drill — Probability

Check

Slides closed first.

Check your understanding

Minimizing mean squared error is maximum likelihood under which assumption?

  • A. Gaussian noise with constant variance (correct)
  • B. Laplace noise
  • C. Bernoulli labels
  • D. no assumption — MSE is always the likelihood

Answer: A

Why: With y = f(x) + ε, ε ~ N(0,σ²), the negative log-likelihood reduces to (1/2σ²)Σ(y−f)² — exactly MSE (L8). We saw the NLL bottom out at the OLS w = [2.2, 0.6].

Why B tempts people
Laplace noise gives the L1 / mean-absolute-error loss, not squared error.
Why C tempts people
Bernoulli labels give cross-entropy (classification), not MSE.
Why D tempts people
MSE corresponds to a specific (Gaussian) noise model; other models give other losses.

86. Answer it before you see the options: Drill — Optimization

Prediction

Predict first

You run gradient descent on the OLS loss with learning rate 0.02 and the weights explode to infinity. The largest eigenvalue of the Hessian 2XᵀX is about 118. What is the fix?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Lower the learning rate below 2/L ≈ 0.017

Why: Gradient descent on a quadratic is stable only when lr < 2/L, where L is the largest Hessian eigenvalue. Here 2/L ≈ 0.017, so lr = 0.02 overshoots every step and diverges; lr = 0.001 converges to [2.2, 0.6].

87. Drill — Optimization

Check

Attempt before clicking.

Check your understanding

You run gradient descent on the OLS loss with learning rate 0.02 and the weights explode to infinity. The largest eigenvalue of the Hessian 2XᵀX is about 118. What is the fix?

  • A. Lower the learning rate below 2/L ≈ 0.017 (correct)
  • B. Add more iterations — it will eventually converge
  • C. Switch the loss to cross-entropy
  • D. The loss is non-convex, so use a random restart

Answer: A

Why: Gradient descent on a quadratic is stable only when lr < 2/L, where L is the largest Hessian eigenvalue. Here 2/L ≈ 0.017, so lr = 0.02 overshoots every step and diverges; lr = 0.001 converges to [2.2, 0.6].

Why B tempts people
More steps of a divergent iteration diverge faster, not slower — the iterate grows geometrically each step.
Why C tempts people
Cross-entropy is for classification; it does not fit this regression and does not address the step-size instability.
Why D tempts people
The OLS loss is convex (Hessian 2XᵀX is PSD everywhere) — there are no bad local minima. The problem is purely the step size.

88. Rule out three: Drill — ML Foundations

Elimination

Eliminate the wrong options

On a 2%-positive dataset your model gets ROC-AUC 0.64 but PR-AUC 0.18. Which metric better reflects minority-class performance, and why?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. PR-AUC — its precision axis exposes false positives among the rare positives, which ROC-AUC's huge true-negative denominator hides
  • B. ROC-AUC — it is always the industry standard
  • C. Accuracy — it summarizes both in one number
  • D. They are interchangeable; 0.64 and 0.18 describe the same thing

Survives elimination: A

Why: Under heavy imbalance ROC-AUC stays optimistic because the false-positive rate is diluted by many true negatives. PR-AUC watches precision on the positives and drops to 0.18, honestly reflecting the minority performance (verified on 105/5000 positives).

89. Drill — ML Foundations

Check

Last one — then we find your gaps.

Check your understanding

On a 2%-positive dataset your model gets ROC-AUC 0.64 but PR-AUC 0.18. Which metric better reflects minority-class performance, and why?

  • A. PR-AUC — its precision axis exposes false positives among the rare positives, which ROC-AUC's huge true-negative denominator hides (correct)
  • B. ROC-AUC — it is always the industry standard
  • C. Accuracy — it summarizes both in one number
  • D. They are interchangeable; 0.64 and 0.18 describe the same thing

Answer: A

Why: Under heavy imbalance ROC-AUC stays optimistic because the false-positive rate is diluted by many true negatives. PR-AUC watches precision on the positives and drops to 0.18, honestly reflecting the minority performance (verified on 105/5000 positives).

Why B tempts people
ROC-AUC being common does not make it appropriate here — it is precisely the metric that misleads under imbalance.
Why C tempts people
Accuracy scores 98% by predicting all-negative and finding zero positives — it hides the failure entirely.
Why D tempts people
They diverge sharply here (0.64 vs 0.18); that gap is the whole point — they measure different things under imbalance.

90. Find the gaps & build it

Section

Part 5 of 5

91. The gap-identification protocol

Concept

Turn this session into a plan: for every topic, rate your confidence 1–5 and log anything that needed a hint.

confidenceaction
5 — cold, fastretire from active rotation
3–4 — slow / shakyre-quiz in 3 days, then 1 week
1–2 — needed a hintre-derive now, re-quiz next session

The prioritized 1–2 list IS your mock-exam study plan. A concept is 'done' only when you reproduce it cold on a later day.

92. Your turn: reproduce OLS-four-ways

Concept

The capstone build: fit the five-student line and prove three independent routes agree. You derived every piece — now assemble it yourself, typing each line and predicting each output before you run it.

#milestoneroute
1Design matrix + normal equationsnp.linalg.solve
2Pseudoinverse routenp.linalg.pinv
3Score R² and match sklearnLinearRegression

Build rules: type every line yourself, run after each line, and when something errors, read the array shapes first — don't delete the error.

93. Milestone 1 — design matrix + normal equations

Worked example

Your turn: build X with a bias column and solve the normal equations. Say the shape (5, 2) and the answer [2.2, 0.6] out loud before printing.

Hint: np.c_[np.ones_like(x), x] stacks the bias column, then np.linalg.solve(X.T @ X, X.T @ y). Never use inv.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
XtX = X.T @ X
Xty = X.T @ y
w = np.linalg.solve(XtX, Xty)
print(X.shape)   # (5, 2)
print(XtX)       # [[5 15] [15 55]]
print(w)         # [2.2 0.6]
checkvalue (verified)
X.shape(5, 2)
XᵀX[[5, 15], [15, 55]]
Xᵀy[20, 66]
w[2.2, 0.6]

94. Milestone 2 — the pseudoinverse route

Worked example

Your turn: get the same w a second, independent way — the SVD-based pseudoinverse. Predict it will match milestone 1 before printing.

Hint: np.linalg.pinv(X) @ y computes X⁺y directly via the SVD, no XᵀX formed.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w_normal = np.linalg.solve(X.T @ X, X.T @ y)
w_pinv   = np.linalg.pinv(X) @ y
print(w_normal.round(4))   # [2.2 0.6]
print(w_pinv.round(4))     # [2.2 0.6]
print(np.allclose(w_normal, w_pinv))  # True
routew (verified)
normal equations[2.2, 0.6]
pseudoinverse X⁺y[2.2, 0.6]
allcloseTrue

95. Restore the missing line: Milestone 3 — score R² and match sklearn

Fill the middle

Fill in the blanks

From Milestone 3 — score R² and match sklearn — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
yhat = X @ w
r2 = 1 - np.sum((y - yhat)2) / np.sum((y - y.mean())2)
print('w =', w.round(4), ' R2 =', round(r2, 4)) # [2.2 0.6] 0.6
m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))

Why: X is what everything below it consumes, so the wrong expression here fails later and somewhere else. Your from-scratch line explains 60% of the score variance, and sklearn's fitted intercept and slope match your hand-built coefficients exactly.

96. Milestone 3 — score R² and match sklearn

Worked example

Your turn: compute R² from scratch, then fit sklearn and confirm intercept, slope, and score all agree. Predict whether they match exactly.

Hint: ŷ = X @ w; R² = 1 − Σ(y−ŷ)² / Σ(y−ȳ)². Fit sklearn on x.reshape(-1, 1) and compare intercept_, coef_.

import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
w = np.linalg.solve(X.T @ X, X.T @ y)
yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print('w =', w.round(4), ' R2 =', round(r2, 4))   # [2.2 0.6] 0.6
m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))

R² = 0.6, and sklearn prints 2.2 0.6 — everything agrees

Why: Your from-scratch line explains 60% of the score variance, and sklearn's fitted intercept and slope match your hand-built coefficients exactly. You reproduced linear regression from the math up.

sourceinterceptslopeR²
your OLS2.20.60.6
sklearn2.20.60.6

97. Work backwards from the answer: Milestone 3 — score R² and match sklearn

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

R² = 0.6, and sklearn prints 2.2 0.6 — everything agrees

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Your turn: compute R² from scratch, then fit sklearn and confirm intercept, slope, and score all agree. Predict whether they match exactly.

98. When two features are secretly the same

Intuition

Suppose one feature is a copy of another — income and income in cents. The second column points the same direction as the first; it adds no new axis to the column space.

Now infinitely many weight vectors give the identical prediction — shift weight from one twin to the other and nothing changes. There is no unique w, so XᵀX is singular and solve can't pick an answer.

Ridge breaks the tie: by charging for weight size, it prefers the smallest ‖w‖ among all the equivalent fits — a unique answer, always. That's the fallback we verify next.

99. Ridge — the always-invertible fallback

Concept

When a feature is collinear, XᵀX has a zero eigenvalue and solve fails. Ridge adds λI, lifting every eigenvalue by λ so the system is always solvable.

\[ (X^\top X + \lambda I)\, w = X^\top y \]

Add a redundant column 2x to our data: XᵀX becomes rank-2 of 3 with eigenvalues {0, 0.90, 279.1}. That 0 is what kills solve. Ridge with λ = 1 shifts it to 1.

100. Teach it back: Ridge — the always-invertible fallback

Explain it

Discussion prompt

Explain Ridge — the always-invertible fallback to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

When a feature is collinear, XᵀX has a zero eigenvalue and solve fails. Ridge adds λI, lifting every eigenvalue by λ so the system is always solvable.

101. What has to be given first: Ridge lifts the zero eigenvalue

Missing information

Discussion prompt

Same collinear Xc that makes solve raise. Show the singular eigenvalue, add λI, and watch it become solvable — runnable as written.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

rank 2 of 3 gives a zero eigenvalue, so plain solve raises LinAlgError. Adding λI = np.eye(3) lifts every eigenvalue by 1, making the matrix non-singular. Same solve now returns [1.0734, 0.1808, 0.3616].

102. Ridge lifts the zero eigenvalue

Worked example

Same collinear Xc that makes solve raise. Show the singular eigenvalue, add λI, and watch it become solvable — runnable as written.

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x]        # col 3 = 2 * col 2
G = Xc.T @ Xc
print(np.linalg.matrix_rank(Xc))      # 2  (not 3)
print(np.linalg.eigvalsh(G).round(4)) # [-0. 0.8957 279.1043]
wr = np.linalg.solve(G + 1.0*np.eye(3), Xc.T @ y)
print(np.linalg.eigvalsh(G + np.eye(3)).round(4))  # [1. 1.8957 280.1043]
print(wr.round(4))                    # [1.0734 0.1808 0.3616]

The 0 eigenvalue becomes 1 after adding λI, and solve returns weights

Why: rank 2 of 3 gives a zero eigenvalue, so plain solve raises LinAlgError. Adding λI = np.eye(3) lifts every eigenvalue by 1, making the matrix non-singular. Same solve now returns [1.0734, 0.1808, 0.3616].

matrixeigenvalues (verified)solvable?
XcᵀXc (collinear)[0, 0.8957, 279.1]no — singular
XcᵀXc + I (ridge)[1, 1.8957, 280.1]yes
ridge w[1.0734, 0.1808, 0.3616]returned

103. What each one costs: Ridge lifts the zero eigenvalue

Trade off

Comparison matrix

From Ridge lifts the zero eigenvalue: every row here is a choice with a cost. Fill the eigenvalues (verified) column, then say which row you would actually pick and what you give up for it.

matrixeigenvalues (verified)solvable?
XcᵀXc (collinear)[0, 0.8957, 279.1]no — singular
XcᵀXc + I (ridge)[1, 1.8957, 280.1]yes
ridge w[1.0734, 0.1808, 0.3616]returned

104. Guess the shape of the answer: Ridge is MAP with a Gaussian prior

Estimation

Predict first

One more Phase-1 bridge: adding λ‖w‖² is the same as putting a Gaussian prior w ~ N(0, τ²I) on the weights and doing MAP (L12). The penalty is the log-prior. Ridge shrinks ‖w‖; verify it.

Commit before you compute: what does Ridge is MAP with a Gaussian prior come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: λ = 1 shrinks w from [2.2, 0.6] to [1.1712, 0.8649], and ‖w‖ from 2.2804 to 1.4559

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The Gaussian prior pulls weights toward 0, trading a little training fit for stability — the MAP estimate under a zero-mean weight prior.

105. Ridge is MAP with a Gaussian prior

Worked example

One more Phase-1 bridge: adding λ‖w‖² is the same as putting a Gaussian prior w ~ N(0, τ²I) on the weights and doing MAP (L12). The penalty is the log-prior. Ridge shrinks ‖w‖; verify it.

\[ \underbrace{\lVert Xw - y\rVert^2}_{\text{Gaussian likelihood}} + \underbrace{\lambda \lVert w\rVert^2}_{\text{Gaussian log-prior}} \;\Rightarrow\; (X^\top X + \lambda I)w = X^\top y \]

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w_ols   = np.linalg.solve(X.T @ X, X.T @ y)
w_ridge = np.linalg.solve(X.T @ X + 1.0*np.eye(2), X.T @ y)
print(w_ols.round(4))                        # [2.2 0.6]
print(w_ridge.round(4))                       # [1.1712 0.8649]
print(round(np.linalg.norm(w_ols), 4),
      round(np.linalg.norm(w_ridge), 4))       # 2.2804 1.4559

λ = 1 shrinks w from [2.2, 0.6] to [1.1712, 0.8649], and ‖w‖ from 2.2804 to 1.4559

Why: The Gaussian prior pulls weights toward 0, trading a little training fit for stability — the MAP estimate under a zero-mean weight prior. As λ → 0 ridge recovers OLS; as λ → ∞ it forces w → 0.

λw = [w₀, w₁] (verified)‖w‖
0 (OLS / MLE)[2.2, 0.6]2.2804
1 (ridge / MAP)[1.1712, 0.8649]1.4559

106. Fill in: w = [w₀, w₁] (verified) for Ridge is MAP with a Gaussian prior

Comparison

Comparison matrix

From Ridge is MAP with a Gaussian prior: refill the w = [w₀, w₁] (verified) column from what you know. The rest of the table is as it appeared.

λw = [w₀, w₁] (verified)‖w‖
0 (OLS / MLE)[2.2, 0.6]2.2804
1 (ridge / MAP)[1.1712, 0.8649]1.4559

107. Mock-exam readiness

Concept

The mock tests theory AND code at pace. Before it: a clean environment ready, the 1–2 confidence list reviewed, and the retrieval log current.

On the exam: attempt cold first, show every derivation step, and watch the three traps (p-value conditioning, det>0 ⇏ PD, MSE-for-classification). You built the whole foundation — now prove it sticks.

108. By analogy: Mock-exam readiness

Analogy

Discussion prompt

Explain Mock-exam readiness by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The mock tests theory AND code at pace. Before it: a clean environment ready, the 1–2 confidence list reviewed, and the retrieval log current.

109. Show it off

Concept

Slides closed, out loud: explain (1) why the four OLS routes must agree, (2) why softmax+CE gives p − y, and (3) what λI does to the eigenvalues of XᵀX.

Stretch: run lr = 0.02 gradient descent and watch it diverge, then drop to 0.001 and watch it converge to [2.2, 0.6] — the exact stability boundary the exam likes to probe.

110. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why the four OLS routes must agree, (2) why softmax+CE gives p − y, and (3) what λI does to the eigenvalues of XᵀX.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Stretch: run lr = 0.02 gradient descent and watch it diverge, then drop to 0.001 and watch it converge to [2.2, 0.6] — the exact stability boundary the exam likes to probe.

111. Connect it up: Lesson 37: Phase 1 Review & Mock Prep

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — One dataset, four routes to OLS · The four pillars, re-derived · Traps & synthesis · Retrieval drills · Find the gaps & build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

112. Phase 1, consolidated

Recap

pillarthe one keystone (verified)
synthesisfour routes → one w = [2.2, 0.6]
linear algebraSVD: A = UΣVᵀ, σ = √eig(AᵀA)
probabilitylosses are negative log-likelihoods; CE = H + KL
optimizationconvex ⇒ global min; lr < 2/L for stability
ML foundationssoftmax+CE ⇒ ∂L/∂z = p − y; imbalance ⇒ PR-AUC

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 37 (Week 13 — Phase 1 Review & Mock Exam) — Barron · USAAIO Round 2 Preparation, 2026
  2. Revised Edition retention protocol (cold attempt → teach → reconstruct; log gaps) — USAAIO Revised Lesson Plan, June 2026
  3. Gilbert Strang, Introduction to Linear Algebra, Ch. 4 & 7 (Projections, SVD) — Wellesley-Cambridge Press, 5th ed.
  4. Every coefficient, eigenvalue, singular value, gradient, and AUC produced by real execution — numpy 2.2.6 + scikit-learn 1.9.0, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108