Lesson 39: Mock Exam — Phase 1 Coding

USAAIO Lesson 39, from Week 13, fully worked: a timed from-scratch coding mock exam over the four Phase 1 algorithms. PCA is done by both the covariance and the SVD route, derived and traced on a four-point toy and then on digits; gradient descent on (w−5)² is traced step by step with the learning-rate stability bound; Gaussian MLE is derived from the log-likelihood, with the 1/n against 1/(n−1) split; and 5-fold cross-validation is built index by index. Each comes with a theory sub-question, every snippet is self-contained and verified against sklearn, and there is a code-review protocol and a your-turn rebuild. Every number was produced by real execution. The lesson runs to 64 slides.

Subject: Machine Learning · 111 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Mock Exam: Phase 1 Coding

Title

USAAIO · Lesson 39 · Week 13

Four algorithms, from scratch, on the clock — PCA, gradient descent, MLE, cross-validation. We derive each one, trace it line by line on a tiny hand-checkable dataset, then prove our NumPy matches sklearn to the digit.

2. By the end of this exam you can

Objectives

  1. Implement PCA two ways — covariance eigendecomposition and SVD — and report the explained-variance ratio, matching sklearn exactly
  2. Run gradient descent on (w−5)² by hand, trace every step, and state the learning-rate bound that separates convergence from divergence
  3. Derive the Gaussian MLE from the log-likelihood and know cold why it divides by n, not n−1
  4. Build k-fold cross-validation index by index so each point is held out exactly once, and match cross_val_score
  5. Answer the theory sub-question attached to each problem, then pass a code review on centering, ddof, shapes, and vectorization

3. What survived from Mock Exam — Phase 1 Theory?

Warm-up

Discussion prompt

Before we open Lesson 39: Mock Exam — Phase 1 Coding: without looking back, what was the main idea of Mock Exam — Phase 1 Theory, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a timed theory mock exam mirroring the USAAIO format — exam-pace retrieval questions across linear algebra, probability, calculus/optimization, information theory, metrics, SVM/kernels, and generalization, plus a verified answer key and an error-analysis protocol.

4. How the coding exam works

Section

Part 1 of 6

5. The format

Concept

The USAAIO coding section is timed and runs in Colab. You implement each algorithm from scratch in NumPy — the library is allowed only to verify your answer, never to do the algorithm for you.

6. Break it if you can: The format

Counterexample

Discussion prompt

The USAAIO coding section is timed and runs in Colab. You implement each algorithm from scratch in NumPy — the library is allowed only to verify your answer, never to do the algorithm for you.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Why 'from scratch' is the whole point

Intuition

Calling PCA().fit(X) proves you can read docs. Writing the SVD, the centering, and the explained-variance ratio yourself proves you understand what the library is doing — which is the only thing the exam can actually test.

So the graders pick problems where a one-line library call and a from-scratch version can disagree unless you know the internals: centering in PCA, the ddof in variance, the fold boundaries in CV. Those disagreements are exactly where points are won and lost.

8. Rebuild the recipe: The per-problem protocol

Ranking

Put in order

These are the steps of The per-problem protocol, scrambled. Put them back in order before the next slide shows you.

  1. Answer the theory sub-question first — it tells you the formula the code must implement
  2. Predict the output before you run (a wrong prediction catches a bug early)
  3. Implement from scratch, vectorized where NumPy allows
  4. Verify against a known value or the library
  5. Review: shapes, edge cases (centering, ddof, fold split), readability

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

9. The per-problem protocol

Pattern

  1. Answer the theory sub-question first — it tells you the formula the code must implement
  2. Predict the output before you run (a wrong prediction catches a bug early)
  3. Implement from scratch, vectorized where NumPy allows
  4. Verify against a known value or the library
  5. Review: shapes, edge cases (centering, ddof, fold split), readability

10. One running dataset for the derivations

Concept

Each library check uses a real dataset (digits, diabetes), but every derivation first runs on a tiny hand-checkable set so you can verify by eye before you trust the code. For PCA that set is four points on the axes:

pointx₁x₂
120
202
3−20
40−2

Four points, already centered on the origin, spread equally along both axes. Keep this picture — we will predict every PCA number from it by hand, then confirm in NumPy.

11. Fill in: x₁ for One running dataset for the derivations

Comparison

Comparison matrix

From One running dataset for the derivations: refill the x₁ column from what you know. The rest of the table is as it appeared.

pointx₁x₂
120
202
3−20
40−2

12. Problem 1 — PCA

Section

Part 2 of 6 — from scratch

13. What PCA is asking

Concept

PCA finds the directions along which the data varies most. The first principal component is the unit vector v that maximizes the variance of the projected data; the second is the next such direction perpendicular to it, and so on.

\[ v_1 = \arg\max_{\lVert v \rVert = 1} \; \operatorname{Var}(Xv) \]

The theory sub-question for this problem: why must you center the data first? Hold that thought — the derivation answers it in two slides.

14. By analogy: What PCA is asking

Analogy

Discussion prompt

Explain What PCA is asking by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The theory sub-question for this problem: why must you center the data first? Hold that thought — the derivation answers it in two slides.

15. Variance lives in the covariance matrix

Concept

For centered data Xc (each column has mean 0), the sample covariance is C = XcᵀXc / (n−1). Entry (j,k) is the covariance of feature j with feature k.

\[ C = \frac{1}{n-1} X_c^\top X_c \quad\in\; \mathbb{R}^{d\times d} \]

The variance of the projection Xc v is exactly vᵀ C v. So maximizing projected variance means maximizing vᵀ C v over unit vectors — and that maximum is the top eigenvector of C.

16. Teach it back: Variance lives in the covariance matrix

Explain it

Discussion prompt

Explain Variance lives in the covariance matrix to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For centered data Xc (each column has mean 0), the sample covariance is C = XcᵀXc / (n−1). Entry (j,k) is the covariance of feature j with feature k.

17. What has to happen first: Build C by hand on the 4-point toy

Ranking

Put in order

Put the moves of Build C by hand on the 4-point toy into the order they have to happen.

  1. (XᵀX)₁₁ = Σ x₁² = 2² + 0² + (−2)² + 0² = 8
  2. (XᵀX)₂₂ = Σ x₂² = 0² + 2² + 0² + (−2)² = 8
  3. (XᵀX)₁₂ = Σ x₁x₂ = 2·0 + 0·2 + (−2)·0 + 0·(−2) = 0

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Column 1 dotted with itself: the squared first coordinates sum to 8.

18. Build C by hand on the 4-point toy

Worked example

The four points are already centered (mean [0,0]), so Xc = X. Compute XcᵀXc entry by entry, then divide by n−1 = 3:

(XᵀX)₁₁ = Σ x₁² = 2² + 0² + (−2)² + 0² = 8

Why: Column 1 dotted with itself: the squared first coordinates sum to 8.

(XᵀX)₂₂ = Σ x₂² = 0² + 2² + 0² + (−2)² = 8

Why: Column 2 dotted with itself. Symmetric spread means the two diagonal entries match.

(XᵀX)₁₂ = Σ x₁x₂ = 2·0 + 0·2 + (−2)·0 + 0·(−2) = 0

Why: The features never move together on this set, so the off-diagonal covariance is 0.

\[ C = \frac{1}{3}\begin{bmatrix} 8 & 0 \\ 0 & 8 \end{bmatrix} = \begin{bmatrix} 2.667 & 0 \\ 0 & 2.667 \end{bmatrix} \]

import numpy as np
X = np.array([[2., 0.], [0., 2.], [-2., 0.], [0., -2.]])
Xc = X - X.mean(0)                 # center: variance is around the mean
C = (Xc.T @ Xc) / (len(X) - 1)     # sample covariance, ddof=1
print(C)
entryby handprinted C
C[0,0]8/32.66666667
C[1,1]8/32.66666667
C[0,1] = C[1,0]0/30.0

19. What each one costs: Build C by hand on the 4-point toy

Trade off

Comparison matrix

From Build C by hand on the 4-point toy: every row here is a choice with a cost. Fill the by hand column, then say which row you would actually pick and what you give up for it.

entryby handprinted C
C[0,0]8/32.66666667
C[1,1]8/32.66666667
C[0,1] = C[1,0]0/30.0

20. Guess the shape of the answer: Eigendecompose C

Estimation

Predict first

C is a scalar multiple of the identity — 2.667·I. Every vector is an eigenvector, and every eigenvalue equals 2.667. So the variance is the same in every direction: a symmetric cross of points has no single 'most-spread' axis.

Commit before you compute: what does Eigendecompose C come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Eigenvalue = variance captured by that component

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each eigenvalue λ is the variance of the data projected onto its eigenvector.

21. Eigendecompose C

Worked example

C is a scalar multiple of the identity — 2.667·I. Every vector is an eigenvector, and every eigenvalue equals 2.667. So the variance is the same in every direction: a symmetric cross of points has no single 'most-spread' axis.

\[ C\,v = 2.667\,v \quad\text{for every } v \;\Longrightarrow\; \lambda_1 = \lambda_2 = 2.667 \]

Eigenvalue = variance captured by that component

Why: Each eigenvalue λ is the variance of the data projected onto its eigenvector. Here both are 2.667, so the two components explain equal shares — a clean 50/50.

objecthand valuenp.linalg.eigh(C)
λ₁2.6672.66666667
λ₂2.6672.66666667
eigenvectorsany orthonormal pair[[1,0],[0,1]]

22. Work backwards from the answer: Eigendecompose C

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Eigenvalue = variance captured by that component

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

C is a scalar multiple of the identity — 2.667·I. Every vector is an eigenvector, and every eigenvalue equals 2.667. So the variance is the same in every direction: a symmetric cross of points has no single 'most-spread' axis.

23. The SVD shortcut

Concept

You rarely form C explicitly. The SVD of the centered data, Xc = U S Vᵀ, hands you the principal components directly: the rows of Vᵀ are the components, and each singular value relates to an eigenvalue of C by λ = S²/(n−1).

\[ X_c = U S V^\top, \qquad \lambda_i = \frac{S_i^2}{n-1}, \qquad \text{PC}_i = V^\top_{[i]} \]

SVD is preferred because it never forms XcᵀXc — so it never squares the condition number, exactly the numerical-stability point from Lesson 7. Same components, safer arithmetic.

24. What has to be given first: SVD on the toy — singular values to EVR

Missing information

Discussion prompt

Run the SVD and convert. Predict first: λ = 2.667 and λ = S²/(n−1), so S² = 3·2.667 = 8, meaning S = √8 ≈ 2.828. And equal eigenvalues ⇒ equal EVR ⇒ [0.5, 0.5].

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The SVD route and the covariance eigendecomposition agree to the last digit — as they must, because they compute the same spectrum. This equality is the standard exam sanity check.

25. SVD on the toy — singular values to EVR

Worked example

Run the SVD and convert. Predict first: λ = 2.667 and λ = S²/(n−1), so S² = 3·2.667 = 8, meaning S = √8 ≈ 2.828. And equal eigenvalues ⇒ equal EVR ⇒ [0.5, 0.5].

import numpy as np
X = np.array([[2., 0.], [0., 2.], [-2., 0.], [0., -2.]])
Xc = X - X.mean(0)
U, S, Vt = np.linalg.svd(Xc, full_matrices=False)
print("S =", S)                         # singular values
print("S**2/(n-1) =", S**2 / (len(X) - 1))   # = eigenvalues of C
print("EVR =", (S**2 / np.sum(S**2)))   # explained-variance ratio
quantitypredictedprinted
S[2.828, 2.828][2.82842712, 2.82842712]
S²/(n−1)[2.667, 2.667][2.66666667, 2.66666667]
EVR[0.5, 0.5][0.5, 0.5]

S²/(n−1) reproduced the eigenvalues exactly

Why: The SVD route and the covariance eigendecomposition agree to the last digit — as they must, because they compute the same spectrum. This equality is the standard exam sanity check.

26. Predict the next row: Projecting the points onto PC1

Pattern

Predict first

The table runs: 1 | 2 | −2 | −2.0 · 2 | 0 | 0 | 0.0 · 3 | −2 | 2 | 2.0

In Projecting the points onto PC1, given the rows so far: what is the next one — the row where point is 4?

Correct: 4 | 0 | 0 | 0.0

pointx₁score = −x₁printed z
12−2−2.0
2000.0
3−222.0
4000.0

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two points on the x₁-axis get scores ±2; the two on the x₂-axis project to 0.

27. Projecting the points onto PC1

Worked example

The principal scores are the projections Xc · PC1. NumPy returns PC1 = [−1, 0] (a sign choice; [1,0] is equally valid). Projecting each point onto that axis just reads off (the negative of) its first coordinate:

import numpy as np
X = np.array([[2., 0.], [0., 2.], [-2., 0.], [0., -2.]])
Xc = X - X.mean(0)
U, S, Vt = np.linalg.svd(Xc, full_matrices=False)
pc1 = Vt[0]                 # first principal direction
z = Xc @ pc1               # project every point onto PC1
print("PC1 =", pc1)
print("scores =", z)
pointx₁score = −x₁printed z
12−2−2.0
2000.0
3−222.0
4000.0

PC1 = [−1, 0], scores = [−2, 0, 2, 0]

Why: The two points on the x₁-axis get scores ±2; the two on the x₂-axis project to 0. PCA compressed 2-D points to a single number while keeping the biggest spread — the whole purpose of the method.

28. Watch it run: Projecting the points onto PC1

Pattern

Step through it

Step through Projecting the points onto PC1 one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: point is 1
  2. Step 2: point is 2
  3. Step 3: point is 3
  4. Step 4: point is 4

29. Scaling up: the real exam dataset

Concept

The graders run PCA on sklearn's digits — 1797 handwritten digits, each a 64-pixel vector. The task: report the top-3 EVR and how many components reach 95% of the variance.

The pipeline is identical to the toy — center, SVD, S²/ΣS² — just on a (1797, 64) matrix instead of (4, 2). Predict: 64 pixels are highly correlated, so far fewer than 64 components should cover 95%.

30. Restore the missing line: PCA on digits — full pipeline

Fill the middle

Fill in the blanks

From PCA on digits — full pipeline — one line has had its right-hand side removed. Put it back.

import numpy as np
from sklearn.datasets import load_digits
X = load_digits().data # (1797, 64)
Xc = X - X.mean(0) # center
U, S, Vt = np.linalg.svd(Xc, full_matrices=False)
evr = S2 / np.sum(S2)
k95 = int(np.searchsorted(np.cumsum(evr), 0.95) + 1)
print("EVR top 3 =", evr[:3].round(4))
print("k for 95% =", k95, "of", X.shape[1])

Why: Xc is what everything below it consumes, so the wrong expression here fails later and somewhere else. More than half the pixel dimensions are redundant — the 64-D images live near a 29-D subspace.

31. PCA on digits — full pipeline

Worked example

Center, SVD, EVR, then find the first cumulative-variance index past 0.95. Self-contained and runnable:

import numpy as np
from sklearn.datasets import load_digits
X = load_digits().data                       # (1797, 64)
Xc = X - X.mean(0)                            # center
U, S, Vt = np.linalg.svd(Xc, full_matrices=False)
evr = S**2 / np.sum(S**2)
k95 = int(np.searchsorted(np.cumsum(evr), 0.95) + 1)
print("EVR top 3 =", evr[:3].round(4))
print("k for 95% =", k95, "of", X.shape[1])
quantityvalue (verified)
EVR component 10.1489
EVR component 20.1362
EVR component 30.1179
k for 95% variance29 of 64

29 of 64 components hold 95% of the variance

Why: More than half the pixel dimensions are redundant — the 64-D images live near a 29-D subspace. That compression is the entire value proposition of PCA.

32. Finish it with less help: Verify against sklearn

Faded example

Fill in the blanks

Verify against sklearn, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
X = load_digits().data
# from scratch
Xc = X - X.mean(0)
S = np.linalg.svd(Xc, full_matrices=False)[1]
evr_scratch = (S2 / np.sum(S2))[:3].round(4)
# library
evr_sklearn = PCA().fit(X).explained_variance_ratio_[:3].round(4)
print("scratch:", evr_scratch)
print("sklearn:", evr_sklearn)
print("match:", np.allclose(evr_scratch, evr_sklearn))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. sklearn subtracts the column means internally before its SVD.

33. Verify against sklearn

Worked example

The from-scratch pipeline must match sklearn.decomposition.PCA to the digit — that is the graded correctness check. Compare the top-3 EVR both ways:

import numpy as np
from sklearn.datasets import load_digits
from sklearn.decomposition import PCA
X = load_digits().data
# from scratch
Xc = X - X.mean(0)
S = np.linalg.svd(Xc, full_matrices=False)[1]
evr_scratch = (S**2 / np.sum(S**2))[:3].round(4)
# library
evr_sklearn = PCA().fit(X).explained_variance_ratio_[:3].round(4)
print("scratch:", evr_scratch)
print("sklearn:", evr_sklearn)
print("match:", np.allclose(evr_scratch, evr_sklearn))
sourceEVR top 3
from scratch (SVD)[0.1489, 0.1362, 0.1179]
sklearn PCA[0.1489, 0.1362, 0.1179]
np.allcloseTrue

Exact match — because both center first

Why: sklearn subtracts the column means internally before its SVD. Reproduce that one step and the results are identical; skip it and they diverge. That single line is the theory sub-question in disguise.

34. Why centering matters — measured

Worked example

Shift the toy points by [+10, +10] so their mean is [10, 10]. Run SVD without centering (wrong) and with centering (right) and watch PC1 flip:

import numpy as np
X = np.array([[2., 0.], [0., 2.], [-2., 0.], [0., -2.]]) + 10.0  # shifted
# WRONG: skip centering
Sw = np.linalg.svd(X, full_matrices=False)[1]
Vw = np.linalg.svd(X, full_matrices=False)[2]
print("uncentered S   =", Sw.round(4), " PC1 =", Vw[0].round(4))
# RIGHT: center first
Xc = X - X.mean(0)
Sr = np.linalg.svd(Xc, full_matrices=False)[1]
Vr = np.linalg.svd(Xc, full_matrices=False)[2]
print("centered   S   =", Sr.round(4), " PC1 =", Vr[0].round(4))
variantS₁PC1
uncentered (wrong)28.43[−0.707, −0.707]
centered (right)2.83[−1.0, 0.0]
true spread axis—along an axis, not the diagonal

Uncentered PC1 points at the mean, not the spread

Why: Without centering, the dominant singular direction chases the offset [10,10] (the diagonal) and its singular value balloons to 28.43. Centering removes the offset so PC1 finds the real axis of variation. This is the classic PCA fail.

35. Something is wrong here: PCA without centering

Anomaly

Predict first

A student writes this, and it looks reasonable:

PCA is 'just the SVD of the data', so decompose X directly — centering is an extra line I can skip.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: On the shifted toy this returns S₁ = 28.43 and PC1 = [−0.707, −0.707].

Center first — variance is defined as spread around the mean.

Why: On the shifted toy this returns S₁ = 28.43 and PC1 = [−0.707, −0.707]. The top component points toward the data's MEAN (the origin offset), not its spread. Every EVR is wrong — an instant rubric fail.

36. Trap: PCA without centering

Trap

The trap

PCA is 'just the SVD of the data', so decompose X directly — centering is an extra line I can skip.

U, S, Vt = svd(X); PC1 = Vt[0]

Why: On the shifted toy this returns S₁ = 28.43 and PC1 = [−0.707, −0.707]. The top component points toward the data's MEAN (the origin offset), not its spread. Every EVR is wrong — an instant rubric fail.

The fix

Center first — variance is defined as spread around the mean.

Xc = X − X.mean(0); U, S, Vt = svd(Xc)

Why: Now S₁ = 2.83 and PC1 = [−1, 0], the true axis of variation, and the EVR matches sklearn. The reviewer checks exactly this line — it is the single most common PCA point-loser.

37. Problem 2 — Gradient descent

Section

Part 3 of 6

38. The problem

Concept

Minimize f(w) = (w − 5)² starting from w = 0, using plain gradient descent. The theory sub-question: write the update rule and the gradient, and say what value w converges to.

\[ f(w) = (w-5)^2 \qquad\Longrightarrow\qquad \min_w f(w) \text{ at } w = 5 \]

The minimum is obvious by eye — a parabola with vertex at w = 5. That is exactly why it's a good exam problem: you know the answer, so you can prove your iteration actually reaches it.

39. Plan first: The gradient, then the update rule

Step zero

Discussion prompt

The gradient, then the update rule — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Differentiate f(w) = (w − 5)²

Answer:

  1. Differentiate f(w) = (w − 5)²
  2. Step opposite the gradient
  3. At w < 5 the gradient is negative, so we step right

40. The gradient, then the update rule

Worked example

Differentiate f(w) = (w − 5)²

Why: Chain rule: d/dw (w−5)² = 2(w−5)·1. The gradient is zero exactly at w = 5, the vertex.

\[ f'(w) = 2(w - 5) \]

Step opposite the gradient

Why: Gradient descent moves w a small distance −η·f'(w): downhill, because the gradient points uphill. η is the learning rate.

\[ w \leftarrow w - \eta\, f'(w) = w - \eta\cdot 2(w-5) \]

At w < 5 the gradient is negative, so we step right

Why: e.g. at w = 0, f'(0) = 2(0−5) = −10; the update w − η(−10) = w + 10η increases w, marching it toward 5. The sign logic is the heart of the method.

wf'(w) = 2(w−5)direction of step
0−10increase w (→ 5)
50stay (at the minimum)
8+6decrease w (→ 5)

41. Decode the notation: The gradient, then the update rule

Notation

Annotate

From The gradient, then the update rule — read this one piece at a time. What is each part doing?

On: \( w \leftarrow w - \eta\, f'(w) = w - \eta\cdot 2(w-5) \)

  • Chain rule: d/dw (w−5)² = 2(w−5)·1. The gradient is zero exactly at w = 5, the vertex.
  • Gradient descent moves w a small distance −η·f'(w): downhill, because the gradient points uphill. η is the learning rate.
  • e.g. at w = 0, f'(0) = 2(0−5) = −10; the update w − η(−10) = w + 10η increases w, marching it toward 5. The sign logic is the heart of the method.

42. Trace the first six steps by hand

Worked example

With η = 0.1 the update is w ← w − 0.1·2(w−5) = 0.8w + 1. Apply it from w = 0 and record w, f(w), and the gradient at each step:

import numpy as np
grad = lambda w: 2 * (w - 5)     # gradient of f(w) = (w - 5)**2
w, lr = 0.0, 0.1
for t in range(6):
    print(f"t={t}  w={w:.4f}  f={(w-5)**2:.4f}  grad={grad(w):+.4f}")
    w -= lr * grad(w)            # descent step
twf(w)grad = 2(w−5)
00.000025.0000−10.0000
11.000016.0000−8.0000
21.800010.2400−6.4000
32.44006.5536−5.1200
42.95204.1943−4.0960
53.36162.6844−3.2768

Each step multiplies the distance-to-5 by 0.8

Why: From w=0: 0.8·0 + 1 = 1, then 0.8·1 + 1 = 1.8, then 2.44… The gap 5 − w shrinks by a factor 0.8 every step (5 − wₜ = 5·0.8ᵗ), so f falls geometrically toward 0. Every printed value matches the by-hand recurrence.

43. Watch it run: Trace the first six steps by hand

Pattern

Step through it

Step through Trace the first six steps by hand one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: t is 0
  2. Step 2: t is 1
  3. Step 3: t is 2
  4. Step 4: t is 3
  5. Step 5: t is 4
  6. Step 6: t is 5

44. Guess the shape of the answer: Run it to convergence

Estimation

Predict first

Sixty steps drive the 0.8ᵗ gap to essentially zero. Predict w → 5.0 before running:

Commit before you compute: what does Run it to convergence come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: w converges to 5.0 — the analytic minimum

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The iteration reaches the vertex we found by calculus, confirming the code implements the math.

45. Run it to convergence

Worked example

Sixty steps drive the 0.8ᵗ gap to essentially zero. Predict w → 5.0 before running:

import numpy as np
grad = lambda w: 2 * (w - 5)
w, lr = 0.0, 0.1
for _ in range(60):
    w -= lr * grad(w)
print(round(w, 4))               # -> 5.0
steps5 − w = 5·0.8ᵗw
100.5374.463
300.00624.994
600.00000775.0000

w converges to 5.0 — the analytic minimum

Why: The iteration reaches the vertex we found by calculus, confirming the code implements the math. round(w, 4) prints 5.0. Answer to the theory sub-question: converges to 5.

46. The learning rate is a tightrope

Intuition

Too small and you crawl — thousands of steps to reach the bottom. Too big and each step overshoots the minimum, landing farther up the opposite wall; the iterates bounce outward and diverge.

For a quadratic with curvature L (here f''(w) = 2, so L = 2), plain gradient descent converges only when η < 2/L = 1. At η = 1 it oscillates forever; above 1 it explodes. The next slide measures exactly that boundary.

47. Restore the missing line: Learning-rate sweep — the stability edge

Fill the middle

Fill in the blanks

From Learning-rate sweep — the stability edge — one line has had its right-hand side removed. Put it back.

import numpy as np
grad = lambda w: 2 * (w - 5)
for lr in [0.1, 0.5, 1.0, 1.1]:
w = 0.0
for _ in range(50):
w -= lr * grad(w)
tag = f"___" if abs(w) < 1e6 else f"diverged (___)"
print(f"lr=___ -> w=___")

Why: w is what everything below it consumes, so the wrong expression here fails later and somewhere else. At η = 1 the multiplier is |1 − 2η| = 1, so w bounces 0 → 10 → 0 forever and lands back at 0.

48. Learning-rate sweep — the stability edge

Worked example

Run 50 steps at four learning rates and watch the 2/L = 1 threshold act as a hard wall:

import numpy as np
grad = lambda w: 2 * (w - 5)
for lr in [0.1, 0.5, 1.0, 1.1]:
    w = 0.0
    for _ in range(50):
        w -= lr * grad(w)
    tag = f"{w:.4f}" if abs(w) < 1e6 else f"diverged ({w:.2e})"
    print(f"lr={lr:<4} -> w={tag}")
ηbehaviorw after 50 steps
0.1converges (slow-ish)4.9999
0.5converges (fast)5.0000
1.0oscillates, never settles0.0000
1.1diverges−45497.19

η = 1.0 is the knife-edge; η = 1.1 explodes

Why: At η = 1 the multiplier is |1 − 2η| = 1, so w bounces 0 → 10 → 0 forever and lands back at 0. At η = 1.1 the multiplier is 1.2 > 1, so the distance grows every step to −45497. This is the 2/L bound made concrete.

49. Something is wrong here: cranking the learning rate for speed

Anomaly

Predict first

A student writes this, and it looks reasonable:

Convergence is slow at η = 0.1, so push it way up to η = 1.1 to finish in fewer steps.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The step multiplier |1 − 2η| = 1.2 exceeds 1, so every update grows the distance to the minimum.

Stay below the stability bound η < 2/L; tune within it for speed.

Why: The step multiplier |1 − 2η| = 1.2 exceeds 1, so every update grows the distance to the minimum. Bigger is NOT faster past the 2/L bound — it's catastrophic. The loss shoots to infinity.

50. Trap: cranking the learning rate for speed

Trap

The trap

Convergence is slow at η = 0.1, so push it way up to η = 1.1 to finish in fewer steps.

η = 1.1 → w = −45497 (diverged)

Why: The step multiplier |1 − 2η| = 1.2 exceeds 1, so every update grows the distance to the minimum. Bigger is NOT faster past the 2/L bound — it's catastrophic. The loss shoots to infinity.

The fix

Stay below the stability bound η < 2/L; tune within it for speed.

η = 0.5 → w = 5.0 in far fewer steps

Why: 0.5 is safely under 1 = 2/L and near the sweet spot, so it converges fast AND stably. The right move is a bigger-but-bounded rate, not an unbounded one.

51. Problem 3 — MLE

Section

Part 4 of 6 — Gaussian

52. The problem

Concept

Given a sample assumed drawn from a Gaussian, find the maximum-likelihood estimates of the mean μ and variance σ². The theory sub-question: what are the closed-form MLEs, and does σ² divide by n or n−1?

The running sample — eight numbers chosen so every quantity is a clean integer:

the 8 data points
values2, 4, 4, 4, 5, 5, 7, 9
n8
sum40

53. Fill in: the 8 data points for The problem

Comparison

Comparison matrix

From The problem: refill the the 8 data points column from what you know. The rest of the table is as it appeared.

the 8 data points
values2, 4, 4, 4, 5, 5, 7, 9
n8
sum40

54. The likelihood we are maximizing

Concept

For independent Gaussian samples, the likelihood is the product of the per-point densities. We maximize its log (a sum, easier to differentiate) — the log-likelihood as a function of μ and σ²:

\[ \ell(\mu, \sigma^2) = \sum_{i=1}^{n} \left[ -\tfrac{1}{2}\log(2\pi\sigma^2) - \frac{(x_i - \mu)^2}{2\sigma^2} \right] \]

Maximizing ℓ means setting its partial derivatives with respect to μ and σ² to zero. Two derivatives, two estimators — derived next, one step per beat.

55. Teach it back: The likelihood we are maximizing

Explain it

Discussion prompt

Explain The likelihood we are maximizing to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For independent Gaussian samples, the likelihood is the product of the per-point densities. We maximize its log (a sum, easier to differentiate) — the log-likelihood as a function of μ and σ²:

56. What has to happen first: Derive μ̂ — differentiate in μ

Ranking

Put in order

Put the moves of Derive μ̂ — differentiate in μ into the order they have to happen.

  1. ∂ℓ/∂μ keeps only the squared-error term
  2. Set to zero and cancel σ²
  3. Solve for μ̂ = the sample mean

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The log(2πσ²) piece has no μ in it, so it drops.

57. Derive μ̂ — differentiate in μ

Worked example

∂ℓ/∂μ keeps only the squared-error term

Why: The log(2πσ²) piece has no μ in it, so it drops. Differentiating −(xᵢ−μ)²/(2σ²) in μ gives +(xᵢ−μ)/σ².

\[ \frac{\partial \ell}{\partial \mu} = \sum_{i=1}^{n} \frac{x_i - \mu}{\sigma^2} \]

Set to zero and cancel σ²

Why: σ² > 0 is a common nonzero factor, so it divides out. What remains is Σ(xᵢ − μ) = 0.

\[ \sum_{i=1}^{n} (x_i - \mu) = 0 \]

Solve for μ̂ = the sample mean

Why: Σxᵢ − nμ = 0 ⇒ μ = (1/n)Σxᵢ. The MLE of the mean is just the average — for our sample 40/8 = 5.

\[ \hat\mu = \frac{1}{n}\sum_{i=1}^{n} x_i = \frac{40}{8} = 5 \]

58. Draw the shape of it: Derive μ̂ — differentiate in μ

Blank canvas

Draw it

Draw what Derive μ̂ — differentiate in μ just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

59. What has to happen first: Derive σ̂² — differentiate in σ²

Ranking

Put in order

Put the moves of Derive σ̂² — differentiate in σ² into the order they have to happen.

  1. ∂ℓ/∂σ² of the two σ²-terms
  2. Set to zero, multiply through by 2σ⁴
  3. Solve for σ̂² — divide by n

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. d/dσ² of −½log(2πσ²) is −1/(2σ²); d/dσ² of −(xᵢ−μ)²/(2σ²) is +(xᵢ−μ)²/(2σ⁴).

60. Derive σ̂² — differentiate in σ²

Worked example

∂ℓ/∂σ² of the two σ²-terms

Why: d/dσ² of −½log(2πσ²) is −1/(2σ²); d/dσ² of −(xᵢ−μ)²/(2σ²) is +(xᵢ−μ)²/(2σ⁴). Sum over i.

\[ \frac{\partial \ell}{\partial \sigma^2} = \sum_{i=1}^{n} \left[ -\frac{1}{2\sigma^2} + \frac{(x_i - \mu)^2}{2\sigma^4} \right] \]

Set to zero, multiply through by 2σ⁴

Why: Clearing denominators turns the equation into −nσ² + Σ(xᵢ−μ)² = 0.

\[ -n\,\sigma^2 + \sum_{i=1}^{n}(x_i - \mu)^2 = 0 \]

Solve for σ̂² — divide by n

Why: σ² = (1/n)Σ(xᵢ − μ)². The MLE divides by n, NOT n−1. It plugs in μ̂, so it is a biased-low estimate — this is the whole point of the sub-question.

\[ \boxed{\;\hat\sigma^2 = \frac{1}{n}\sum_{i=1}^{n}(x_i - \hat\mu)^2\;} \]

61. Say it in words: Derive σ̂² — differentiate in σ²

Translation

\( \frac{\partial \ell}{\partial \sigma^2} = \sum_{i=1}^{n} \left[ -\frac{1}{2\sigma^2} + \frac{(x_i - \mu)^2}{2\sigma^4} \right] \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

62. Compute the MLEs by hand

Worked example

With μ̂ = 5, form each deviation, square it, sum, and divide by n = 8:

Deviations xᵢ − 5 = [−3, −1, −1, −1, 0, 0, 2, 4]

Why: Subtract the mean 5 from each of 2,4,4,4,5,5,7,9.

Squares = [9, 1, 1, 1, 0, 0, 4, 16], sum = 32

Why: Square each deviation and add: 9+1+1+1+0+0+4+16 = 32.

\[ \hat\sigma^2 = \frac{32}{8} = 4, \qquad \hat\sigma = \sqrt{4} = 2 \]

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean()                          # MLE of mean
var = ((d - mu)**2).mean()             # MLE of variance: divide by n
print("mu =", mu)
print("var_MLE =", var)
print("sigma_MLE =", np.sqrt(var))
estimatorby handprinted
μ̂40/8 = 55.0
σ̂² = 32/844.0
σ̂√4 = 22.0

63. Decode the notation: Compute the MLEs by hand

Notation

Annotate

From Compute the MLEs by hand — read this one piece at a time. What is each part doing?

On: \( \hat\sigma^2 = \frac{32}{8} = 4, \qquad \hat\sigma = \sqrt{4} = 2 \)

  • Subtract the mean 5 from each of 2,4,4,4,5,5,7,9.
  • Square each deviation and add: 9+1+1+1+0+0+4+16 = 32.

64. Finish it with less help: n vs n−1 — verify against NumPy

Faded example

Fill in the blanks

n vs n−1 — verify against NumPy, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean()
var_mle = ((d - mu)2).mean()
print("var_MLE (/n) =", var_mle, "== np.var:", np.var(d))
print("var_unbiased =", ((d-mu)
2).sum()/(len(d)-1),
"== np.var ddof=1:", np.var(d, ddof=1))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. They are not rounding variants; they answer different questions.

65. n vs n−1 — verify against NumPy

Worked example

The unbiased sample variance divides by n−1 = 7 instead, giving a different number. np.var defaults to the MLE (ddof=0); ddof=1 gives the unbiased one. Confirm both:

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean()
var_mle = ((d - mu)**2).mean()
print("var_MLE (/n)     =", var_mle,          "== np.var:", np.var(d))
print("var_unbiased     =", ((d-mu)**2).sum()/(len(d)-1),
      "== np.var ddof=1:", np.var(d, ddof=1))
estimatordivisorvalue
σ̂²_MLE (np.var ddof=0)n = 84.0
unbiased (np.var ddof=1)n−1 = 74.5714
difference—0.5714 (never zero for finite n)

MLE = 4.0, unbiased = 4.5714 — two different estimators

Why: They are not rounding variants; they answer different questions. The MLE maximizes likelihood but is biased low because it reuses μ̂; the n−1 version is the unbiased correction. Know which one the prompt wants.

66. Restore the missing line: Confirm σ̂² = 4 is really the maximum

Fill the middle

Fill in the blanks

From Confirm σ̂² = 4 is really the maximum — one line has had its right-hand side removed. Put it back.

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean()
def loglik(v):
return np.sum(-0.5np.log(2np.piv) - (d - mu)2 / (2v))
for v in [3.0, 4.0, 6.0]: # 4.0 is the MLE
print(f"var=___: loglik=___")

Why: mu is what everything below it consumes, so the wrong expression here fails later and somewhere else. Both neighbors (3 and 6) score lower log-likelihood than 4.

67. Confirm σ̂² = 4 is really the maximum

Worked example

The derivation says σ² = 4 maximizes the log-likelihood (with μ = 5 fixed). Evaluate ℓ at σ² = 3, 4, 6 — the peak should sit at 4:

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean()
def loglik(v):
    return np.sum(-0.5*np.log(2*np.pi*v) - (d - mu)**2 / (2*v))
for v in [3.0, 4.0, 6.0]:            # 4.0 is the MLE
    print(f"var={v}: loglik={loglik(v):.4f}")
σ²log-likelihoodvs MLE
3.0−17.0793lower
4.0−16.8967highest ✓
6.0−17.1852lower

ℓ peaks at σ² = 4, exactly the derived MLE

Why: Both neighbors (3 and 6) score lower log-likelihood than 4. The calculus and the numerics agree: our closed form is the maximizer, not a saddle or a min.

68. Watch it run: Confirm σ̂² = 4 is really the maximum

Pattern

Step through it

Step through Confirm σ̂² = 4 is really the maximum one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: σ² is 3.0
  2. Step 2: σ² is 4.0
  3. Step 3: σ² is 6.0

69. Something is wrong here: using n−1 for the MLE

Anomaly

Predict first

A student writes this, and it looks reasonable:

'Sample variance' always divides by n−1, so use np.var(d, ddof=1) for the MLE too.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That is the UNBIASED estimator, a different quantity.

The Gaussian MLE divides by n — that is what maximizing the likelihood gives.

Why: That is the UNBIASED estimator, a different quantity. The prompt asked for the MLE, which the likelihood derivation pins to divisor n. Answering 4.5714 loses the point even though the number is 'a variance'.

70. Trap: using n−1 for the MLE

Trap

The trap

'Sample variance' always divides by n−1, so use np.var(d, ddof=1) for the MLE too.

σ̂² = 32/7 = 4.5714

Why: That is the UNBIASED estimator, a different quantity. The prompt asked for the MLE, which the likelihood derivation pins to divisor n. Answering 4.5714 loses the point even though the number is 'a variance'.

The fix

The Gaussian MLE divides by n — that is what maximizing the likelihood gives.

σ̂² = 32/8 = 4.0 (np.var, ddof=0)

Why: The derivation ended in σ² = (1/n)Σ(xᵢ−μ)². Divisor n is the MLE by construction; n−1 is a separate unbiased correction. Match the estimator to the word 'MLE' in the prompt.

71. Break it on purpose: using n−1 for the MLE

Break the constraint

Discussion prompt

The rule this trap just fixed:

The derivation ended in σ² = (1/n)Σ(xᵢ−μ)². Divisor n is the MLE by construction; n−1 is a separate unbiased correction. Match the estimator to the word 'MLE' in the prompt.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

That is the UNBIASED estimator, a different quantity. The prompt asked for the MLE, which the likelihood derivation pins to divisor n. Answering 4.5714 loses the point even though the number is 'a variance'.

72. Problem 4 — Cross-validation

Section

Part 5 of 6 — k-fold

73. The problem

Concept

Implement 5-fold cross-validation from scratch and match sklearn's cross_val_score. The theory sub-question: across the whole procedure, how many times is each data point used for validation?

The idea: split the data into k equal blocks (folds). Train on k−1 of them, score on the held-out one, rotate which fold is held out, and average the k scores. This estimates how the model does on unseen data.

74. Each point validates exactly once

Intuition

Every fold is held out once across the k rounds, and the folds partition the data — no point is in two folds. So each point is a validation point in exactly one round and a training point in the other k−1.

That is the answer to the sub-question: exactly once. It's what makes k-fold efficient — every point contributes to both training and testing, without ever being tested on data it trained on.

75. Predict the next row: The fold split, made visible

Pattern

Predict first

The table runs: 0 | [0, 1] | 2,3,4,5,6,7,8,9 · 1 | [2, 3] | 0,1,4,5,6,7,8,9 · 2 | [4, 5] | 0,1,2,3,6,7,8,9 · 3 | [6, 7] | 0,1,2,3,4,5,8,9

In The fold split, made visible, given the rows so far: what is the next one — the row where round is 4?

Correct: 4 | [8, 9] | 0,1,2,3,4,5,6,7

roundvalidation foldtraining indices
0[0, 1]2,3,4,5,6,7,8,9
1[2, 3]0,1,4,5,6,7,8,9
2[4, 5]0,1,2,3,6,7,8,9
3[6, 7]0,1,2,3,4,5,8,9
4[8, 9]0,1,2,3,4,5,6,7

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The five folds tile the range with no overlap and no gap, so each point is held out once — the property the sub-question is testing.

76. The fold split, made visible

Worked example

np.array_split cuts the index range into k contiguous blocks. On 10 points with k = 5 you get five clean pairs — easy to check by eye:

import numpy as np
idx = np.arange(10)                    # 10 data points, indices 0..9
folds = np.array_split(idx, 5)        # 5 contiguous folds
for i, fo in enumerate(folds):
    print("fold", i, "=", fo.tolist())
roundvalidation foldtraining indices
0[0, 1]2,3,4,5,6,7,8,9
1[2, 3]0,1,4,5,6,7,8,9
2[4, 5]0,1,2,3,6,7,8,9
3[6, 7]0,1,2,3,4,5,8,9
4[8, 9]0,1,2,3,4,5,6,7

Every index 0–9 appears in exactly one validation fold

Why: The five folds tile the range with no overlap and no gap, so each point is held out once — the property the sub-question is testing. array_split also handles non-divisible sizes by making the first folds one larger.

77. What has to be given first: Implement 5-fold CV

Missing information

Discussion prompt

Split indices, loop the held-out fold, train on the concatenated rest, score, average. Uses Ridge on diabetes — self-contained:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Averaging [0.3217, 0.4405, 0.4221, 0.4247, 0.4420] gives 0.4102 — the cross-validated estimate of how Ridge generalizes. The concatenate line rebuilds the training set from every fold except the held-out one.

78. Implement 5-fold CV

Worked example

Split indices, loop the held-out fold, train on the concatenated rest, score, average. Uses Ridge on diabetes — self-contained:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
X, y = load_diabetes(return_X_y=True)

def kfold(model, X, y, k=5):
    folds = np.array_split(np.arange(len(X)), k)
    scores = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        model.fit(X[tr], y[tr])
        scores.append(model.score(X[te], y[te]))
    return np.array(scores)

sc = kfold(Ridge(alpha=1.0), X, y, 5)
print("per-fold R2 =", sc.round(4))
print("mean R2 =", round(sc.mean(), 4))
roundheld-out foldR² (verified)
0fold 00.3217
1fold 10.4405
2fold 20.4221
3fold 30.4247
4fold 40.4420

Mean R² = 0.4102 across the five folds

Why: Averaging [0.3217, 0.4405, 0.4221, 0.4247, 0.4420] gives 0.4102 — the cross-validated estimate of how Ridge generalizes. The concatenate line rebuilds the training set from every fold except the held-out one.

79. Watch it run: Implement 5-fold CV

Pattern

Step through it

Step through Implement 5-fold CV one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: round is 0
  2. Step 2: round is 1
  3. Step 3: round is 2
  4. Step 4: round is 3
  5. Step 5: round is 4

80. Guess the shape of the answer: Verify against cross_val_score

Estimation

Predict first

sklearn's KFold(shuffle=False) uses the same contiguous split, so the scores must match fold for fold. That equality is the graded check:

Commit before you compute: what does Verify against cross_val_score come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Fold-for-fold identical — because shuffle=False matches array_split

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. sklearn's default KFold does NOT shuffle and splits contiguously, exactly like np.array_split.

81. Verify against cross_val_score

Worked example

sklearn's KFold(shuffle=False) uses the same contiguous split, so the scores must match fold for fold. That equality is the graded check:

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)
cvs = cross_val_score(Ridge(alpha=1.0), X, y,
                      cv=KFold(n_splits=5, shuffle=False))
print("sklearn per-fold =", cvs.round(4))
print("sklearn mean     =", round(cvs.mean(), 4))
sourceper-fold R²mean
from scratch[0.3217, 0.4405, 0.4221, 0.4247, 0.4420]0.4102
sklearn[0.3217, 0.4405, 0.4221, 0.4247, 0.4420]0.4102
match?identicalyes

Fold-for-fold identical — because shuffle=False matches array_split

Why: sklearn's default KFold does NOT shuffle and splits contiguously, exactly like np.array_split. Match that and the numbers agree; if you had shuffled, they'd differ even though both are 'correct' 5-fold CV.

82. Work backwards from the answer: Verify against cross_val_score

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Fold-for-fold identical — because shuffle=False matches array_split

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

sklearn's KFold(shuffle=False) uses the same contiguous split, so the scores must match fold for fold. That equality is the graded check:

83. Something is wrong here: shuffling one side but not the other

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compare your contiguous-split CV against cross_val_score with KFold(shuffle=True) and call it a mismatch bug.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Shuffling reorders points before splitting, so the folds hold different data — a legitimately different estimate.

Match the split protocol on both sides before comparing.

Why: Shuffling reorders points before splitting, so the folds hold different data — a legitimately different estimate. You'll 'fix' a bug that isn't there and waste exam minutes chasing a phantom.

84. Trap: shuffling one side but not the other

Trap

The trap

Compare your contiguous-split CV against cross_val_score with KFold(shuffle=True) and call it a mismatch bug.

Your 0.4102 ≠ sklearn's shuffled 0.48…

Why: Shuffling reorders points before splitting, so the folds hold different data — a legitimately different estimate. You'll 'fix' a bug that isn't there and waste exam minutes chasing a phantom.

The fix

Match the split protocol on both sides before comparing.

Compare against KFold(shuffle=False)

Why: Your np.array_split is contiguous, so verify against the contiguous sklearn split. Then both give 0.4102. When comparing implementations, hold the randomness fixed — mismatched shuffling is the usual false alarm.

85. Which of these survive contact with Lesson 39: Mock Exam — Phase 1 Coding?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
The theory sub-question for this problem: why must you center the data first? Hold that thought — the derivation answers it in two slides.; For centered data Xc (each column has mean 0), the sample covariance is C = XcᵀXc / (n−1). Entry (j,k) is the covariance of feature j with feature k.; The running sample — eight numbers chosen so every quantity is a clean integer:
Breaks
PCA is 'just the SVD of the data', so decompose X directly — centering is an extra line I can skip.; Convergence is slow at η = 0.1, so push it way up to η = 1.1 to finish in fewer steps.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 39: Mock Exam — Phase 1 Coding puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

86. Theory sub-questions

Section

Part 6 of 6 — and code review

87. Rebuild the recipe: The four gotchas, one line each

Ranking

Put in order

These are the steps of The four gotchas, one line each, scrambled. Put them back in order before the next slide shows you.

  1. PCA: center before decomposing — uncentered PC1 chases the mean, not the spread
  2. Gradient descent: keep η < 2/L — bigger overshoots and diverges, it is not 'faster'
  3. MLE variance: divide by n (biased MLE), not n−1 (unbiased) — match the word in the prompt
  4. k-fold CV: each point is held out exactly once; match shuffle=False when verifying

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

88. The four gotchas, one line each

Pattern

  1. PCA: center before decomposing — uncentered PC1 chases the mean, not the spread
  2. Gradient descent: keep η < 2/L — bigger overshoots and diverges, it is not 'faster'
  3. MLE variance: divide by n (biased MLE), not n−1 (unbiased) — match the word in the prompt
  4. k-fold CV: each point is held out exactly once; match shuffle=False when verifying

89. Where does it stop working: The four gotchas, one line each

Edge cases

Discussion prompt

The four gotchas, one line each works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. PCA: center before decomposing — uncentered PC1 chases the mean, not the spread
  2. Gradient descent: keep η < 2/L — bigger overshoots and diverges, it is not 'faster'
  3. MLE variance: divide by n (biased MLE), not n−1 (unbiased) — match the word in the prompt
  4. k-fold CV: each point is held out exactly once; match shuffle=False when verifying

90. Rule out three: Check yourself — the PCA sub-question

Elimination

Eliminate the wrong options

Your from-scratch PCA must match sklearn's explained_variance_ratio_. The most likely cause of a qualitative mismatch is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. forgetting to center the data before decomposing
  • B. using SVD instead of eigendecomposition of the covariance
  • C. computing more components than you need
  • D. using float64 instead of float32

Survives elimination: A

Why: sklearn subtracts the column means internally. A from-scratch version must center explicitly, or the top component chases the mean offset (on the shifted toy, PC1 flipped to [−0.707,−0.707] and S₁ ballooned to 28.43) and every EVR is wrong.

91. Check yourself — the PCA sub-question

Check

The theory question graders pair with the PCA implementation.

Check your understanding

Your from-scratch PCA must match sklearn's explained_variance_ratio_. The most likely cause of a qualitative mismatch is:

  • A. forgetting to center the data before decomposing (correct)
  • B. using SVD instead of eigendecomposition of the covariance
  • C. computing more components than you need
  • D. using float64 instead of float32

Answer: A

Why: sklearn subtracts the column means internally. A from-scratch version must center explicitly, or the top component chases the mean offset (on the shifted toy, PC1 flipped to [−0.707,−0.707] and S₁ ballooned to 28.43) and every EVR is wrong.

Why B tempts people
SVD of the centered data and eigendecomposition of the covariance give the SAME components — λ = S²/(n−1). We verified they agree to the digit. Not a source of mismatch.
Why C tempts people
Extra components don't disturb the leading ones; you just compute more of them. The top-3 EVR is unchanged.
Why D tempts people
float precision moves results by ~1e-15, not the qualitative flip that skipping centering causes.

92. Answer it before you see the options: Check yourself — the MLE sub-question

Prediction

Predict first

For the 8-point sample the sum of squared deviations from the mean is 32. What is the maximum-likelihood estimate of σ²?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 32/8 = 4.0 — divide by n

Why: Maximizing the Gaussian log-likelihood gives σ̂² = (1/n)Σ(xᵢ−μ̂)². With Σ = 32 and n = 8 that is 32/8 = 4.0. It is biased low because it reuses μ̂, but it IS the MLE.

93. Check yourself — the MLE sub-question

Check

Divisor discipline — the classic point-loser.

Check your understanding

For the 8-point sample the sum of squared deviations from the mean is 32. What is the maximum-likelihood estimate of σ²?

  • A. 32/8 = 4.0 — divide by n (correct)
  • B. 32/7 ≈ 4.571 — divide by n−1
  • C. 32/9 ≈ 3.556 — divide by n+1
  • D. √(32/8) = 2.0

Answer: A

Why: Maximizing the Gaussian log-likelihood gives σ̂² = (1/n)Σ(xᵢ−μ̂)². With Σ = 32 and n = 8 that is 32/8 = 4.0. It is biased low because it reuses μ̂, but it IS the MLE.

Why B tempts people
32/7 = 4.571 is the UNBIASED sample variance (Bessel's n−1 correction), a different estimator from the MLE. np.var(d, ddof=1) returns it.
Why C tempts people
n+1 minimizes a different (mean-squared-error) criterion; it is neither the MLE nor the unbiased estimator.
Why D tempts people
2.0 is σ̂ (the standard deviation), not σ̂² (the variance). The question asked for the variance.

94. Answer it before you see the options: Check yourself — gradient descent…

Prediction

Predict first

Minimizing f(w) = (w−5)² (so f''= 2, L = 2) with plain gradient descent, which learning rate makes the iterates diverge?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: η = 1.1, because 1.1 > 2/L = 1

Why: The step multiplier is |1 − 2η|. Convergence needs it below 1, i.e. η < 2/L = 1. At η = 1.1 the multiplier is 1.2 > 1, so the distance to 5 grows each step — we measured w → −45497 after 50 steps.

95. Check yourself — gradient descent stability

Check

The learning-rate bound in one question.

Check your understanding

Minimizing f(w) = (w−5)² (so f''= 2, L = 2) with plain gradient descent, which learning rate makes the iterates diverge?

  • A. η = 1.1, because 1.1 > 2/L = 1 (correct)
  • B. η = 0.5, because larger steps always overshoot
  • C. η = 0.1, because tiny steps accumulate error
  • D. none — gradient descent converges for every η > 0

Answer: A

Why: The step multiplier is |1 − 2η|. Convergence needs it below 1, i.e. η < 2/L = 1. At η = 1.1 the multiplier is 1.2 > 1, so the distance to 5 grows each step — we measured w → −45497 after 50 steps.

Why B tempts people
η = 0.5 is safely under the bound of 1 and in fact converges fastest of the rates we tried (w → 5.0). Larger-but-bounded is good, not divergent.
Why C tempts people
η = 0.1 converges (slowly) to 4.9999. Small steps are inefficient, not divergent — they don't accumulate error on a convex quadratic.
Why D tempts people
False: above η = 2/L the method diverges. η = 1 already oscillates forever without settling; η = 1.1 explodes.

96. Rule out three: Check yourself — cross-validation

Elimination

Eliminate the wrong options

In standard k-fold cross-validation, how many times is each data point used for validation, and what makes your from-scratch scores match cross_val_score exactly?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Exactly once; match sklearn's contiguous split with KFold(shuffle=False)
  • B. k times; the point is validated in every fold
  • C. Exactly once; but the scores can never match sklearn from scratch
  • D. k−1 times; once per training fold

Survives elimination: A

Why: The folds partition the data, and each is held out in exactly one round, so every point validates once. np.array_split is contiguous, so it matches KFold(shuffle=False) fold-for-fold — both gave mean R² = 0.4102.

97. Check yourself — cross-validation

Check

The CV sub-question, and the usual verification pitfall.

Check your understanding

In standard k-fold cross-validation, how many times is each data point used for validation, and what makes your from-scratch scores match cross_val_score exactly?

  • A. Exactly once; match sklearn's contiguous split with KFold(shuffle=False) (correct)
  • B. k times; the point is validated in every fold
  • C. Exactly once; but the scores can never match sklearn from scratch
  • D. k−1 times; once per training fold

Answer: A

Why: The folds partition the data, and each is held out in exactly one round, so every point validates once. np.array_split is contiguous, so it matches KFold(shuffle=False) fold-for-fold — both gave mean R² = 0.4102.

Why B tempts people
A point is held out (validated) in only ONE round; in the other k−1 rounds it is a training point. k times would mean testing on training data.
Why C tempts people
They match exactly when the split protocol is the same — we verified [0.3217,…,0.4420] identically. 'Never' is wrong.
Why D tempts people
k−1 is how many rounds a point is used for TRAINING, not validation. Validation happens exactly once.

98. Code-review criteria

Concept

criterionwhat the reviewer checks
correctnessmatches the library / known value to the digit
vectorizationNumPy ops, not Python loops where avoidable
edge casescentering, ddof, fold boundaries, array shapes
readabilitynamed variables, one brief comment per block

Correct-but-slow still passes; correct-and-clean is the bar. A subtle edge-case miss — uncentered PCA, wrong ddof, a shuffled comparison — is the usual point-loser, not a wrong formula.

99. By analogy: Code-review criteria

Analogy

Discussion prompt

Explain Code-review criteria by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Correct-but-slow still passes; correct-and-clean is the bar. A subtle edge-case miss — uncentered PCA, wrong ddof, a shuffled comparison — is the usual point-loser, not a wrong formula.

100. Your turn: rebuild the four

Section

The project

101. Project: the four from memory

Concept

Now do it yourself, slides closed. Re-implement each algorithm from scratch and verify it against the library. You've derived every piece — assemble it under exam conditions.

#milestoneverify against
1Gradient descent on (w−5)²analytic minimum w = 5
2PCA on digits (EVR, k@95%)sklearn PCA
35-fold CV (Ridge, diabetes)cross_val_score

Build rules: type every line yourself, predict the output before running, and when something errors, read the shapes in the message — don't delete the error.

102. Break it if you can: Project: the four from memory

Counterexample

Discussion prompt

Build rules: type every line yourself, predict the output before running, and when something errors, read the shapes in the message — don't delete the error.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

103. Milestone 1 — gradient descent

Worked example

Your turn: minimize (w−5)² from w = 0 with η = 0.1 for 60 steps. Predict the final w out loud before printing.

Hint: gradient is 2(w−5); the update is w -= lr * grad. The final w should round to the vertex you found by calculus.

import numpy as np
# Your turn: minimize f(w) = (w - 5)**2 from w = 0 with lr = 0.1
grad = lambda w: 2 * (w - 5)
w, lr, hist = 0.0, 0.1, []
for _ in range(60):
    w -= lr * grad(w)
    hist.append(w)
print("final w =", round(w, 4))
print("f(final) =", round((w - 5)**2, 6))
checkexpected
final w5.0
f(final)0.0
matches analytic min?yes (vertex at 5)

104. Milestone 2 — PCA on digits

Worked example

Your turn: center digits, SVD, and report top-3 EVR plus the count of components for 95%. Predict whether it beats or ties the earlier k = 29.

Hint: Xc = X − X.mean(0), take S from np.linalg.svd(...)[1], then evr = S²/ΣS² and np.searchsorted(np.cumsum(evr), 0.95) + 1.

import numpy as np
from sklearn.datasets import load_digits
# Your turn: PCA on digits, report EVR top-3 and k for 95%
X = load_digits().data
Xc = X - X.mean(0)
S = np.linalg.svd(Xc, full_matrices=False)[1]
evr = S**2 / np.sum(S**2)
print("EVR[:3] =", evr[:3].round(4))
print("k95 =", int(np.searchsorted(np.cumsum(evr), 0.95) + 1))
checkexpected
EVR[:3][0.1489, 0.1362, 0.1179]
k9529
centered first?yes (required)

105. Milestone 3 — 5-fold CV and verify

Worked example

Your turn: build 5-fold CV for Ridge on diabetes and confirm your mean matches cross_val_score. Predict the match before running.

Hint: split with np.array_split, train on the concatenated other folds, average the scores. Verify against KFold(5, shuffle=False).

import numpy as np
from sklearn.linear_model import Ridge
from sklearn.datasets import load_diabetes
from sklearn.model_selection import cross_val_score, KFold
X, y = load_diabetes(return_X_y=True)

def kfold(model, X, y, k=5):
    folds = np.array_split(np.arange(len(X)), k)
    s = []
    for i in range(k):
        te = folds[i]
        tr = np.concatenate([folds[j] for j in range(k) if j != i])
        model.fit(X[tr], y[tr]); s.append(model.score(X[te], y[te]))
    return np.mean(s)

mine = kfold(Ridge(1.0), X, y, 5)
lib = cross_val_score(Ridge(1.0), X, y, cv=KFold(5, shuffle=False)).mean()
print("mine   =", round(mine, 4))
print("sklearn=", round(lib, 4))
print("match  =", np.isclose(mine, lib))
sourcemean R²
mine (from scratch)0.4102
sklearn0.4102
np.iscloseTrue

106. What each one costs: Milestone 3 — 5-fold CV and verify

Trade off

Comparison matrix

From Milestone 3 — 5-fold CV and verify: every row here is a choice with a cost. Fill the mean R² column, then say which row you would actually pick and what you give up for it.

sourcemean R²
mine (from scratch)0.4102
sklearn0.4102
np.iscloseTrue

107. The full answer key

Concept

import numpy as np
from sklearn.datasets import load_digits, load_diabetes
from sklearn.linear_model import Ridge

# 1. PCA
X = load_digits().data
S = np.linalg.svd(X - X.mean(0), full_matrices=False)[1]
evr = S**2 / np.sum(S**2)
k95 = int(np.searchsorted(np.cumsum(evr), 0.95) + 1)

# 2. Gradient descent on (w-5)**2
w = 0.0
for _ in range(60):
    w -= 0.1 * 2 * (w - 5)

# 3. MLE Gaussian
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu, var = d.mean(), ((d - d.mean())**2).mean()

# 4. 5-fold CV
Xd, yd = load_diabetes(return_X_y=True)
folds = np.array_split(np.arange(len(Xd)), 5)
r2 = []
for i in range(5):
    tr = np.concatenate([folds[j] for j in range(5) if j != i])
    m = Ridge(1.0).fit(Xd[tr], yd[tr]); r2.append(m.score(Xd[folds[i]], yd[folds[i]]))

print("PCA  k@95% =", k95, " EVR[:3] =", evr[:3].round(4))
print("GD   w =", round(w, 4))
print("MLE  mu =", mu, " var =", var)
print("CV   R2 =", round(np.mean(r2), 4))
problemresultverified vs
PCAk@95% = 29, EVR [0.1489, 0.1362, 0.1179]sklearn PCA
gradient descentw → 5.0analytic minimum
MLE Gaussianμ = 5.0, σ² = 4.0np.var (ddof=0)
k-fold CVR² = 0.4102cross_val_score

Four algorithms, four library matches — the coding foundation Phase 2 builds every model on.

108. Fill in: result for The full answer key

Comparison

Comparison matrix

From The full answer key: refill the result column from what you know. The rest of the table is as it appeared.

problemresultverified vs
PCAk@95% = 29, EVR [0.1489, 0.1362, 0.1179]sklearn PCA
gradient descentw → 5.0analytic minimum
MLE Gaussianμ = 5.0, σ² = 4.0np.var (ddof=0)
k-fold CVR² = 0.4102cross_val_score

109. Show it off & Phase 2 prep

Concept

Slides closed, out loud: explain (1) why PCA must center, (2) the η < 2/L bound and what happens above it, (3) why the MLE variance divides by n, and (4) why each CV point is held out exactly once.

Before Phase 2: re-implement any of the four that still needed a hint, add a shape assertion to each, and set up a clean PyTorch environment — Phase 2 moves all coding to PyTorch starting Lesson 40.

110. Connect it up: Lesson 39: Mock Exam — Phase 1 Coding

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — How the coding exam works · Problem 1 — PCA · Problem 2 — Gradient descent · Problem 3 — MLE · Problem 4 — Cross-validation · Theory sub-questions. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

111. Coding-exam debrief

Recap

algorithmthe one gotchaverified number
PCAcenter before decomposingk@95% = 29 of 64
gradient descentη < 2/L or it divergesw → 5.0
MLE σ²divide by n, not n−1σ² = 4.0
k-fold CVeach point held out exactly onceR² = 0.4102

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 39 (Week 13 — Coding Mock Exam) — Barron · USAAIO Round 2 Preparation, 2026
  2. scikit-learn: PCA, Ridge, cross_val_score / KFold
  3. NumPy linalg.svd
  4. Every EVR, gradient-descent trace, MLE, and CV score produced by real execution — numpy 2.2.6 + scikit-learn 1.9, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108