Lesson 38: Mock Exam — Phase 1 Theory

USAAIO Lesson 38, from Week 13, fully worked: a timed Phase 1 theory mock exam turned into an olympiad-grade walkthrough. Nine exam questions span linear algebra with symmetric eigendecomposition, MLE, Hessian curvature, KL divergence, AUC, kernels and Mercer's theorem, generalization, the softmax-with-cross-entropy gradient, and the p-value trap. Each is derived one move per beat with nothing skipped, every number was produced by real numpy, torch, and sklearn execution, and there is a from-scratch metrics-evaluator project verified against sklearn. The lesson runs to 69 slides.

Subject: Machine Learning · 125 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Mock Exam — Phase 1 Theory

Title

USAAIO · Lesson 38 · Week 13

Nine exam questions, every one derived one move at a time with nothing skipped, then confirmed in numpy/torch/sklearn. Attempt each cold, watch the full derivation, then log every miss. This is the whole Phase-1 toolkit under one roof.

2. By the end of this mock you can

Objectives

  1. State the spectral theorem and find the eigenpairs of a 2×2 symmetric matrix by hand, checking trace and orthogonality
  2. Derive the Gaussian MLE μ̂, σ̂² from the log-likelihood and know why σ̂² divides by n
  3. Classify a stationary point from its Hessian eigenvalues (min / max / saddle) and derive the softmax+cross-entropy gradient p − y
  4. State the sign, asymmetry, and CE = H + KL identity for KL divergence, and read AUC as a ranking probability
  5. Give Mercer's PSD-Gram condition for a valid kernel, name the overfitting signature, and never misread a p-value
  6. Build a metrics evaluator (precision / recall / F1) from scratch and match sklearn exactly

3. What survived from Phase 1 Review & Mock Prep?

Warm-up

Discussion prompt

Before we open Lesson 38: Mock Exam — Phase 1 Theory: without looking back, what was the main idea of Phase 1 Review & Mock Prep, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a retrieval-practice consolidation of Phase 1 — the four pillars (linear algebra, probability/information, calculus/optimization, ML foundations), a synthesis capstone (OLS four equivalent ways), exam-pace retrieval drills, and a gap-identification protocol before the first mock exam.

4. How this mock is built

Concept

A real USAAIO theory section mixes multiple choice, short derivations, and explanations under time. This mock mirrors it — but for each question we do the full derivation you'd have to produce cold, then confirm every number by execution.

for each exam questionwhat you'll do
1. attemptderive on paper, slides closed
2. derivewatch it built one move per beat
3. commitpick an answer on the check slide
4. verifysee the numpy/torch/sklearn output
5. logrecord every miss for the debrief

Rule of the room: no notes, no calculator, attempt cold. The struggle is the encoding — revealing early throws the value away.

5. Fill in: what you'll do for How this mock is built

Comparison

Comparison matrix

From How this mock is built: refill the what you'll do column from what you know. The rest of the table is as it appeared.

for each exam questionwhat you'll do
1. attemptderive on paper, slides closed
2. derivewatch it built one move per beat
3. commitpick an answer on the check slide
4. verifysee the numpy/torch/sklearn output
5. logrecord every miss for the debrief

6. Q1 — Linear Algebra

Section

Part 1 of 11 — symmetric eigenpairs

7. Why symmetric matrices are special

Intuition

A general square matrix can rotate space — its eigenvalues may be complex and its eigenvectors need not be perpendicular. A symmetric matrix (A = Aᵀ) only ever stretches along fixed perpendicular axes. It never rotates.

So a symmetric matrix always has real eigenvalues and a full set of orthogonal eigenvectors — you can pick an axis-aligned frame in which A is pure scaling. That is the spectral theorem, and it powers PCA, covariance, kernels, and Hessians.

Our running matrix for this question is A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.

8. Break it if you can: Why symmetric matrices are special

Counterexample

Discussion prompt

Our running matrix for this question is A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. The eigenvalue equation

Concept

An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.

\[ A v = \lambda v \;\Longleftrightarrow\; (A - \lambda I)\,v = 0 \]

For a nonzero v to exist, A − λI must be singular — its determinant must be zero. That single condition is the characteristic equation.

\[ \det(A - \lambda I) = 0 \]

10. By analogy: The eigenvalue equation

Analogy

Discussion prompt

Explain The eigenvalue equation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.

11. What has to happen first: Solve the characteristic equation

Ranking

Put in order

Put the moves of Solve the characteristic equation into the order they have to happen.

  1. Write A − λI for our matrix
  2. Take the 2×2 determinant ad − bc
  3. Expand (2−λ)² and set to zero
  4. Factor the quadratic

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Subtract λ from each diagonal entry of [[2,1],[1,2]].

12. Solve the characteristic equation

Worked example

Write A − λI for our matrix

Why: Subtract λ from each diagonal entry of [[2,1],[1,2]].

\[ A - \lambda I = \begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} \]

Take the 2×2 determinant ad − bc

Why: det = (2−λ)(2−λ) − (1)(1). The off-diagonal product is 1·1 = 1.

\[ \det(A-\lambda I) = (2-\lambda)^2 - 1 \]

Expand (2−λ)² and set to zero

Why: (2−λ)² = 4 − 4λ + λ², so the equation is λ² − 4λ + 4 − 1 = 0.

\[ \lambda^2 - 4\lambda + 3 = 0 \]

Factor the quadratic

Why: λ² − 4λ + 3 = (λ − 1)(λ − 3), so the roots are λ = 1 and λ = 3 — both real, as the spectral theorem promised.

\[ (\lambda - 1)(\lambda - 3) = 0 \;\Longrightarrow\; \lambda = 1,\; 3 \]

13. Decode the notation: Solve the characteristic equation

Notation

Annotate

From Solve the characteristic equation — read this one piece at a time. What is each part doing?

On: \( A - \lambda I = \begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} \)

  • Subtract λ from each diagonal entry of [[2,1],[1,2]].
  • det = (2−λ)(2−λ) − (1)(1). The off-diagonal product is 1·1 = 1.
  • (2−λ)² = 4 − 4λ + λ², so the equation is λ² − 4λ + 4 − 1 = 0.

14. Guess the shape of the answer: Sanity-check with trace and determinant

Estimation

Predict first

Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.

Commit before you compute: what does Sanity-check with trace and determinant come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Determinant check: product of eigenvalues = det

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. det(A) = 2·2 − 1·1 = 3, and 1·3 = 3.

15. Sanity-check with trace and determinant

Worked example

Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.

Trace check: sum of eigenvalues = trace

Why: trace(A) = 2 + 2 = 4, and 1 + 3 = 4. ✓

\[ \lambda_1 + \lambda_2 = 1 + 3 = 4 = \operatorname{trace}(A) \]

Determinant check: product of eigenvalues = det

Why: det(A) = 2·2 − 1·1 = 3, and 1·3 = 3. ✓ Both invariants agree, so λ = {1, 3} is correct.

\[ \lambda_1 \lambda_2 = 1 \cdot 3 = 3 = \det(A) \]

16. Work backwards from the answer: Sanity-check with trace and determinant

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Determinant check: product of eigenvalues = det

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.

17. Plan first: Find the eigenvectors, entry by entry

Step zero

Discussion prompt

Find the eigenvectors, entry by entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: For λ = 3: solve (A − 3I)v = 0

Answer:

  1. For λ = 3: solve (A − 3I)v = 0
  2. For λ = 1: solve (A − I)v = 0
  3. Check they're orthogonal

18. Find the eigenvectors, entry by entry

Worked example

For λ = 3: solve (A − 3I)v = 0

Why: A − 3I = [[−1, 1], [1, −1]]. The first row says −v₁ + v₂ = 0, i.e. v₂ = v₁.

\[ \begin{bmatrix} -1 & 1 \\ 1 & -1 \end{bmatrix}\begin{bmatrix} v_1 \\ v_2 \end{bmatrix} = 0 \;\Longrightarrow\; v = \tfrac{1}{\sqrt2}\begin{bmatrix} 1 \\ 1 \end{bmatrix} \]

For λ = 1: solve (A − I)v = 0

Why: A − I = [[1, 1], [1, 1]]. The row says v₁ + v₂ = 0, i.e. v₂ = −v₁.

\[ \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}\begin{bmatrix} v_1 \\ v_2 \end{bmatrix} = 0 \;\Longrightarrow\; v = \tfrac{1}{\sqrt2}\begin{bmatrix} 1 \\ -1 \end{bmatrix} \]

Check they're orthogonal

Why: [1,1]·[1,−1] = 1 − 1 = 0. Perpendicular eigenvectors — exactly what the spectral theorem guarantees for a symmetric matrix.

19. Draw the shape of it: Find the eigenvectors, entry by entry

Blank canvas

Draw it

Draw what Find the eigenvectors, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

20. What has to be given first: Confirm in numpy

Missing information

Discussion prompt

Use eigh (the symmetric solver — it returns sorted real eigenvalues and orthonormal eigenvectors). Runnable as written:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

eigh sorts ascending, so column 0 pairs with λ = 1 and column 1 with λ = 3.

21. Confirm in numpy

Worked example

Use eigh (the symmetric solver — it returns sorted real eigenvalues and orthonormal eigenvectors). Runnable as written:

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
vals, vecs = np.linalg.eigh(A)   # symmetric solver
print("eigenvalues:", vals)               # [1. 3.]
print("eigenvector cols:\n", vecs.round(4))
print("orthogonal? v0.v1 =", round(float(vecs[:,0] @ vecs[:,1]), 12))
print("trace", np.trace(A), "= sum eig", vals.sum())

eigenvalues come back [1., 3.] — matches the hand factoring

Why: eigh sorts ascending, so column 0 pairs with λ = 1 and column 1 with λ = 3.

quantityvalue (verified)
eigenvalues[1., 3.]
eigenvector for λ=1[−0.7071, 0.7071] (∝ [1, −1])
eigenvector for λ=3[0.7071, 0.7071] (∝ [1, 1])
v0 · v1 (orthogonality)0.0
trace = Σ eigenvalues4.0 = 4.0

22. What each one costs: Confirm in numpy

Trade off

Comparison matrix

From Confirm in numpy: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.

quantityvalue (verified)
eigenvalues[1., 3.]
eigenvector for λ=1[−0.7071, 0.7071] (∝ [1, −1])
eigenvector for λ=3[0.7071, 0.7071] (∝ [1, 1])
v0 · v1 (orthogonality)0.0
trace = Σ eigenvalues4.0 = 4.0

23. Something is wrong here: reading eigenvalues off the diagonal

Anomaly

Predict first

A student writes this, and it looks reasonable:

A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix.

Solve det(A − λI) = 0 — the off-diagonals matter.

Why: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix. Here the off-diagonal 1s couple the axes, so the true spectrum is {1, 3}, not {2, 2}. It even fails the det check: 2·2 = 4 ≠ det = 3.

24. Trap: reading eigenvalues off the diagonal

Trap

The trap

A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.

λ = 2, 2 ✗

Why: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix. Here the off-diagonal 1s couple the axes, so the true spectrum is {1, 3}, not {2, 2}. It even fails the det check: 2·2 = 4 ≠ det = 3.

The fix

Solve det(A − λI) = 0 — the off-diagonals matter.

λ = 1, 3 ✓

Why: The characteristic equation λ² − 4λ + 3 = 0 factors to (λ−1)(λ−3). Sum 4 = trace and product 3 = det both check out. Only trust 'diagonal = eigenvalues' when the matrix is triangular.

25. Rebuild the recipe: The eigenpair recipe (any 2×2)

Ranking

Put in order

These are the steps of The eigenpair recipe (any 2×2), scrambled. Put them back in order before the next slide shows you.

  1. Characteristic equation: det(A − λI) = 0 → λ² − (trace)λ + (det) = 0
  2. Solve for λ (factor or quadratic formula); confirm Σλ = trace, Πλ = det
  3. Each λ: solve (A − λI)v = 0 for the direction v, then normalize
  4. Symmetric A: eigenvalues are real and eigenvectors come out orthogonal — check v_i · v_j = 0
  5. In code: np.linalg.eigh for symmetric, np.linalg.eig for general matrices

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

26. The eigenpair recipe (any 2×2)

Pattern

  1. Characteristic equation: det(A − λI) = 0 → λ² − (trace)λ + (det) = 0
  2. Solve for λ (factor or quadratic formula); confirm Σλ = trace, Πλ = det
  3. Each λ: solve (A − λI)v = 0 for the direction v, then normalize
  4. Symmetric A: eigenvalues are real and eigenvectors come out orthogonal — check v_i · v_j = 0
  5. In code: np.linalg.eigh for symmetric, np.linalg.eig for general matrices

27. Where does it stop working: The eigenpair recipe (any 2×2)

Edge cases

Discussion prompt

The eigenpair recipe (any 2×2) works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Characteristic equation: det(A − λI) = 0 → λ² − (trace)λ + (det) = 0
  2. Solve for λ (factor or quadratic formula); confirm Σλ = trace, Πλ = det
  3. Each λ: solve (A − λI)v = 0 for the direction v, then normalize
  4. Symmetric A: eigenvalues are real and eigenvectors come out orthogonal — check v_i · v_j = 0
  5. In code: np.linalg.eigh for symmetric, np.linalg.eig for general matrices

28. Rule out three: Q1 — the exam question

Elimination

Eliminate the wrong options

Which property is guaranteed for a real symmetric matrix but NOT for a general square matrix?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Real eigenvalues and orthogonal eigenvectors
  • B. All eigenvalues strictly positive
  • C. Eigenvalues equal to the diagonal entries
  • D. Exactly one eigenvector

Survives elimination: A

Why: The spectral theorem: a real symmetric matrix has real eigenvalues and an orthonormal eigenbasis. We saw it directly — [[2,1],[1,2]] gave real λ = {1,3} with perpendicular eigenvectors. A general matrix can have complex eigenvalues and non-orthogonal eigenvectors.

29. Q1 — the exam question

Check

Attempt cold, then commit.

Check your understanding

Which property is guaranteed for a real symmetric matrix but NOT for a general square matrix?

  • A. Real eigenvalues and orthogonal eigenvectors (correct)
  • B. All eigenvalues strictly positive
  • C. Eigenvalues equal to the diagonal entries
  • D. Exactly one eigenvector

Answer: A

Why: The spectral theorem: a real symmetric matrix has real eigenvalues and an orthonormal eigenbasis. We saw it directly — [[2,1],[1,2]] gave real λ = {1,3} with perpendicular eigenvectors. A general matrix can have complex eigenvalues and non-orthogonal eigenvectors.

Why B tempts people
Strictly positive eigenvalues is positive-DEFINITENESS — a stronger, separate condition. A symmetric matrix can have negative eigenvalues (e.g. [[0,1],[1,0]] has λ = {−1, 1}).
Why C tempts people
Diagonal entries equal the eigenvalues only for triangular matrices. Our symmetric [[2,1],[1,2]] has diagonal {2,2} but eigenvalues {1,3} — the exact trap from the previous slide.
Why D tempts people
A symmetric n×n matrix has a FULL set of n orthogonal eigenvectors, not one. Even repeated eigenvalues still yield an orthogonal eigenbasis.

30. Q2 — MLE (Probability)

Section

Part 2 of 11 — the Gaussian fit

31. What maximum likelihood asks

Intuition

You have data and a family of models (here: Gaussians with unknown mean μ and variance σ²). Maximum likelihood picks the parameters under which the data you actually observed was most probable.

Our sample for this question: d = [2, 4, 4, 4, 5, 5, 7, 9] (eight numbers). We'll find the μ̂, σ̂² that make this exact sample most likely under a normal model.

Spoiler you should predict: the MLE mean is just the sample average, and the MLE variance is the average squared deviation. But deriving it — and seeing the ÷n — is the exam skill.

32. Teach it back: What maximum likelihood asks

Explain it

Discussion prompt

Explain What maximum likelihood asks to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Spoiler you should predict: the MLE mean is just the sample average, and the MLE variance is the average squared deviation. But deriving it — and seeing the ÷n — is the exam skill.

33. Likelihood → log-likelihood

Concept

Assuming the n points are independent, the likelihood is the product of Gaussian densities. Products are painful to differentiate, so take the log — it turns the product into a sum and doesn't move the maximum (log is increasing).

\[ \mathcal{L}(\mu,\sigma^2) = \prod_{i=1}^{n} \frac{1}{\sqrt{2\pi\sigma^2}}\exp\!\Big(-\frac{(d_i-\mu)^2}{2\sigma^2}\Big) \]

\[ \ell(\mu,\sigma^2) = -\frac{n}{2}\log(2\pi) - \frac{n}{2}\log\sigma^2 - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(d_i-\mu)^2 \]

34. Plan first: Maximize over μ

Step zero

Discussion prompt

Maximize over μ — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Differentiate ℓ with respect to μ

Answer:

  1. Differentiate ℓ with respect to μ
  2. Set to zero and solve
  3. Plug in our sample

35. Maximize over μ

Worked example

Differentiate ℓ with respect to μ

Why: Only the last term depends on μ. Chain rule on (dᵢ − μ)²: derivative is −2(dᵢ − μ), and the −1/(2σ²) out front flips signs.

\[ \frac{\partial \ell}{\partial \mu} = \frac{1}{\sigma^2}\sum_{i=1}^{n}(d_i - \mu) \]

Set to zero and solve

Why: σ² > 0, so the bracket must vanish: Σ(dᵢ − μ) = 0 ⇒ Σdᵢ = nμ.

\[ \sum_{i=1}^{n} d_i - n\mu = 0 \;\Longrightarrow\; \hat\mu = \frac{1}{n}\sum_{i=1}^{n} d_i \]

Plug in our sample

Why: Σd = 2+4+4+4+5+5+7+9 = 40, and n = 8, so μ̂ = 40/8 = 5. The MLE mean is exactly the sample average.

\[ \hat\mu = \frac{40}{8} = 5 \]

36. State the rule before it runs: Maximize over σ²

Hypothesis

Predict first

Maximize over σ² is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Differentiate ℓ with respect to σ²

Why: Treat σ² as one variable s. The −(n/2)log s term gives −n/(2s); the last term gives +Σ(dᵢ−μ)²/(2s²).

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

37. Maximize over σ²

Worked example

Differentiate ℓ with respect to σ²

Why: Treat σ² as one variable s. The −(n/2)log s term gives −n/(2s); the last term gives +Σ(dᵢ−μ)²/(2s²).

\[ \frac{\partial \ell}{\partial \sigma^2} = -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}\sum_{i=1}^{n}(d_i-\mu)^2 \]

Set to zero, multiply through by 2σ⁴

Why: 0 = −nσ² + Σ(dᵢ−μ)². Solving for σ² isolates the estimator.

\[ \hat\sigma^2 = \frac{1}{n}\sum_{i=1}^{n}(d_i-\hat\mu)^2 \]

The MLE divides by n, not n−1

Why: Maximum likelihood gives the ÷n estimator (biased low). The ÷(n−1) 'sample variance' is Bessel's correction for unbiasedness — a DIFFERENT estimator. On an MLE question, use ÷n.

38. Predict the next row: Compute σ̂² by hand, then verify

Pattern

Predict first

The table runs: Σ dᵢ / n = μ̂ | 40 / 8 = 5 | 5.0 · Σ(dᵢ−μ̂)² | 32 | 32.0 · σ̂² = Σ(dᵢ−μ̂)²/n | 32 / 8 = 4 | 4.0

In Compute σ̂² by hand, then verify, given the rows so far: what is the next one — the row where quantity is σ̂ = √σ̂²?

Correct: σ̂ = √σ̂² | 2 | 2.0

quantityhandnumpy (verified)
Σ dᵢ / n = μ̂40 / 8 = 55.0
Σ(dᵢ−μ̂)²3232.0
σ̂² = Σ(dᵢ−μ̂)²/n32 / 8 = 44.0
σ̂ = √σ̂²22.0

Why: The relationship between the columns, not the individual numbers, is what generates the next row. d.mean() is ÷n and (( )**2).mean() is also ÷n — exactly the MLE, matching the hand result 40/8 and 32/8.

39. Compute σ̂² by hand, then verify

Worked example

Deviations from μ̂ = 5 are [−3, −1, −1, −1, 0, 0, 2, 4]. Square and sum:

\[ \sum (d_i-5)^2 = 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32 \;\Longrightarrow\; \hat\sigma^2 = \frac{32}{8} = 4 \]

import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu  = d.mean()                     # MLE mean
var = ((d - mu)**2).mean()         # MLE variance (/ n, not n-1)
print("mu =", mu, " var =", var, " sigma =", var**0.5)

Output: mu = 5.0, var = 4.0, sigma = 2.0

Why: d.mean() is ÷n and (( )**2).mean() is also ÷n — exactly the MLE, matching the hand result 40/8 and 32/8.

quantityhandnumpy (verified)
Σ dᵢ / n = μ̂40 / 8 = 55.0
Σ(dᵢ−μ̂)²3232.0
σ̂² = Σ(dᵢ−μ̂)²/n32 / 8 = 44.0
σ̂ = √σ̂²22.0

40. Say it in words: Compute σ̂² by hand, then verify

Translation

\( \sum (d_i-5)^2 = 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32 \;\Longrightarrow\; \hat\sigma^2 = \frac{32}{8} = 4 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

41. Something is wrong here: using n−1 on an MLE question

Anomaly

Predict first

A student writes this, and it looks reasonable:

Variance means the sample variance, so divide the sum of squares by n − 1 = 7.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: That's the unbiased sample variance (Bessel's correction), a different estimator.

The MLE of the variance divides by n = 8.

Why: That's the unbiased sample variance (Bessel's correction), a different estimator. The question asked for the MLE, whose derivative gives exactly ÷n. Mixing them up is the most common MLE slip.

42. Trap: using n−1 on an MLE question

Trap

The trap

Variance means the sample variance, so divide the sum of squares by n − 1 = 7.

σ² = 32 / 7 ≈ 4.571 ✗

Why: That's the unbiased sample variance (Bessel's correction), a different estimator. The question asked for the MLE, whose derivative gives exactly ÷n. Mixing them up is the most common MLE slip.

The fix

The MLE of the variance divides by n = 8.

σ̂² = 32 / 8 = 4 ✓

Why: Setting ∂ℓ/∂σ² = 0 produces Σ(dᵢ−μ̂)²/n with no −1. Reach for ÷(n−1) only when the question explicitly wants an unbiased estimate.

43. Answer it before you see the options: Q2 — the exam question

Prediction

Predict first

For d = [2,4,4,4,5,5,7,9], the maximum-likelihood Gaussian estimates (μ̂, σ̂²) are:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: (5.0, 4.0)

Why: μ̂ = Σd/n = 40/8 = 5. σ̂² = Σ(d−μ̂)²/n = 32/8 = 4. The MLE variance divides by n. Verified: numpy returns (5.0, 4.0).

44. Q2 — the exam question

Check

Derive, then choose.

Check your understanding

For d = [2,4,4,4,5,5,7,9], the maximum-likelihood Gaussian estimates (μ̂, σ̂²) are:

  • A. (5.0, 4.0) (correct)
  • B. (5.0, 4.571)
  • C. (5.0, 2.0)
  • D. (40, 32)

Answer: A

Why: μ̂ = Σd/n = 40/8 = 5. σ̂² = Σ(d−μ̂)²/n = 32/8 = 4. The MLE variance divides by n. Verified: numpy returns (5.0, 4.0).

Why B tempts people
32/7 ≈ 4.571 is the ÷(n−1) unbiased sample variance, not the MLE. The MLE derivative gives ÷n — the exact trap from the previous slide.
Why C tempts people
2.0 is σ̂ (the standard deviation), not σ̂² (the variance). The question asked for the variance: σ̂² = σ̂² = 4, and √4 = 2.
Why D tempts people
40 = Σd (unnormalized) and 32 = Σ(d−μ̂)² (unnormalized). Both still need dividing by n = 8 to become the estimates 5 and 4.

45. Q3 — Optimization

Section

Part 3 of 11 — reading the Hessian

46. Zero gradient isn't enough

Intuition

At a stationary point ∇f = 0 the surface is flat — but flat can mean a valley bottom (minimum), a hilltop (maximum), or a saddle (up one way, down another, like a mountain pass). The gradient alone can't tell them apart.

The Hessian (matrix of second derivatives) measures curvature in every direction. Its eigenvalues are the curvatures along its eigenvector axes — positive = curves up, negative = curves down.

Our exam point has Hessian eigenvalues {+5, −2}. One axis curves up, one curves down. Before computing anything, you can already smell a saddle.

47. The second-derivative test, by eigenvalue sign

Concept

Near a stationary point, f(x) ≈ f₀ + ½(x−x₀)ᵀ H (x−x₀). The quadratic form zᵀHz decides the shape, and its sign is governed entirely by H's eigenvalues.

Hessian eigenvaluesdefinitenesspoint is
all > 0positive definitelocal minimum
all < 0negative definitelocal maximum
mixed signsindefinitesaddle point
some = 0 (rest one sign)semidefinitetest inconclusive

This is the multivariable version of the single-variable rule f'' > 0 ⇒ min. For {+5, −2} the signs are mixed, so H is indefinite.

48. What has to happen first: Why mixed signs mean a saddle

Ranking

Put in order

Put the moves of Why mixed signs mean a saddle into the order they have to happen.

  1. Walk along the first eigenaxis z = [1, 0]
  2. Walk along the second eigenaxis z = [0, 1]
  3. Up one axis, down another ⇒ saddle

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. f ≈ ½·5·1² = +2.5 > 0 — moving this way INCREASES f.

49. Why mixed signs mean a saddle

Worked example

Take the clean representative H = [[5, 0], [0, −2]] (eigenvalues sit on the diagonal for a diagonal matrix). The local quadratic is f ≈ ½(5z₁² − 2z₂²).

Walk along the first eigenaxis z = [1, 0]

Why: f ≈ ½·5·1² = +2.5 > 0 — moving this way INCREASES f. It curves up.

\[ \tfrac12\,[1,0]\,H\,[1,0]^\top = \tfrac12(5) = +2.5 \]

Walk along the second eigenaxis z = [0, 1]

Why: f ≈ ½·(−2)·1² = −1 < 0 — moving this way DECREASES f. It curves down.

\[ \tfrac12\,[0,1]\,H\,[0,1]^\top = \tfrac12(-2) = -1 \]

Up one axis, down another ⇒ saddle

Why: No neighborhood is entirely higher or entirely lower, so it's neither a min nor a max. The negative eigenvalue is literally a downhill escape direction — why saddles stall naive optimizers but not momentum/SGD noise.

50. Where does each piece belong: Lesson 38: Mock Exam — Phase 1 Theory

Sorting

Sort into buckets

These are the pieces of Lesson 38: Mock Exam — Phase 1 Theory, out of order. Put each one back under the part of the lesson it belongs to.

Q1 — Linear Algebra
Why symmetric matrices are special; The eigenvalue equation; Solve the characteristic equation
Q2 — MLE (Probability)
What maximum likelihood asks; Likelihood → log-likelihood; Maximize over μ
Q3 — Optimization
Zero gradient isn't enough; The second-derivative test, by eigenvalue sign; Why mixed signs mean a saddle
s1
Q1 — Linear Algebra is where Lesson 38: Mock Exam — Phase 1 Theory puts Why symmetric matrices are special, The eigenvalue equation, Solve the characteristic equation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Q2 — MLE (Probability) is where Lesson 38: Mock Exam — Phase 1 Theory puts What maximum likelihood asks, Likelihood → log-likelihood, Maximize over μ. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Q3 — Optimization is where Lesson 38: Mock Exam — Phase 1 Theory puts Zero gradient isn't enough, The second-derivative test, by eigenvalue sign, Why mixed signs mean a saddle. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

51. Restore the missing line: Classify from eigenvalues in code

Fill the middle

Fill in the blanks

From Classify from eigenvalues in code — one line has had its right-hand side removed. Put it back.

import numpy as np
H = np.array([[5., 0.], [0., -2.]]) # a stationary point's Hessian
vals = np.linalg.eigvalsh(H)
print("Hessian eigenvalues:", vals) # [-2. 5.]
sign = ("min" if (vals > 0).all() else
"max" if (vals < 0).all() else "saddle")
print("classification:", sign)

Why: vals is what everything below it consumes, so the wrong expression here fails later and somewhere else. eigvalsh sorts ascending. (vals>0).all() is False (−2 fails) and (vals<0).all() is False (5 fails), so the classifier falls through to 'saddle'.

52. Classify from eigenvalues in code

Worked example

A tiny classifier that reads the sign pattern of the Hessian's eigenvalues. Runnable as written:

import numpy as np
H = np.array([[5., 0.], [0., -2.]])   # a stationary point's Hessian
vals = np.linalg.eigvalsh(H)
print("Hessian eigenvalues:", vals)   # [-2.  5.]
sign = ("min" if (vals > 0).all() else
        "max" if (vals < 0).all() else "saddle")
print("classification:", sign)

eigvalsh returns [-2., 5.] → not all one sign → 'saddle'

Why: eigvalsh sorts ascending. (vals>0).all() is False (−2 fails) and (vals<0).all() is False (5 fails), so the classifier falls through to 'saddle'.

stepvalue (verified)
eigvalsh(H)[-2., 5.]
(vals > 0).all()False
(vals < 0).all()False
classificationsaddle

53. Which is which, by value (verified)

Discrimination

Sort into buckets

Sort these by value (verified), from memory, without looking back at Classify from eigenvalues in code. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

[-2., 5.]
eigvalsh(H)
False
(vals > 0).all(); (vals < 0).all()
saddle
classification
g1
value (verified) is "[-2., 5.]" for eigvalsh(H) — that is what the table on "Classify from eigenvalues in code" records, and it is the single property separating this group from the rest.
g2
value (verified) is "False" for (vals > 0).all(), (vals < 0).all() — that is what the table on "Classify from eigenvalues in code" records, and it is the single property separating this group from the rest.
g3
value (verified) is "saddle" for classification — that is what the table on "Classify from eigenvalues in code" records, and it is the single property separating this group from the rest.

54. Q3 — the exam question

Check

Think curvature.

Check your understanding

At a point where ∇f = 0, the Hessian has eigenvalues {+5, −2}. The point is:

  • A. A saddle point (indefinite Hessian) (correct)
  • B. A local minimum
  • C. A local maximum
  • D. A global minimum

Answer: A

Why: Mixed-sign eigenvalues make the Hessian indefinite: the surface curves up along the +5 axis and down along the −2 axis. That's a saddle — verified, our classifier returned 'saddle'.

Why B tempts people
A local minimum needs ALL eigenvalues positive (positive-definite Hessian). The −2 is a genuine downhill direction, so it can't be a min.
Why C tempts people
A local maximum needs all eigenvalues negative. The +5 is an uphill direction, so it can't be a max either.
Why D tempts people
It isn't even a local minimum, let alone global — the negative eigenvalue gives an immediate descent direction away from the point.

55. Q4 — Information Theory

Section

Part 4 of 11 — KL divergence

56. KL divergence: the cost of the wrong code

Intuition

Suppose reality follows distribution P but you build your compression/model around a wrong distribution Q. KL divergence D(P‖Q) is the average number of extra bits you pay for using Q instead of the true P.

Because it's a penalty for being wrong, it can never be negative — and it hits 0 only when Q = P (no penalty for the truth). It's also directional: mistaking P for Q costs differently than mistaking Q for P.

Running example: P = [0.5, 0.5] (a fair coin) and Q = [0.25, 0.75] (a biased guess). We'll compute the penalty both ways and watch it stay positive but change.

57. Definition and the CE = H + KL identity

Concept

KL is the P-weighted average log-ratio. Cross-entropy H(P,Q) splits cleanly into the true entropy plus the KL penalty:

\[ D_{\mathrm{KL}}(P\Vert Q) = \sum_i P_i \log_2\frac{P_i}{Q_i} \;\ge\; 0 \]

\[ \underbrace{H(P,Q)}_{\text{cross-entropy}} = \underbrace{H(P)}_{\text{entropy}} + \underbrace{D_{\mathrm{KL}}(P\Vert Q)}_{\ge\,0} \]

This is why minimizing cross-entropy loss = minimizing KL to the true labels: H(P) is fixed by the data, so all the tunable cost lives in the KL term.

58. Plan first: Compute both directions by hand

Step zero

Discussion prompt

Compute both directions by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: D(P‖Q) with P = [.5,.5], Q = [.25,.75]

Answer:

  1. D(P‖Q) with P = [.5,.5], Q = [.25,.75]
  2. D(Q‖P) — swap the roles
  3. Both positive, but 0.2075 ≠ 0.1887 → asymmetric

59. Compute both directions by hand

Worked example

D(P‖Q) with P = [.5,.5], Q = [.25,.75]

Why: Each term is Pᵢ log₂(Pᵢ/Qᵢ). First: 0.5·log₂(0.5/0.25) = 0.5·log₂2 = 0.5. Second: 0.5·log₂(0.5/0.75) = 0.5·(−0.585) = −0.2925.

\[ D(P\Vert Q) = 0.5\log_2 2 + 0.5\log_2\tfrac{2}{3} = 0.5 - 0.2925 = 0.2075 \]

D(Q‖P) — swap the roles

Why: Now weight by Q: 0.25·log₂(0.25/0.5) + 0.75·log₂(0.75/0.5) = 0.25·(−1) + 0.75·(0.585) = −0.25 + 0.4387 = 0.1887.

\[ D(Q\Vert P) = 0.25\log_2\tfrac12 + 0.75\log_2\tfrac32 = 0.1887 \]

Both positive, but 0.2075 ≠ 0.1887 → asymmetric

Why: Neither is negative (the ≥ 0 property holds), yet they differ — KL is NOT a symmetric distance. That asymmetry is the whole point of the exam question.

60. Restore the missing line: Verify KL in numpy

Fill the middle

Fill in the blanks

From Verify KL in numpy — one line has had its right-hand side removed. Put it back.

import numpy as np
P = np.array([0.5, 0.5])
Q = np.array([0.25, 0.75])
kl = lambda p, q: float(np.sum(p * np.log2(p / q)))
print("D(P||Q) =", round(kl(P, Q), 6))
print("D(Q||P) =", round(kl(Q, P), 6))
print("symmetric?", kl(P, Q) == kl(Q, P))

Why: Q is what everything below it consumes, so the wrong expression here fails later and somewhere else. The numbers match the hand computation to the digit, and the equality test returns False — confirming asymmetry.

61. Verify KL in numpy

Worked example

import numpy as np
P = np.array([0.5, 0.5])
Q = np.array([0.25, 0.75])
kl = lambda p, q: float(np.sum(p * np.log2(p / q)))
print("D(P||Q) =", round(kl(P, Q), 6))
print("D(Q||P) =", round(kl(Q, P), 6))
print("symmetric?", kl(P, Q) == kl(Q, P))

D(P‖Q)=0.207519, D(Q‖P)=0.188722, symmetric? False

Why: The numbers match the hand computation to the digit, and the equality test returns False — confirming asymmetry.

quantityhandnumpy (verified)
D(P‖Q)0.20750.207519
D(Q‖P)0.18870.188722
both ≥ 0yesyes
D(P‖Q) = D(Q‖P)?noFalse

62. Q4 — the exam question

Check

Recall the properties.

Check your understanding

Which statement about the KL divergence D(P‖Q) is correct?

  • A. It is ≥ 0 and generally asymmetric (correct)
  • B. It is a symmetric distance metric
  • C. It can be negative
  • D. It equals H(P) − H(Q)

Answer: A

Why: KL ≥ 0 by Gibbs' inequality (zero only when P = Q) and is generally asymmetric — we computed D(P‖Q) = 0.2075 ≠ 0.1887 = D(Q‖P) on the same pair.

Why B tempts people
KL is not symmetric (our two directions differed) and violates the triangle inequality, so it is a divergence, not a metric.
Why C tempts people
Individual terms Pᵢlog(Pᵢ/Qᵢ) can be negative (our second term was −0.2925), but the full P-weighted SUM is always ≥ 0.
Why D tempts people
That's backwards and mismatched: cross-entropy H(P,Q) = H(P) + KL, so KL = H(P,Q) − H(P), not H(P) − H(Q).

63. Q5 — Evaluation Metrics

Section

Part 5 of 11 — AUC as a ranking

64. AUC is about ranking, not thresholds

Intuition

Accuracy needs you to pick a threshold and then count correct labels. AUC-ROC asks a threshold-free question: if I grab one random positive and one random negative, how often does the model score the positive higher?

That equivalence — area under the ROC curve equals P(score₊ > score₋) — is the Mann–Whitney identity. It means AUC only cares about the ordering of scores, never their absolute values.

Running example: six scored examples, three positive and three negative. We'll count concordant pairs directly and confirm it equals the AUC.

65. Count concordant pairs by hand

Worked example

examplescorelabel
a0.90+
b0.80+
c0.60−
d0.55+
e0.40−
f0.20−

Positives {0.90, 0.80, 0.55} vs negatives {0.60, 0.40, 0.20}: 9 pairs

Why: 3 positives × 3 negatives = 9 ordered comparisons. Count how many the positive wins (scores strictly higher).

0.90 beats all 3, 0.80 beats all 3, 0.55 beats {0.40, 0.20} but loses to 0.60

Why: 3 + 3 + 2 = 8 wins out of 9. The single loss is the positive scored 0.55 sitting below the negative scored 0.60.

\[ \mathrm{AUC} = \frac{\#\{\text{pos} > \text{neg}\}}{n_+ \, n_-} = \frac{8}{9} \approx 0.889 \]

66. Watch it run: Count concordant pairs by hand

Pattern

Step through it

Step through Count concordant pairs by hand one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: example is a
  2. Step 2: example is b
  3. Step 3: example is c
  4. Step 4: example is d
  5. Step 5: example is e
  6. Step 6: example is f

67. Verify against sklearn

Worked example

import numpy as np
scores = np.array([0.9, 0.8, 0.6, 0.55, 0.4, 0.2])
labels = np.array([1,   1,   0,   1,    0,   0])
pos, neg = scores[labels == 1], scores[labels == 0]
wins = sum((p > n) + 0.5*(p == n) for p in pos for n in neg)
print("concordant pairs:", wins, "of", len(pos)*len(neg))
print("AUC =", wins / (len(pos)*len(neg)))

Output: concordant pairs 8.0 of 9, AUC = 0.8889

Why: The pair-counting matches the hand total exactly. Ties would count as half a win — here there are none.

methodAUC (verified)
hand pair-count 8/90.8889
numpy loop above0.8889
sklearn roc_auc_score0.8889

68. Something is wrong here: reading AUC = 0.9 as accuracy

Anomaly

Predict first

A student writes this, and it looks reasonable:

AUC-ROC = 0.9 means the model gets 90% of predictions correct.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING.

AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.

Why: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.

69. Trap: reading AUC = 0.9 as accuracy

Trap

The trap

AUC-ROC = 0.9 means the model gets 90% of predictions correct.

Interpret AUC as accuracy ✗

Why: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.

The fix

AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.

Interpret AUC as a ranking probability ✓

Why: AUC = P(score₊ > score₋). It says nothing about how many labels are right at any single cutoff — only how well the model orders positives above negatives.

70. Break it on purpose: reading AUC = 0.9 as accuracy

Break the constraint

Discussion prompt

The rule this trap just fixed:

AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.

71. Q5 — the exam question

Check

Ranking, not labels.

Check your understanding

A model has AUC-ROC = 0.9. This means:

  • A. A random positive ranks above a random negative 90% of the time (correct)
  • B. The model is 90% accurate
  • C. Precision is 0.9
  • D. 90% of predictions exceed probability 0.9

Answer: A

Why: AUC-ROC equals P(random positive scored above random negative) — the Mann–Whitney identity we counted directly (8/9 concordant pairs = 0.889). It's a threshold-free ranking measure.

Why B tempts people
Accuracy depends on a chosen threshold and counts correct labels; AUC is threshold-free and measures ordering. The two can differ wildly, especially under class imbalance.
Why C tempts people
Precision is a single-threshold quantity (TP / predicted-positive). AUC aggregates ranking across ALL thresholds, so it isn't any one precision value.
Why D tempts people
AUC ignores absolute probability values entirely — only the relative ordering of positive vs negative scores matters, so it says nothing about how many exceed 0.9.

72. Q6 — SVM & Kernels

Section

Part 6 of 11 — Mercer's condition

73. A kernel is a disguised inner product

Intuition

A kernel k(x, z) promises to equal φ(x)·φ(z) — an inner product in some (possibly huge, possibly infinite) feature space φ, without ever building φ. That's the 'kernel trick': all the SVM math only ever touches inner products.

But not every symmetric function is secretly an inner product. Mercer's theorem gives the exact test: k is a valid kernel iff, for every dataset, the Gram matrix G with Gᵢⱼ = k(xᵢ, xⱼ) is positive semidefinite (all eigenvalues ≥ 0).

Why PSD? Because a Gram matrix of real vectors always satisfies zᵀGz = ‖Σzᵢφ(xᵢ)‖² ≥ 0. If some Gram matrix has a negative eigenvalue, no feature map can exist.

74. Guess the shape of the answer: The linear kernel passes; a fake one fails

Estimation

Predict first

Take three points and the linear kernel k(x,z) = xᵀz, whose Gram matrix is XXᵀ. Then contrast with a symmetric matrix that is NOT a valid Gram matrix.

Commit before you compute: what does The linear kernel passes; a fake one fails come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: The 'bad' symmetric matrix has eigenvalues [−1, −1, 2] — a negative → NOT a kernel

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Symmetry is not enough. Two negative eigenvalues mean no feature map φ can reproduce it, so this symmetric function fails Mercer.

75. The linear kernel passes; a fake one fails

Worked example

Take three points and the linear kernel k(x,z) = xᵀz, whose Gram matrix is XXᵀ. Then contrast with a symmetric matrix that is NOT a valid Gram matrix.

import numpy as np
X = np.array([[1., 0.], [0., 1.], [1., 1.]])
G = X @ X.T                         # linear-kernel Gram matrix
print("Gram eigenvalues:", np.linalg.eigvalsh(G).round(4))
print("valid kernel (PSD)?", bool(np.all(np.linalg.eigvalsh(G) >= -1e-9)))
bad = np.array([[0.,1.,1.],[1.,0.,1.],[1.,1.,0.]])
print("bad matrix eigenvalues:", np.linalg.eigvalsh(bad))

Linear Gram eigenvalues [0, 1, 3] — all ≥ 0 → PSD → valid

Why: Every eigenvalue is non-negative, so the Gram matrix is PSD. The linear kernel is genuinely an inner product (it literally IS xᵀz).

The 'bad' symmetric matrix has eigenvalues [−1, −1, 2] — a negative → NOT a kernel

Why: Symmetry is not enough. Two negative eigenvalues mean no feature map φ can reproduce it, so this symmetric function fails Mercer.

matrixeigenvalues (verified)valid kernel?
linear Gram XXᵀ[0, 1, 3]yes — PSD
symmetric 'bad'[−1, −1, 2]no — indefinite

76. Inspect it line by line: The linear kernel passes; a fake one fails

Error analysis

Annotate

Walk the callouts on The linear kernel passes; a fake one fails. Each one is a place this is easy to get subtly wrong.

  • Every eigenvalue is non-negative, so the Gram matrix is PSD. The linear kernel is genuinely an inner product (it literally IS xᵀz).
  • Symmetry is not enough. Two negative eigenvalues mean no feature map φ can reproduce it, so this symmetric function fails Mercer.

77. Something is wrong here: 'symmetric ⇒ valid kernel'

Anomaly

Predict first

A student writes this, and it looks reasonable:

k(x, z) = k(z, x), so it's symmetric — therefore it's a valid kernel.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Symmetry is necessary but NOT sufficient.

Check that every Gram matrix is positive semidefinite, not just symmetric.

Why: Symmetry is necessary but NOT sufficient. Our 'bad' matrix is perfectly symmetric yet has eigenvalues {−1, −1, 2} — a negative one — so no φ produces it and it fails Mercer.

78. Trap: 'symmetric ⇒ valid kernel'

Trap

The trap

k(x, z) = k(z, x), so it's symmetric — therefore it's a valid kernel.

Symmetric ⇒ kernel ✗

Why: Symmetry is necessary but NOT sufficient. Our 'bad' matrix is perfectly symmetric yet has eigenvalues {−1, −1, 2} — a negative one — so no φ produces it and it fails Mercer.

The fix

Check that every Gram matrix is positive semidefinite, not just symmetric.

Symmetric AND every Gram PSD ⇒ kernel ✓

Why: Mercer requires all Gram-matrix eigenvalues ≥ 0. The linear kernel passes (eigenvalues [0,1,3]); a symmetric-but-indefinite function does not. PSD is the real test.

79. How sure are you: Q6 — the exam question

Commit first

Predict first

A function k(x, z) is a valid kernel if and only if:

Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.

Correct: Its Gram matrix is positive semidefinite for every dataset

Why: Mercer's theorem: k corresponds to an inner product φ(x)·φ(z) exactly when every Gram matrix is PSD (all eigenvalues ≥ 0). We saw the linear kernel pass ([0,1,3]) and a symmetric non-kernel fail ([−1,−1,2]).

The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.

80. Q6 — the exam question

Check

Mercer's condition.

Check your understanding

A function k(x, z) is a valid kernel if and only if:

  • A. Its Gram matrix is positive semidefinite for every dataset (correct)
  • B. It is symmetric in x and z
  • C. It is always non-negative
  • D. It equals a Euclidean distance

Answer: A

Why: Mercer's theorem: k corresponds to an inner product φ(x)·φ(z) exactly when every Gram matrix is PSD (all eigenvalues ≥ 0). We saw the linear kernel pass ([0,1,3]) and a symmetric non-kernel fail ([−1,−1,2]).

Why B tempts people
Symmetry is necessary but not sufficient — our 'bad' matrix was symmetric yet indefinite, so it's not a kernel. PSD is the missing requirement.
Why C tempts people
Kernel values can be negative — the linear kernel xᵀz is negative for opposing vectors. Non-negativity is neither required nor sufficient.
Why D tempts people
A distance is not a kernel; the RBF kernel puts a squared distance INSIDE an exponential (e^{−γ‖x−z‖²}) precisely to manufacture a PSD Gram matrix.

81. Q7 — Generalization

Section

Part 7 of 11 — the overfit signature

82. The two ways a model fails

Intuition

A model can fail two opposite ways. Underfitting (high bias): too simple to capture the pattern — it's wrong on both training and test data. Overfitting (high variance): so flexible it memorizes the training noise — near-perfect on train, but it falls apart on unseen test data.

So the diagnostic is the train–test gap. Both errors high ⇒ underfit. Train tiny but test large ⇒ overfit. Both low ⇒ well-tuned.

We'll fit polynomials of degree 1, 3, 9 to noisy sine data and watch degree 9 achieve zero training error while its test error explodes — the overfit signature made visible.

83. Predict the next row: Watch overfitting happen

Pattern

Predict first

The table runs: 1 | 0.2891 | 0.4525 | underfit (high bias) · 3 | 0.0200 | 0.0629 | well-tuned

In Watch overfitting happen, given the rows so far: what is the next one — the row where degree is 9?

Correct: 9 | 0.0000 | 1.6056 | overfit (high variance)

degreetrain MSEtest MSEdiagnosis
10.28910.4525underfit (high bias)
30.02000.0629well-tuned
90.00001.6056overfit (high variance)

Why: The relationship between the columns, not the individual numbers, is what generates the next row. A line can't bend to a sine, so it's wrong everywhere.

84. Watch overfitting happen

Worked example

Twelve noisy points from a sine, split into train (even indices) and test (odd). Fit three polynomial degrees. Seeded, runnable as written:

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
np.random.seed(0)
x = np.linspace(0, 1, 12)
y = np.sin(2*np.pi*x) + np.random.normal(0, 0.2, 12)
xtr, ytr = x[::2].reshape(-1,1), y[::2]
xte, yte = x[1::2].reshape(-1,1), y[1::2]
for deg in (1, 3, 9):
    m = make_pipeline(PolynomialFeatures(deg), LinearRegression()).fit(xtr, ytr)
    tr = np.mean((m.predict(xtr) - ytr)**2)
    te = np.mean((m.predict(xte) - yte)**2)
    print(f"deg={deg}  train={tr:.4f}  test={te:.4f}")

deg=1 underfits: train 0.2891, test 0.4525 — both high

Why: A line can't bend to a sine, so it's wrong everywhere. High bias.

deg=9 overfits: train 0.0000, test 1.6056 — the gap explodes

Why: A degree-9 polynomial threads all 6 training points exactly (zero train error) but oscillates wildly between them, so test error blows up 25× over deg 3. THIS is the low-train/high-test signature.

degreetrain MSEtest MSEdiagnosis
10.28910.4525underfit (high bias)
30.02000.0629well-tuned
90.00001.6056overfit (high variance)

85. Fill in: train MSE for Watch overfitting happen

Comparison

Comparison matrix

From Watch overfitting happen: refill the train MSE column from what you know. The rest of the table is as it appeared.

degreetrain MSEtest MSEdiagnosis
10.28910.4525underfit (high bias)
30.02000.0629well-tuned
90.00001.6056overfit (high variance)

86. Q7 — the exam question

Check

Read the gap.

Check your understanding

A model has very low training error but high test error. This indicates:

  • A. High variance (overfitting) (correct)
  • B. High bias (underfitting)
  • C. High irreducible noise
  • D. The model is well-tuned

Answer: A

Why: Low train / high test error is the classic high-variance signature — the model fits training noise and fails to generalize. Our degree-9 polynomial did exactly this: train 0.0000, test 1.6056.

Why B tempts people
High bias (underfitting) shows up as high error on BOTH train and test — like our degree-1 fit (0.2891 and 0.4525), not a low-train/high-test split.
Why C tempts people
Irreducible noise raises both train and test error together; it does not by itself create a train–test GAP. The gap is what signals variance.
Why D tempts people
A large train–test gap is the opposite of well-tuned. The tuned model here was degree 3, with both errors low and close (0.0200 vs 0.0629).

87. Q8 — The softmax + CE gradient

Section

Part 8 of 11 — the p − y result

88. Why the gradient is just p − y

Intuition

Almost every classifier ends in softmax (logits → probabilities) followed by cross-entropy loss. The magic that makes training cheap: the gradient of the loss with respect to the logits collapses to the beautifully simple p − y — predicted probabilities minus the one-hot truth.

Intuitively: the gradient is the error signal. Where you over-predicted a class, pₖ − yₖ > 0 pushes that logit down; on the true class, p_true − 1 < 0 pushes it up. No probability, no error.

Running example: logits z = [1, 2, 3], true class 2 (the last, 0-indexed). We'll derive p − y and confirm it against PyTorch autograd to six decimals.

89. Teach it back: Why the gradient is just p − y

Explain it

Discussion prompt

Explain Why the gradient is just p − y to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Intuitively: the gradient is the error signal. Where you over-predicted a class, pₖ − yₖ > 0 pushes that logit down; on the true class, p_true − 1 < 0 pushes it up. No probability, no error.

90. Plan first: Softmax the logits

Step zero

Discussion prompt

Softmax the logits — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Softmax: exponentiate then normalize

Answer:

  1. Softmax: exponentiate then normalize
  2. Compute the three probabilities
  3. They sum to 1 — a valid distribution

91. Softmax the logits

Worked example

Softmax: exponentiate then normalize

Why: pₖ = e^{zₖ} / Σⱼ e^{zⱼ}. Subtracting the max (here 3) inside the exponent is a numerical-stability trick that leaves p unchanged.

\[ p_k = \frac{e^{z_k}}{\sum_j e^{z_j}}, \qquad z = [1,2,3] \]

Compute the three probabilities

Why: With z − max = [−2, −1, 0]: e^{−2}=0.1353, e^{−1}=0.3679, e^{0}=1. Sum = 1.5032. Divide each through.

\[ p = [\,0.0900,\; 0.2447,\; 0.6652\,] \]

They sum to 1 — a valid distribution

Why: 0.0900 + 0.2447 + 0.6652 = 0.9999 ≈ 1 (rounding). Softmax always outputs a probability vector.

92. Draw the shape of it: Softmax the logits

Blank canvas

Draw it

Draw what Softmax the logits just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

93. What has to happen first: Subtract the one-hot truth

Ranking

Put in order

Put the moves of Subtract the one-hot truth into the order they have to happen.

  1. One-hot the true class 2
  2. Gradient = p − y, componentwise
  3. Result: [0.0900, 0.2447, −0.3348]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Class 2 (last position) is the truth, so y = [0, 0, 1].

94. Subtract the one-hot truth

Worked example

One-hot the true class 2

Why: Class 2 (last position) is the truth, so y = [0, 0, 1].

\[ y = [\,0,\; 0,\; 1\,] \]

Gradient = p − y, componentwise

Why: The full derivation (softmax Jacobian times the CE derivative) telescopes to exactly p − y. Subtract y from p entry by entry.

\[ \nabla_z L = p - y = [\,0.0900,\; 0.2447,\; 0.6652 - 1\,] \]

Result: [0.0900, 0.2447, −0.3348]

Why: Only the true-class entry is negative (push its logit UP); the two wrong classes are positive (push DOWN). The three components sum to 0, as they must.

\[ \nabla_z L = [\,0.0900,\; 0.2447,\; -0.3348\,] \]

95. Decode the notation: Subtract the one-hot truth

Notation

Annotate

From Subtract the one-hot truth — read this one piece at a time. What is each part doing?

On: \( \nabla_z L = [\,0.0900,\; 0.2447,\; -0.3348\,] \)

  • Class 2 (last position) is the truth, so y = [0, 0, 1].
  • The full derivation (softmax Jacobian times the CE derivative) telescopes to exactly p − y. Subtract y from p entry by entry.
  • Only the true-class entry is negative (push its logit UP); the two wrong classes are positive (push DOWN). The three components sum to 0, as they must.

96. Confirm with PyTorch autograd

Worked example

import torch, numpy as np
z = torch.tensor([1., 2., 3.], requires_grad=True)
loss = torch.nn.functional.cross_entropy(z.unsqueeze(0), torch.tensor([2]))
loss.backward()
print("loss:", round(loss.item(), 6))
print("grad (autograd):", z.grad.numpy().round(6))
p = np.exp([1.,2.,3.]) / np.exp([1.,2.,3.]).sum()
print("p - y (by hand):", (p - np.array([0.,0.,1.])).round(6))

autograd grad = [0.090031, 0.244728, −0.334759]

Why: PyTorch's backward pass through softmax+CE returns exactly p − y, matching the hand vector to six decimals.

sourcegrad component 012
hand p − y0.0900310.244728−0.334759
torch autograd0.0900310.244728−0.334759
sum of components≈ 0

97. Answer it before you see the options: Q8 — the exam question

Prediction

Predict first

For softmax outputs p and one-hot labels y, the gradient of cross-entropy loss with respect to the logits z is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: p − y

Why: The softmax Jacobian times the CE derivative telescopes to ∇_z L = p − y. Verified: z=[1,2,3], class 2 gave [0.0900, 0.2447, −0.3348] from both hand and PyTorch autograd.

98. Q8 — the exam question

Check

The result you must know cold.

Check your understanding

For softmax outputs p and one-hot labels y, the gradient of cross-entropy loss with respect to the logits z is:

  • A. p − y (correct)
  • B. y − p
  • C. p(1 − p)
  • D. −log p

Answer: A

Why: The softmax Jacobian times the CE derivative telescopes to ∇_z L = p − y. Verified: z=[1,2,3], class 2 gave [0.0900, 0.2447, −0.3348] from both hand and PyTorch autograd.

Why B tempts people
y − p is the negation — the direction of ASCENT. Gradient DESCENT uses p − y (or equivalently steps along −(p−y) = y−p); the gradient itself is p − y.
Why C tempts people
p(1−p) is the derivative of a single sigmoid's output, not the logit-gradient of softmax cross-entropy. It confuses the activation's slope with the loss gradient.
Why D tempts people
−log p is (up to indexing) the cross-entropy LOSS value itself, not its gradient. The question asked for the derivative, which is p − y.

99. Q9 — The p-value trap

Section

Part 9 of 11 — the classic reversal

100. The most-tested statistics mistake

Intuition

A p-value is P(data this extreme | H₀ true) — a statement about the data, assuming the null. It is emphatically not P(H₀ true | data) — a statement about the hypothesis. The two conditionals point opposite ways.

Confusing them is the single most common — and most-examined — stats error. p = 0.04 does not mean 'a 4% chance the null is true'. Getting P(H₀ | data) would require a prior and Bayes' rule, which the p-value never uses.

101. What has to be given first: What p actually computes

Missing information

Discussion prompt

Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

This is the tail area under the null distribution — how surprising the data is IF the null holds. It never touches the probability the null itself is true.

102. What p actually computes

Worked example

Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:

from scipy import stats
z = 2.05
p_two = 2 * (1 - stats.norm.cdf(abs(z)))   # two-sided
print("z =", z, " p-value =", round(p_two, 4))
print("this is P(|Z| >= 2.05 | H0), NOT P(H0 true)")

p = 0.0404 = P(|Z| ≥ 2.05 | H₀)

Why: This is the tail area under the null distribution — how surprising the data is IF the null holds. It never touches the probability the null itself is true.

expressionmeaningvalue
P(data extreme | H₀)the p-value0.0404
P(H₀ | data)needs a prior + BayesNOT computed
conclusion from small preject H₀data is surprising under H₀

103. Work backwards from the answer: What p actually computes

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

p = 0.0404 = P(|Z| ≥ 2.05 | H₀)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:

104. Something is wrong here: The trap question (it catches everyone)

Anomaly

Predict first

A student writes this, and it looks reasonable:

A study reports p = 0.04. Conclusion: there's a 4% probability the null hypothesis is true.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This reverses the conditioning.

p = P(data this extreme | H₀ true) — a statement about the data, not the hypothesis.

Why: This reverses the conditioning. The p-value is P(data this extreme | H₀), a statement about the data given the null — not P(H₀ | data). Our example computed p = 0.0404 as a tail area, not a hypothesis probability.

105. The trap question (it catches everyone)

Trap

The trap

A study reports p = 0.04. Conclusion: there's a 4% probability the null hypothesis is true.

Read p as P(H₀ true) ✗

Why: This reverses the conditioning. The p-value is P(data this extreme | H₀), a statement about the data given the null — not P(H₀ | data). Our example computed p = 0.0404 as a tail area, not a hypothesis probability.

The fix

p = P(data this extreme | H₀ true) — a statement about the data, not the hypothesis.

Small p ⇒ data is surprising under H₀ ⇒ reject H₀ ✓

Why: A small p means the observed data would be unlikely if the null were true, so we reject the null. It never gives the probability H₀ is true — that would require a prior and Bayes' rule.

106. Which of these survive contact with Lesson 38: Mock Exam — Phase 1 Theory?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Rule of the room: no notes, no calculator, attempt cold. The struggle is the encoding — revealing early throws the value away.; Our running matrix for this question is A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.; An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.
Breaks
A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.; Variance means the sample variance, so divide the sum of squares by n − 1 = 7.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 38: Mock Exam — Phase 1 Theory puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

107. Rule out three: Q9 — the exam question

Elimination

Eliminate the wrong options

A study reports p = 0.04. Which interpretation is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. If H₀ were true, data this extreme would occur only 4% of the time
  • B. There is a 4% probability the null hypothesis is true
  • C. There is a 96% probability the alternative hypothesis is true
  • D. The effect size is 0.04

Survives elimination: A

Why: The p-value is P(data this extreme | H₀ true) — a tail probability under the null. Our z = 2.05 example gave p = 0.0404 as exactly that tail area, a statement about the data, not the hypothesis.

108. Q9 — the exam question

Check

Mind the direction of the conditional.

Check your understanding

A study reports p = 0.04. Which interpretation is correct?

  • A. If H₀ were true, data this extreme would occur only 4% of the time (correct)
  • B. There is a 4% probability the null hypothesis is true
  • C. There is a 96% probability the alternative hypothesis is true
  • D. The effect size is 0.04

Answer: A

Why: The p-value is P(data this extreme | H₀ true) — a tail probability under the null. Our z = 2.05 example gave p = 0.0404 as exactly that tail area, a statement about the data, not the hypothesis.

Why B tempts people
This reverses the conditional. p is P(data | H₀), not P(H₀ | data); getting the latter requires a prior and Bayes' rule, which the p-value never uses — the exact trap.
Why C tempts people
Same reversal as B plus a complement error: 1 − p is not P(H₁ true). The p-value says nothing about the posterior probability of either hypothesis.
Why D tempts people
The p-value is not an effect size. A tiny effect can give a small p with enough data, and a large effect can give a large p with too little — they are different quantities.

109. Your turn: build a metrics evaluator

Section

Part 10 of 11 — the project

110. Project: precision, recall, F1 from scratch

Concept

Q5 tested metrics conceptually — now implement them. Given true labels and predictions, build the confusion matrix by hand, derive precision / recall / F1, and prove your numbers match sklearn.

#milestonetool
1Count TP, FP, FN, TN with boolean masksnp.sum((a==1)&(b==1))
2Derive precision, recall, F1 from the countsarithmetic
3Verify against sklearnprecision_score, recall_score, f1_score

The fixed dataset for all three milestones — y_true = [1,1,1,1,0,0,0,0,0,0], y_pred = [1,1,1,0,1,1,0,0,0,0]. Build rules: type every line yourself, run after each, and if a mask misbehaves, print the boolean array and look at it.

111. By analogy: Project: precision, recall, F1 from scratch

Analogy

Discussion prompt

Explain Project: precision, recall, F1 from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Q5 tested metrics conceptually — now implement them. Given true labels and predictions, build the confusion matrix by hand, derive precision / recall / F1, and prove your numbers match sklearn.

112. Milestone 1 — the confusion matrix

Worked example

Your turn: count the four cells with boolean masks. Predict each count out loud from the two label arrays before you print.

Hint: a positive prediction on a true positive is (y_true==1) & (y_pred==1) — sum the mask to count. TP is a hit, FP is a false alarm, FN is a miss.

import numpy as np
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
TP = int(np.sum((y_true==1)&(y_pred==1)))
FP = int(np.sum((y_true==0)&(y_pred==1)))
FN = int(np.sum((y_true==1)&(y_pred==0)))
TN = int(np.sum((y_true==0)&(y_pred==0)))
print("TP FP FN TN:", TP, FP, FN, TN)
cellmeaningvalue (verified)
TPpredicted +, truly + (hits)3
FPpredicted +, truly − (false alarms)2
FNpredicted −, truly + (misses)1
TNpredicted −, truly − (correct rejects)4

113. Milestone 2 — precision, recall, F1

Worked example

Your turn: from TP=3, FP=2, FN=1, compute the three metrics. Say aloud which one a spam filter should maximize before printing.

Hint: precision = TP/(TP+FP) (of what you flagged, how much was right). recall = TP/(TP+FN) (of the real positives, how many you caught). F1 is their harmonic mean 2PR/(P+R).

TP, FP, FN, TN = 3, 2, 1, 4
precision = TP / (TP + FP)
recall    = TP / (TP + FN)
f1 = 2 * precision * recall / (precision + recall)
print("precision:", round(precision, 4))
print("recall:   ", round(recall, 4))
print("F1:       ", round(f1, 4))
metricformulavalue (verified)
precision3/(3+2)0.6
recall3/(3+1)0.75
F12·0.6·0.75/(0.6+0.75)0.6667

114. Milestone 3 — verify against sklearn

Worked example

Your turn: feed the same arrays to sklearn and confirm the three numbers match your from-scratch values exactly. Predict whether they'll agree to the last digit.

Hint: import precision_score, recall_score, f1_score from sklearn.metrics and call each on (y_true, y_pred).

import numpy as np
from sklearn.metrics import precision_score, recall_score, f1_score
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
print("precision:", precision_score(y_true, y_pred))
print("recall:   ", recall_score(y_true, y_pred))
print("F1:       ", round(f1_score(y_true, y_pred), 4))
metricyour valuesklearn (verified)
precision0.60.6
recall0.750.75
F10.66670.6667

115. Watch it run: Milestone 3 — verify against sklearn

Pattern

Step through it

Step through Milestone 3 — verify against sklearn one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: metric is precision
  2. Step 2: metric is recall
  3. Step 3: metric is F1

116. Guess the shape of the answer: The full evaluator, assembled

Estimation

Predict first

All three milestones in one program — build the counts, derive the metrics, and check sklearn in the same run:

Commit before you compute: what does The full evaluator, assembled come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Both lines print 0.6 0.75 0.6667

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Your hand-built confusion-matrix metrics and sklearn agree to the last digit — you built classification evaluation from the counts up.

117. The full evaluator, assembled

Worked example

All three milestones in one program — build the counts, derive the metrics, and check sklearn in the same run:

import numpy as np
from sklearn.metrics import precision_score, recall_score, f1_score
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
TP = int(np.sum((y_true==1)&(y_pred==1)))
FP = int(np.sum((y_true==0)&(y_pred==1)))
FN = int(np.sum((y_true==1)&(y_pred==0)))
prec = TP/(TP+FP); rec = TP/(TP+FN)
f1 = 2*prec*rec/(prec+rec)
print("mine:   ", round(prec,4), round(rec,4), round(f1,4))
print("sklearn:", precision_score(y_true,y_pred), recall_score(y_true,y_pred), round(f1_score(y_true,y_pred),4))

Both lines print 0.6 0.75 0.6667

Why: Your hand-built confusion-matrix metrics and sklearn agree to the last digit — you built classification evaluation from the counts up.

printed lineprecisionrecallF1
mine:0.60.750.6667
sklearn:0.60.750.6667

118. What each one costs: The full evaluator, assembled

Trade off

Comparison matrix

From The full evaluator, assembled: every row here is a choice with a cost. Fill the F1 column, then say which row you would actually pick and what you give up for it.

printed lineprecisionrecallF1
mine:0.60.750.6667
sklearn:0.60.750.6667

119. Error analysis & debrief

Section

Part 11 of 11

120. Analyze every miss

Concept

The exam's value is in the debrief. For each wrong answer, write three sentences: the correct reasoning, the exact step where your path went wrong, and the one-line result to memorize.

Add every miss to your retrieval log with a re-quiz date. A concept is only 'done' when you reproduce it cold on a later day — not right after seeing the answer.

121. Break it if you can: Analyze every miss

Counterexample

Discussion prompt

The exam's value is in the debrief. For each wrong answer, write three sentences: the correct reasoning, the exact step where your path went wrong, and the one-line result to memorize.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

122. Rank your gaps for Phase 2

Concept

Rank your misses and pick the top 5 weakest areas for targeted review before Phase 2. Don't carry shaky foundations into classical ML — every classical algorithm rests on the linear algebra, probability, and optimization you just tested.

if you missed…the fastest fix
Q1 eigenre-derive det(A−λI)=0 on three 2×2 matrices
Q2 MLEdifferentiate the Gaussian log-likelihood from scratch
Q3 Hessianclassify 5 random symmetric 2×2 matrices by eigen-sign
Q4/Q8 info+gradre-derive p − y and the CE = H + KL identity
Q5–Q7, Q9one flashcard each: AUC, Mercer, overfit gap, p-value

123. Fill in: the fastest fix for Rank your gaps for Phase 2

Comparison

Comparison matrix

From Rank your gaps for Phase 2: refill the the fastest fix column from what you know. The rest of the table is as it appeared.

if you missed…the fastest fix
Q1 eigenre-derive det(A−λI)=0 on three 2×2 matrices
Q2 MLEdifferentiate the Gaussian log-likelihood from scratch
Q3 Hessianclassify 5 random symmetric 2×2 matrices by eigen-sign
Q4/Q8 info+gradre-derive p − y and the CE = H + KL identity
Q5–Q7, Q9one flashcard each: AUC, Mercer, overfit gap, p-value

124. Connect it up: Lesson 38: Mock Exam — Phase 1 Theory

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Q1 — Linear Algebra · Q2 — MLE (Probability) · Q3 — Optimization · Q4 — Information Theory · Q5 — Evaluation Metrics · Q6 — SVM & Kernels. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

125. What you can do now

Recap

domainthe keystone to know cold
linear algebrasymmetric ⇒ real λ, orthogonal eigenvectors
MLEσ̂² divides by n; μ̂ is the sample mean
optimizationmixed-sign Hessian eigenvalues ⇒ saddle
informationCE = H + KL; KL ≥ 0 and asymmetric
metricsAUC = P(pos ranked above neg)
kernelsvalid ⟺ every Gram matrix is PSD (Mercer)
generalizationlow train / high test ⇒ overfit
deep learningsoftmax+CE logit-gradient = p − y
statisticsp = P(data | H₀), never P(H₀ | data)

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 38 (Week 13 — Theory Mock Exam) — Barron · USAAIO Round 2 Preparation, 2026
  2. Gilbert Strang, Introduction to Linear Algebra, Ch. 6 (Eigenvalues & the Spectral Theorem) — Wellesley-Cambridge Press, 5th ed.
  3. Cover & Thomas, Elements of Information Theory, Ch. 2 (Entropy, Relative Entropy) — Wiley, 2nd ed.
  4. scikit-learn metrics: roc_auc_score, precision/recall/f1
  5. PyTorch cross_entropy / autograd
  6. Every eigenvalue, gradient, AUC, and metric produced by real execution — numpy 2.2.6 + torch 2.7.1 + scikit-learn 1.9.0 + scipy 1.16, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108