USAAIO Lesson 38, from Week 13, fully worked: a timed Phase 1 theory mock exam turned into an olympiad-grade walkthrough. Nine exam questions span linear algebra with symmetric eigendecomposition, MLE, Hessian curvature, KL divergence, AUC, kernels and Mercer's theorem, generalization, the softmax-with-cross-entropy gradient, and the p-value trap. Each is derived one move per beat with nothing skipped, every number was produced by real numpy, torch, and sklearn execution, and there is a from-scratch metrics-evaluator project verified against sklearn. The lesson runs to 69 slides.
Subject: Machine Learning · 125 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 38 · Week 13
Nine exam questions, every one derived one move at a time with nothing skipped, then confirmed in numpy/torch/sklearn. Attempt each cold, watch the full derivation, then log every miss. This is the whole Phase-1 toolkit under one roof.
Objectives
2×2 symmetric matrix by hand, checking trace and orthogonalityμ̂, σ̂² from the log-likelihood and know why σ̂² divides by np − yCE = H + KL identity for KL divergence, and read AUC as a ranking probabilitysklearn exactlyWarm-up
Discussion prompt
Before we open Lesson 38: Mock Exam — Phase 1 Theory: without looking back, what was the main idea of Phase 1 Review & Mock Prep, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
a retrieval-practice consolidation of Phase 1 — the four pillars (linear algebra, probability/information, calculus/optimization, ML foundations), a synthesis capstone (OLS four equivalent ways), exam-pace retrieval drills, and a gap-identification protocol before the first mock exam.
Concept
A real USAAIO theory section mixes multiple choice, short derivations, and explanations under time. This mock mirrors it — but for each question we do the full derivation you'd have to produce cold, then confirm every number by execution.
| for each exam question | what you'll do |
|---|---|
| 1. attempt | derive on paper, slides closed |
| 2. derive | watch it built one move per beat |
| 3. commit | pick an answer on the check slide |
| 4. verify | see the numpy/torch/sklearn output |
| 5. log | record every miss for the debrief |
Rule of the room: no notes, no calculator, attempt cold. The struggle is the encoding — revealing early throws the value away.
Comparison
Comparison matrix
From How this mock is built: refill the what you'll do column from what you know. The rest of the table is as it appeared.
| for each exam question | what you'll do |
|---|---|
| 1. attempt | derive on paper, slides closed |
| 2. derive | watch it built one move per beat |
| 3. commit | pick an answer on the check slide |
| 4. verify | see the numpy/torch/sklearn output |
| 5. log | record every miss for the debrief |
Section
Part 1 of 11 — symmetric eigenpairs
Intuition
A general square matrix can rotate space — its eigenvalues may be complex and its eigenvectors need not be perpendicular. A symmetric matrix (A = Aᵀ) only ever stretches along fixed perpendicular axes. It never rotates.
So a symmetric matrix always has real eigenvalues and a full set of orthogonal eigenvectors — you can pick an axis-aligned frame in which A is pure scaling. That is the spectral theorem, and it powers PCA, covariance, kernels, and Hessians.
Our running matrix for this question is A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.
Counterexample
Discussion prompt
Our running matrix for this question is A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.
\[ A v = \lambda v \;\Longleftrightarrow\; (A - \lambda I)\,v = 0 \]
For a nonzero v to exist, A − λI must be singular — its determinant must be zero. That single condition is the characteristic equation.
\[ \det(A - \lambda I) = 0 \]
Analogy
Discussion prompt
Explain The eigenvalue equation by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.
Ranking
Put in order
Put the moves of Solve the characteristic equation into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Subtract λ from each diagonal entry of [[2,1],[1,2]].
Worked example
Write A − λI for our matrix
Why: Subtract λ from each diagonal entry of [[2,1],[1,2]].
\[ A - \lambda I = \begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} \]
Take the 2×2 determinant ad − bc
Why: det = (2−λ)(2−λ) − (1)(1). The off-diagonal product is 1·1 = 1.
\[ \det(A-\lambda I) = (2-\lambda)^2 - 1 \]
Expand (2−λ)² and set to zero
Why: (2−λ)² = 4 − 4λ + λ², so the equation is λ² − 4λ + 4 − 1 = 0.
\[ \lambda^2 - 4\lambda + 3 = 0 \]
Factor the quadratic
Why: λ² − 4λ + 3 = (λ − 1)(λ − 3), so the roots are λ = 1 and λ = 3 — both real, as the spectral theorem promised.
\[ (\lambda - 1)(\lambda - 3) = 0 \;\Longrightarrow\; \lambda = 1,\; 3 \]
Notation
Annotate
From Solve the characteristic equation — read this one piece at a time. What is each part doing?
On: \( A - \lambda I = \begin{bmatrix} 2-\lambda & 1 \\ 1 & 2-\lambda \end{bmatrix} \)
Estimation
Predict first
Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.
Commit before you compute: what does Sanity-check with trace and determinant come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Determinant check: product of eigenvalues = det
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. det(A) = 2·2 − 1·1 = 3, and 1·3 = 3.
Worked example
Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.
Trace check: sum of eigenvalues = trace
Why: trace(A) = 2 + 2 = 4, and 1 + 3 = 4. ✓
\[ \lambda_1 + \lambda_2 = 1 + 3 = 4 = \operatorname{trace}(A) \]
Determinant check: product of eigenvalues = det
Why: det(A) = 2·2 − 1·1 = 3, and 1·3 = 3. ✓ Both invariants agree, so λ = {1, 3} is correct.
\[ \lambda_1 \lambda_2 = 1 \cdot 3 = 3 = \det(A) \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Determinant check: product of eigenvalues = det
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Two free checks pin eigenvalues before you find eigenvectors: they must sum to the trace and multiply to the determinant.
Step zero
Discussion prompt
Find the eigenvectors, entry by entry — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: For λ = 3: solve (A − 3I)v = 0
Answer:
Worked example
For λ = 3: solve (A − 3I)v = 0
Why: A − 3I = [[−1, 1], [1, −1]]. The first row says −v₁ + v₂ = 0, i.e. v₂ = v₁.
\[ \begin{bmatrix} -1 & 1 \\ 1 & -1 \end{bmatrix}\begin{bmatrix} v_1 \\ v_2 \end{bmatrix} = 0 \;\Longrightarrow\; v = \tfrac{1}{\sqrt2}\begin{bmatrix} 1 \\ 1 \end{bmatrix} \]
For λ = 1: solve (A − I)v = 0
Why: A − I = [[1, 1], [1, 1]]. The row says v₁ + v₂ = 0, i.e. v₂ = −v₁.
\[ \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}\begin{bmatrix} v_1 \\ v_2 \end{bmatrix} = 0 \;\Longrightarrow\; v = \tfrac{1}{\sqrt2}\begin{bmatrix} 1 \\ -1 \end{bmatrix} \]
Check they're orthogonal
Why: [1,1]·[1,−1] = 1 − 1 = 0. Perpendicular eigenvectors — exactly what the spectral theorem guarantees for a symmetric matrix.
Blank canvas
Draw it
Draw what Find the eigenvectors, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Missing information
Discussion prompt
Use eigh (the symmetric solver — it returns sorted real eigenvalues and orthonormal eigenvectors). Runnable as written:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
eigh sorts ascending, so column 0 pairs with λ = 1 and column 1 with λ = 3.
Worked example
Use eigh (the symmetric solver — it returns sorted real eigenvalues and orthonormal eigenvectors). Runnable as written:
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
vals, vecs = np.linalg.eigh(A) # symmetric solver
print("eigenvalues:", vals) # [1. 3.]
print("eigenvector cols:\n", vecs.round(4))
print("orthogonal? v0.v1 =", round(float(vecs[:,0] @ vecs[:,1]), 12))
print("trace", np.trace(A), "= sum eig", vals.sum())eigenvalues come back [1., 3.] — matches the hand factoring
Why: eigh sorts ascending, so column 0 pairs with λ = 1 and column 1 with λ = 3.
| quantity | value (verified) |
|---|---|
| eigenvalues | [1., 3.] |
| eigenvector for λ=1 | [−0.7071, 0.7071] (∝ [1, −1]) |
| eigenvector for λ=3 | [0.7071, 0.7071] (∝ [1, 1]) |
| v0 · v1 (orthogonality) | 0.0 |
| trace = Σ eigenvalues | 4.0 = 4.0 |
Trade off
Comparison matrix
From Confirm in numpy: every row here is a choice with a cost. Fill the value (verified) column, then say which row you would actually pick and what you give up for it.
| quantity | value (verified) |
|---|---|
| eigenvalues | [1., 3.] |
| eigenvector for λ=1 | [−0.7071, 0.7071] (∝ [1, −1]) |
| eigenvector for λ=3 | [0.7071, 0.7071] (∝ [1, 1]) |
| v0 · v1 (orthogonality) | 0.0 |
| trace = Σ eigenvalues | 4.0 = 4.0 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix.
Solve det(A − λI) = 0 — the off-diagonals matter.
Why: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix. Here the off-diagonal 1s couple the axes, so the true spectrum is {1, 3}, not {2, 2}. It even fails the det check: 2·2 = 4 ≠ det = 3.
Trap
A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.
λ = 2, 2 ✗
Why: The diagonal equals the eigenvalues ONLY for a triangular (or diagonal) matrix. Here the off-diagonal 1s couple the axes, so the true spectrum is {1, 3}, not {2, 2}. It even fails the det check: 2·2 = 4 ≠ det = 3.
Solve det(A − λI) = 0 — the off-diagonals matter.
λ = 1, 3 ✓
Why: The characteristic equation λ² − 4λ + 3 = 0 factors to (λ−1)(λ−3). Sum 4 = trace and product 3 = det both check out. Only trust 'diagonal = eigenvalues' when the matrix is triangular.
Ranking
Put in order
These are the steps of The eigenpair recipe (any 2×2), scrambled. Put them back in order before the next slide shows you.
det(A − λI) = 0 → λ² − (trace)λ + (det) = 0Σλ = trace, Πλ = det(A − λI)v = 0 for the direction v, then normalizev_i · v_j = 0np.linalg.eigh for symmetric, np.linalg.eig for general matricesWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
det(A − λI) = 0 → λ² − (trace)λ + (det) = 0Σλ = trace, Πλ = det(A − λI)v = 0 for the direction v, then normalizev_i · v_j = 0np.linalg.eigh for symmetric, np.linalg.eig for general matricesEdge cases
Discussion prompt
The eigenpair recipe (any 2×2) works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
det(A − λI) = 0 → λ² − (trace)λ + (det) = 0Σλ = trace, Πλ = det(A − λI)v = 0 for the direction v, then normalizev_i · v_j = 0np.linalg.eigh for symmetric, np.linalg.eig for general matricesElimination
Eliminate the wrong options
Which property is guaranteed for a real symmetric matrix but NOT for a general square matrix?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The spectral theorem: a real symmetric matrix has real eigenvalues and an orthonormal eigenbasis. We saw it directly — [[2,1],[1,2]] gave real λ = {1,3} with perpendicular eigenvectors. A general matrix can have complex eigenvalues and non-orthogonal eigenvectors.
Check
Attempt cold, then commit.
Check your understanding
Which property is guaranteed for a real symmetric matrix but NOT for a general square matrix?
Answer: A
Why: The spectral theorem: a real symmetric matrix has real eigenvalues and an orthonormal eigenbasis. We saw it directly — [[2,1],[1,2]] gave real λ = {1,3} with perpendicular eigenvectors. A general matrix can have complex eigenvalues and non-orthogonal eigenvectors.
Section
Part 2 of 11 — the Gaussian fit
Intuition
You have data and a family of models (here: Gaussians with unknown mean μ and variance σ²). Maximum likelihood picks the parameters under which the data you actually observed was most probable.
Our sample for this question: d = [2, 4, 4, 4, 5, 5, 7, 9] (eight numbers). We'll find the μ̂, σ̂² that make this exact sample most likely under a normal model.
Spoiler you should predict: the MLE mean is just the sample average, and the MLE variance is the average squared deviation. But deriving it — and seeing the ÷n — is the exam skill.
Explain it
Discussion prompt
Explain What maximum likelihood asks to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Spoiler you should predict: the MLE mean is just the sample average, and the MLE variance is the average squared deviation. But deriving it — and seeing the ÷n — is the exam skill.
Concept
Assuming the n points are independent, the likelihood is the product of Gaussian densities. Products are painful to differentiate, so take the log — it turns the product into a sum and doesn't move the maximum (log is increasing).
\[ \mathcal{L}(\mu,\sigma^2) = \prod_{i=1}^{n} \frac{1}{\sqrt{2\pi\sigma^2}}\exp\!\Big(-\frac{(d_i-\mu)^2}{2\sigma^2}\Big) \]
\[ \ell(\mu,\sigma^2) = -\frac{n}{2}\log(2\pi) - \frac{n}{2}\log\sigma^2 - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(d_i-\mu)^2 \]
Step zero
Discussion prompt
Maximize over μ — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Differentiate ℓ with respect to μ
Answer:
Worked example
Differentiate ℓ with respect to μ
Why: Only the last term depends on μ. Chain rule on (dᵢ − μ)²: derivative is −2(dᵢ − μ), and the −1/(2σ²) out front flips signs.
\[ \frac{\partial \ell}{\partial \mu} = \frac{1}{\sigma^2}\sum_{i=1}^{n}(d_i - \mu) \]
Set to zero and solve
Why: σ² > 0, so the bracket must vanish: Σ(dᵢ − μ) = 0 ⇒ Σdᵢ = nμ.
\[ \sum_{i=1}^{n} d_i - n\mu = 0 \;\Longrightarrow\; \hat\mu = \frac{1}{n}\sum_{i=1}^{n} d_i \]
Plug in our sample
Why: Σd = 2+4+4+4+5+5+7+9 = 40, and n = 8, so μ̂ = 40/8 = 5. The MLE mean is exactly the sample average.
\[ \hat\mu = \frac{40}{8} = 5 \]
Hypothesis
Predict first
Maximize over σ² is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Differentiate ℓ with respect to σ²
Why: Treat σ² as one variable s. The −(n/2)log s term gives −n/(2s); the last term gives +Σ(dᵢ−μ)²/(2s²).
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Differentiate ℓ with respect to σ²
Why: Treat σ² as one variable s. The −(n/2)log s term gives −n/(2s); the last term gives +Σ(dᵢ−μ)²/(2s²).
\[ \frac{\partial \ell}{\partial \sigma^2} = -\frac{n}{2\sigma^2} + \frac{1}{2\sigma^4}\sum_{i=1}^{n}(d_i-\mu)^2 \]
Set to zero, multiply through by 2σ⁴
Why: 0 = −nσ² + Σ(dᵢ−μ)². Solving for σ² isolates the estimator.
\[ \hat\sigma^2 = \frac{1}{n}\sum_{i=1}^{n}(d_i-\hat\mu)^2 \]
The MLE divides by n, not n−1
Why: Maximum likelihood gives the ÷n estimator (biased low). The ÷(n−1) 'sample variance' is Bessel's correction for unbiasedness — a DIFFERENT estimator. On an MLE question, use ÷n.
Pattern
Predict first
The table runs: Σ dᵢ / n = μ̂ | 40 / 8 = 5 | 5.0 · Σ(dᵢ−μ̂)² | 32 | 32.0 · σ̂² = Σ(dᵢ−μ̂)²/n | 32 / 8 = 4 | 4.0
In Compute σ̂² by hand, then verify, given the rows so far: what is the next one — the row where quantity is σ̂ = √σ̂²?
Correct: σ̂ = √σ̂² | 2 | 2.0
| quantity | hand | numpy (verified) |
|---|---|---|
| Σ dᵢ / n = μ̂ | 40 / 8 = 5 | 5.0 |
| Σ(dᵢ−μ̂)² | 32 | 32.0 |
| σ̂² = Σ(dᵢ−μ̂)²/n | 32 / 8 = 4 | 4.0 |
| σ̂ = √σ̂² | 2 | 2.0 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. d.mean() is ÷n and (( )**2).mean() is also ÷n — exactly the MLE, matching the hand result 40/8 and 32/8.
Worked example
Deviations from μ̂ = 5 are [−3, −1, −1, −1, 0, 0, 2, 4]. Square and sum:
\[ \sum (d_i-5)^2 = 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32 \;\Longrightarrow\; \hat\sigma^2 = \frac{32}{8} = 4 \]
import numpy as np
d = np.array([2., 4., 4., 4., 5., 5., 7., 9.])
mu = d.mean() # MLE mean
var = ((d - mu)**2).mean() # MLE variance (/ n, not n-1)
print("mu =", mu, " var =", var, " sigma =", var**0.5)Output: mu = 5.0, var = 4.0, sigma = 2.0
Why: d.mean() is ÷n and (( )**2).mean() is also ÷n — exactly the MLE, matching the hand result 40/8 and 32/8.
| quantity | hand | numpy (verified) |
|---|---|---|
| Σ dᵢ / n = μ̂ | 40 / 8 = 5 | 5.0 |
| Σ(dᵢ−μ̂)² | 32 | 32.0 |
| σ̂² = Σ(dᵢ−μ̂)²/n | 32 / 8 = 4 | 4.0 |
| σ̂ = √σ̂² | 2 | 2.0 |
Translation
\( \sum (d_i-5)^2 = 9 + 1 + 1 + 1 + 0 + 0 + 4 + 16 = 32 \;\Longrightarrow\; \hat\sigma^2 = \frac{32}{8} = 4 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Variance means the sample variance, so divide the sum of squares by n − 1 = 7.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: That's the unbiased sample variance (Bessel's correction), a different estimator.
The MLE of the variance divides by n = 8.
Why: That's the unbiased sample variance (Bessel's correction), a different estimator. The question asked for the MLE, whose derivative gives exactly ÷n. Mixing them up is the most common MLE slip.
Trap
Variance means the sample variance, so divide the sum of squares by n − 1 = 7.
σ² = 32 / 7 ≈ 4.571 ✗
Why: That's the unbiased sample variance (Bessel's correction), a different estimator. The question asked for the MLE, whose derivative gives exactly ÷n. Mixing them up is the most common MLE slip.
The MLE of the variance divides by n = 8.
σ̂² = 32 / 8 = 4 ✓
Why: Setting ∂ℓ/∂σ² = 0 produces Σ(dᵢ−μ̂)²/n with no −1. Reach for ÷(n−1) only when the question explicitly wants an unbiased estimate.
Prediction
Predict first
For d = [2,4,4,4,5,5,7,9], the maximum-likelihood Gaussian estimates (μ̂, σ̂²) are:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: (5.0, 4.0)
Why: μ̂ = Σd/n = 40/8 = 5. σ̂² = Σ(d−μ̂)²/n = 32/8 = 4. The MLE variance divides by n. Verified: numpy returns (5.0, 4.0).
Check
Derive, then choose.
Check your understanding
For d = [2,4,4,4,5,5,7,9], the maximum-likelihood Gaussian estimates (μ̂, σ̂²) are:
Answer: A
Why: μ̂ = Σd/n = 40/8 = 5. σ̂² = Σ(d−μ̂)²/n = 32/8 = 4. The MLE variance divides by n. Verified: numpy returns (5.0, 4.0).
Section
Part 3 of 11 — reading the Hessian
Intuition
At a stationary point ∇f = 0 the surface is flat — but flat can mean a valley bottom (minimum), a hilltop (maximum), or a saddle (up one way, down another, like a mountain pass). The gradient alone can't tell them apart.
The Hessian (matrix of second derivatives) measures curvature in every direction. Its eigenvalues are the curvatures along its eigenvector axes — positive = curves up, negative = curves down.
Our exam point has Hessian eigenvalues {+5, −2}. One axis curves up, one curves down. Before computing anything, you can already smell a saddle.
Concept
Near a stationary point, f(x) ≈ f₀ + ½(x−x₀)ᵀ H (x−x₀). The quadratic form zᵀHz decides the shape, and its sign is governed entirely by H's eigenvalues.
| Hessian eigenvalues | definiteness | point is |
|---|---|---|
| all > 0 | positive definite | local minimum |
| all < 0 | negative definite | local maximum |
| mixed signs | indefinite | saddle point |
| some = 0 (rest one sign) | semidefinite | test inconclusive |
This is the multivariable version of the single-variable rule f'' > 0 ⇒ min. For {+5, −2} the signs are mixed, so H is indefinite.
Ranking
Put in order
Put the moves of Why mixed signs mean a saddle into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. f ≈ ½·5·1² = +2.5 > 0 — moving this way INCREASES f.
Worked example
Take the clean representative H = [[5, 0], [0, −2]] (eigenvalues sit on the diagonal for a diagonal matrix). The local quadratic is f ≈ ½(5z₁² − 2z₂²).
Walk along the first eigenaxis z = [1, 0]
Why: f ≈ ½·5·1² = +2.5 > 0 — moving this way INCREASES f. It curves up.
\[ \tfrac12\,[1,0]\,H\,[1,0]^\top = \tfrac12(5) = +2.5 \]
Walk along the second eigenaxis z = [0, 1]
Why: f ≈ ½·(−2)·1² = −1 < 0 — moving this way DECREASES f. It curves down.
\[ \tfrac12\,[0,1]\,H\,[0,1]^\top = \tfrac12(-2) = -1 \]
Up one axis, down another ⇒ saddle
Why: No neighborhood is entirely higher or entirely lower, so it's neither a min nor a max. The negative eigenvalue is literally a downhill escape direction — why saddles stall naive optimizers but not momentum/SGD noise.
Sorting
Sort into buckets
These are the pieces of Lesson 38: Mock Exam — Phase 1 Theory, out of order. Put each one back under the part of the lesson it belongs to.
Fill the middle
Fill in the blanks
From Classify from eigenvalues in code — one line has had its right-hand side removed. Put it back.
import numpy as np
H = np.array([[5., 0.], [0., -2.]]) # a stationary point's Hessian
vals = np.linalg.eigvalsh(H)
print("Hessian eigenvalues:", vals) # [-2. 5.]
sign = ("min" if (vals > 0).all() else
"max" if (vals < 0).all() else "saddle")
print("classification:", sign)
Why: vals is what everything below it consumes, so the wrong expression here fails later and somewhere else. eigvalsh sorts ascending. (vals>0).all() is False (−2 fails) and (vals<0).all() is False (5 fails), so the classifier falls through to 'saddle'.
Worked example
A tiny classifier that reads the sign pattern of the Hessian's eigenvalues. Runnable as written:
import numpy as np
H = np.array([[5., 0.], [0., -2.]]) # a stationary point's Hessian
vals = np.linalg.eigvalsh(H)
print("Hessian eigenvalues:", vals) # [-2. 5.]
sign = ("min" if (vals > 0).all() else
"max" if (vals < 0).all() else "saddle")
print("classification:", sign)eigvalsh returns [-2., 5.] → not all one sign → 'saddle'
Why: eigvalsh sorts ascending. (vals>0).all() is False (−2 fails) and (vals<0).all() is False (5 fails), so the classifier falls through to 'saddle'.
| step | value (verified) |
|---|---|
| eigvalsh(H) | [-2., 5.] |
| (vals > 0).all() | False |
| (vals < 0).all() | False |
| classification | saddle |
Discrimination
Sort into buckets
Sort these by value (verified), from memory, without looking back at Classify from eigenvalues in code. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Check
Think curvature.
Check your understanding
At a point where ∇f = 0, the Hessian has eigenvalues {+5, −2}. The point is:
Answer: A
Why: Mixed-sign eigenvalues make the Hessian indefinite: the surface curves up along the +5 axis and down along the −2 axis. That's a saddle — verified, our classifier returned 'saddle'.
Section
Part 4 of 11 — KL divergence
Intuition
Suppose reality follows distribution P but you build your compression/model around a wrong distribution Q. KL divergence D(P‖Q) is the average number of extra bits you pay for using Q instead of the true P.
Because it's a penalty for being wrong, it can never be negative — and it hits 0 only when Q = P (no penalty for the truth). It's also directional: mistaking P for Q costs differently than mistaking Q for P.
Running example: P = [0.5, 0.5] (a fair coin) and Q = [0.25, 0.75] (a biased guess). We'll compute the penalty both ways and watch it stay positive but change.
Concept
KL is the P-weighted average log-ratio. Cross-entropy H(P,Q) splits cleanly into the true entropy plus the KL penalty:
\[ D_{\mathrm{KL}}(P\Vert Q) = \sum_i P_i \log_2\frac{P_i}{Q_i} \;\ge\; 0 \]
\[ \underbrace{H(P,Q)}_{\text{cross-entropy}} = \underbrace{H(P)}_{\text{entropy}} + \underbrace{D_{\mathrm{KL}}(P\Vert Q)}_{\ge\,0} \]
This is why minimizing cross-entropy loss = minimizing KL to the true labels: H(P) is fixed by the data, so all the tunable cost lives in the KL term.
Step zero
Discussion prompt
Compute both directions by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: D(P‖Q) with P = [.5,.5], Q = [.25,.75]
Answer:
Worked example
D(P‖Q) with P = [.5,.5], Q = [.25,.75]
Why: Each term is Pᵢ log₂(Pᵢ/Qᵢ). First: 0.5·log₂(0.5/0.25) = 0.5·log₂2 = 0.5. Second: 0.5·log₂(0.5/0.75) = 0.5·(−0.585) = −0.2925.
\[ D(P\Vert Q) = 0.5\log_2 2 + 0.5\log_2\tfrac{2}{3} = 0.5 - 0.2925 = 0.2075 \]
D(Q‖P) — swap the roles
Why: Now weight by Q: 0.25·log₂(0.25/0.5) + 0.75·log₂(0.75/0.5) = 0.25·(−1) + 0.75·(0.585) = −0.25 + 0.4387 = 0.1887.
\[ D(Q\Vert P) = 0.25\log_2\tfrac12 + 0.75\log_2\tfrac32 = 0.1887 \]
Both positive, but 0.2075 ≠ 0.1887 → asymmetric
Why: Neither is negative (the ≥ 0 property holds), yet they differ — KL is NOT a symmetric distance. That asymmetry is the whole point of the exam question.
Fill the middle
Fill in the blanks
From Verify KL in numpy — one line has had its right-hand side removed. Put it back.
import numpy as np
P = np.array([0.5, 0.5])
Q = np.array([0.25, 0.75])
kl = lambda p, q: float(np.sum(p * np.log2(p / q)))
print("D(P||Q) =", round(kl(P, Q), 6))
print("D(Q||P) =", round(kl(Q, P), 6))
print("symmetric?", kl(P, Q) == kl(Q, P))
Why: Q is what everything below it consumes, so the wrong expression here fails later and somewhere else. The numbers match the hand computation to the digit, and the equality test returns False — confirming asymmetry.
Worked example
import numpy as np
P = np.array([0.5, 0.5])
Q = np.array([0.25, 0.75])
kl = lambda p, q: float(np.sum(p * np.log2(p / q)))
print("D(P||Q) =", round(kl(P, Q), 6))
print("D(Q||P) =", round(kl(Q, P), 6))
print("symmetric?", kl(P, Q) == kl(Q, P))D(P‖Q)=0.207519, D(Q‖P)=0.188722, symmetric? False
Why: The numbers match the hand computation to the digit, and the equality test returns False — confirming asymmetry.
| quantity | hand | numpy (verified) |
|---|---|---|
| D(P‖Q) | 0.2075 | 0.207519 |
| D(Q‖P) | 0.1887 | 0.188722 |
| both ≥ 0 | yes | yes |
| D(P‖Q) = D(Q‖P)? | no | False |
Check
Recall the properties.
Check your understanding
Which statement about the KL divergence D(P‖Q) is correct?
Answer: A
Why: KL ≥ 0 by Gibbs' inequality (zero only when P = Q) and is generally asymmetric — we computed D(P‖Q) = 0.2075 ≠ 0.1887 = D(Q‖P) on the same pair.
Section
Part 5 of 11 — AUC as a ranking
Intuition
Accuracy needs you to pick a threshold and then count correct labels. AUC-ROC asks a threshold-free question: if I grab one random positive and one random negative, how often does the model score the positive higher?
That equivalence — area under the ROC curve equals P(score₊ > score₋) — is the Mann–Whitney identity. It means AUC only cares about the ordering of scores, never their absolute values.
Running example: six scored examples, three positive and three negative. We'll count concordant pairs directly and confirm it equals the AUC.
Worked example
| example | score | label |
|---|---|---|
| a | 0.90 | + |
| b | 0.80 | + |
| c | 0.60 | − |
| d | 0.55 | + |
| e | 0.40 | − |
| f | 0.20 | − |
Positives {0.90, 0.80, 0.55} vs negatives {0.60, 0.40, 0.20}: 9 pairs
Why: 3 positives × 3 negatives = 9 ordered comparisons. Count how many the positive wins (scores strictly higher).
0.90 beats all 3, 0.80 beats all 3, 0.55 beats {0.40, 0.20} but loses to 0.60
Why: 3 + 3 + 2 = 8 wins out of 9. The single loss is the positive scored 0.55 sitting below the negative scored 0.60.
\[ \mathrm{AUC} = \frac{\#\{\text{pos} > \text{neg}\}}{n_+ \, n_-} = \frac{8}{9} \approx 0.889 \]
Pattern
Step through it
Step through Count concordant pairs by hand one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
import numpy as np
scores = np.array([0.9, 0.8, 0.6, 0.55, 0.4, 0.2])
labels = np.array([1, 1, 0, 1, 0, 0])
pos, neg = scores[labels == 1], scores[labels == 0]
wins = sum((p > n) + 0.5*(p == n) for p in pos for n in neg)
print("concordant pairs:", wins, "of", len(pos)*len(neg))
print("AUC =", wins / (len(pos)*len(neg)))Output: concordant pairs 8.0 of 9, AUC = 0.8889
Why: The pair-counting matches the hand total exactly. Ties would count as half a win — here there are none.
| method | AUC (verified) |
|---|---|
| hand pair-count 8/9 | 0.8889 |
| numpy loop above | 0.8889 |
| sklearn roc_auc_score | 0.8889 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
AUC-ROC = 0.9 means the model gets 90% of predictions correct.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING.
AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.
Why: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.
Trap
AUC-ROC = 0.9 means the model gets 90% of predictions correct.
Interpret AUC as accuracy ✗
Why: Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.
AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.
Interpret AUC as a ranking probability ✓
Why: AUC = P(score₊ > score₋). It says nothing about how many labels are right at any single cutoff — only how well the model orders positives above negatives.
Break the constraint
Discussion prompt
The rule this trap just fixed:
AUC-ROC = 0.9 means a random positive outranks a random negative 90% of the time.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Accuracy needs a threshold and counts labels; AUC is threshold-free and measures RANKING. A model can have AUC 0.9 yet 50% accuracy at a badly chosen threshold, or on a 99%-negative set.
Check
Ranking, not labels.
Check your understanding
A model has AUC-ROC = 0.9. This means:
Answer: A
Why: AUC-ROC equals P(random positive scored above random negative) — the Mann–Whitney identity we counted directly (8/9 concordant pairs = 0.889). It's a threshold-free ranking measure.
Section
Part 6 of 11 — Mercer's condition
Intuition
A kernel k(x, z) promises to equal φ(x)·φ(z) — an inner product in some (possibly huge, possibly infinite) feature space φ, without ever building φ. That's the 'kernel trick': all the SVM math only ever touches inner products.
But not every symmetric function is secretly an inner product. Mercer's theorem gives the exact test: k is a valid kernel iff, for every dataset, the Gram matrix G with Gᵢⱼ = k(xᵢ, xⱼ) is positive semidefinite (all eigenvalues ≥ 0).
Why PSD? Because a Gram matrix of real vectors always satisfies zᵀGz = ‖Σzᵢφ(xᵢ)‖² ≥ 0. If some Gram matrix has a negative eigenvalue, no feature map can exist.
Estimation
Predict first
Take three points and the linear kernel k(x,z) = xᵀz, whose Gram matrix is XXᵀ. Then contrast with a symmetric matrix that is NOT a valid Gram matrix.
Commit before you compute: what does The linear kernel passes; a fake one fails come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: The 'bad' symmetric matrix has eigenvalues [−1, −1, 2] — a negative → NOT a kernel
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Symmetry is not enough. Two negative eigenvalues mean no feature map φ can reproduce it, so this symmetric function fails Mercer.
Worked example
Take three points and the linear kernel k(x,z) = xᵀz, whose Gram matrix is XXᵀ. Then contrast with a symmetric matrix that is NOT a valid Gram matrix.
import numpy as np
X = np.array([[1., 0.], [0., 1.], [1., 1.]])
G = X @ X.T # linear-kernel Gram matrix
print("Gram eigenvalues:", np.linalg.eigvalsh(G).round(4))
print("valid kernel (PSD)?", bool(np.all(np.linalg.eigvalsh(G) >= -1e-9)))
bad = np.array([[0.,1.,1.],[1.,0.,1.],[1.,1.,0.]])
print("bad matrix eigenvalues:", np.linalg.eigvalsh(bad))Linear Gram eigenvalues [0, 1, 3] — all ≥ 0 → PSD → valid
Why: Every eigenvalue is non-negative, so the Gram matrix is PSD. The linear kernel is genuinely an inner product (it literally IS xᵀz).
The 'bad' symmetric matrix has eigenvalues [−1, −1, 2] — a negative → NOT a kernel
Why: Symmetry is not enough. Two negative eigenvalues mean no feature map φ can reproduce it, so this symmetric function fails Mercer.
| matrix | eigenvalues (verified) | valid kernel? |
|---|---|---|
| linear Gram XXᵀ | [0, 1, 3] | yes — PSD |
| symmetric 'bad' | [−1, −1, 2] | no — indefinite |
Error analysis
Annotate
Walk the callouts on The linear kernel passes; a fake one fails. Each one is a place this is easy to get subtly wrong.
Anomaly
Predict first
A student writes this, and it looks reasonable:
k(x, z) = k(z, x), so it's symmetric — therefore it's a valid kernel.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Symmetry is necessary but NOT sufficient.
Check that every Gram matrix is positive semidefinite, not just symmetric.
Why: Symmetry is necessary but NOT sufficient. Our 'bad' matrix is perfectly symmetric yet has eigenvalues {−1, −1, 2} — a negative one — so no φ produces it and it fails Mercer.
Trap
k(x, z) = k(z, x), so it's symmetric — therefore it's a valid kernel.
Symmetric ⇒ kernel ✗
Why: Symmetry is necessary but NOT sufficient. Our 'bad' matrix is perfectly symmetric yet has eigenvalues {−1, −1, 2} — a negative one — so no φ produces it and it fails Mercer.
Check that every Gram matrix is positive semidefinite, not just symmetric.
Symmetric AND every Gram PSD ⇒ kernel ✓
Why: Mercer requires all Gram-matrix eigenvalues ≥ 0. The linear kernel passes (eigenvalues [0,1,3]); a symmetric-but-indefinite function does not. PSD is the real test.
Commit first
Predict first
A function k(x, z) is a valid kernel if and only if:
Commit to an answer, then rate it — certain, fairly sure, or guessing — and write the rating down before you turn the page.
Correct: Its Gram matrix is positive semidefinite for every dataset
Why: Mercer's theorem: k corresponds to an inner product φ(x)·φ(z) exactly when every Gram matrix is PSD (all eigenvalues ≥ 0). We saw the linear kernel pass ([0,1,3]) and a symmetric non-kernel fail ([−1,−1,2]).
The rating matters as much as the answer: confident-and-wrong is the combination that survives revision, because nothing about it feels like it needs revisiting.
Check
Mercer's condition.
Check your understanding
A function k(x, z) is a valid kernel if and only if:
Answer: A
Why: Mercer's theorem: k corresponds to an inner product φ(x)·φ(z) exactly when every Gram matrix is PSD (all eigenvalues ≥ 0). We saw the linear kernel pass ([0,1,3]) and a symmetric non-kernel fail ([−1,−1,2]).
Section
Part 7 of 11 — the overfit signature
Intuition
A model can fail two opposite ways. Underfitting (high bias): too simple to capture the pattern — it's wrong on both training and test data. Overfitting (high variance): so flexible it memorizes the training noise — near-perfect on train, but it falls apart on unseen test data.
So the diagnostic is the train–test gap. Both errors high ⇒ underfit. Train tiny but test large ⇒ overfit. Both low ⇒ well-tuned.
We'll fit polynomials of degree 1, 3, 9 to noisy sine data and watch degree 9 achieve zero training error while its test error explodes — the overfit signature made visible.
Pattern
Predict first
The table runs: 1 | 0.2891 | 0.4525 | underfit (high bias) · 3 | 0.0200 | 0.0629 | well-tuned
In Watch overfitting happen, given the rows so far: what is the next one — the row where degree is 9?
Correct: 9 | 0.0000 | 1.6056 | overfit (high variance)
| degree | train MSE | test MSE | diagnosis |
|---|---|---|---|
| 1 | 0.2891 | 0.4525 | underfit (high bias) |
| 3 | 0.0200 | 0.0629 | well-tuned |
| 9 | 0.0000 | 1.6056 | overfit (high variance) |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. A line can't bend to a sine, so it's wrong everywhere.
Worked example
Twelve noisy points from a sine, split into train (even indices) and test (odd). Fit three polynomial degrees. Seeded, runnable as written:
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
np.random.seed(0)
x = np.linspace(0, 1, 12)
y = np.sin(2*np.pi*x) + np.random.normal(0, 0.2, 12)
xtr, ytr = x[::2].reshape(-1,1), y[::2]
xte, yte = x[1::2].reshape(-1,1), y[1::2]
for deg in (1, 3, 9):
m = make_pipeline(PolynomialFeatures(deg), LinearRegression()).fit(xtr, ytr)
tr = np.mean((m.predict(xtr) - ytr)**2)
te = np.mean((m.predict(xte) - yte)**2)
print(f"deg={deg} train={tr:.4f} test={te:.4f}")deg=1 underfits: train 0.2891, test 0.4525 — both high
Why: A line can't bend to a sine, so it's wrong everywhere. High bias.
deg=9 overfits: train 0.0000, test 1.6056 — the gap explodes
Why: A degree-9 polynomial threads all 6 training points exactly (zero train error) but oscillates wildly between them, so test error blows up 25× over deg 3. THIS is the low-train/high-test signature.
| degree | train MSE | test MSE | diagnosis |
|---|---|---|---|
| 1 | 0.2891 | 0.4525 | underfit (high bias) |
| 3 | 0.0200 | 0.0629 | well-tuned |
| 9 | 0.0000 | 1.6056 | overfit (high variance) |
Comparison
Comparison matrix
From Watch overfitting happen: refill the train MSE column from what you know. The rest of the table is as it appeared.
| degree | train MSE | test MSE | diagnosis |
|---|---|---|---|
| 1 | 0.2891 | 0.4525 | underfit (high bias) |
| 3 | 0.0200 | 0.0629 | well-tuned |
| 9 | 0.0000 | 1.6056 | overfit (high variance) |
Check
Read the gap.
Check your understanding
A model has very low training error but high test error. This indicates:
Answer: A
Why: Low train / high test error is the classic high-variance signature — the model fits training noise and fails to generalize. Our degree-9 polynomial did exactly this: train 0.0000, test 1.6056.
Section
Part 8 of 11 — the p − y result
Intuition
Almost every classifier ends in softmax (logits → probabilities) followed by cross-entropy loss. The magic that makes training cheap: the gradient of the loss with respect to the logits collapses to the beautifully simple p − y — predicted probabilities minus the one-hot truth.
Intuitively: the gradient is the error signal. Where you over-predicted a class, pₖ − yₖ > 0 pushes that logit down; on the true class, p_true − 1 < 0 pushes it up. No probability, no error.
Running example: logits z = [1, 2, 3], true class 2 (the last, 0-indexed). We'll derive p − y and confirm it against PyTorch autograd to six decimals.
Explain it
Discussion prompt
Explain Why the gradient is just p − y to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Intuitively: the gradient is the error signal. Where you over-predicted a class, pₖ − yₖ > 0 pushes that logit down; on the true class, p_true − 1 < 0 pushes it up. No probability, no error.
Step zero
Discussion prompt
Softmax the logits — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Softmax: exponentiate then normalize
Answer:
Worked example
Softmax: exponentiate then normalize
Why: pₖ = e^{zₖ} / Σⱼ e^{zⱼ}. Subtracting the max (here 3) inside the exponent is a numerical-stability trick that leaves p unchanged.
\[ p_k = \frac{e^{z_k}}{\sum_j e^{z_j}}, \qquad z = [1,2,3] \]
Compute the three probabilities
Why: With z − max = [−2, −1, 0]: e^{−2}=0.1353, e^{−1}=0.3679, e^{0}=1. Sum = 1.5032. Divide each through.
\[ p = [\,0.0900,\; 0.2447,\; 0.6652\,] \]
They sum to 1 — a valid distribution
Why: 0.0900 + 0.2447 + 0.6652 = 0.9999 ≈ 1 (rounding). Softmax always outputs a probability vector.
Blank canvas
Draw it
Draw what Softmax the logits just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Ranking
Put in order
Put the moves of Subtract the one-hot truth into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Class 2 (last position) is the truth, so y = [0, 0, 1].
Worked example
One-hot the true class 2
Why: Class 2 (last position) is the truth, so y = [0, 0, 1].
\[ y = [\,0,\; 0,\; 1\,] \]
Gradient = p − y, componentwise
Why: The full derivation (softmax Jacobian times the CE derivative) telescopes to exactly p − y. Subtract y from p entry by entry.
\[ \nabla_z L = p - y = [\,0.0900,\; 0.2447,\; 0.6652 - 1\,] \]
Result: [0.0900, 0.2447, −0.3348]
Why: Only the true-class entry is negative (push its logit UP); the two wrong classes are positive (push DOWN). The three components sum to 0, as they must.
\[ \nabla_z L = [\,0.0900,\; 0.2447,\; -0.3348\,] \]
Notation
Annotate
From Subtract the one-hot truth — read this one piece at a time. What is each part doing?
On: \( \nabla_z L = [\,0.0900,\; 0.2447,\; -0.3348\,] \)
Worked example
import torch, numpy as np
z = torch.tensor([1., 2., 3.], requires_grad=True)
loss = torch.nn.functional.cross_entropy(z.unsqueeze(0), torch.tensor([2]))
loss.backward()
print("loss:", round(loss.item(), 6))
print("grad (autograd):", z.grad.numpy().round(6))
p = np.exp([1.,2.,3.]) / np.exp([1.,2.,3.]).sum()
print("p - y (by hand):", (p - np.array([0.,0.,1.])).round(6))autograd grad = [0.090031, 0.244728, −0.334759]
Why: PyTorch's backward pass through softmax+CE returns exactly p − y, matching the hand vector to six decimals.
| source | grad component 0 | 1 | 2 |
|---|---|---|---|
| hand p − y | 0.090031 | 0.244728 | −0.334759 |
| torch autograd | 0.090031 | 0.244728 | −0.334759 |
| sum of components | ≈ 0 |
Prediction
Predict first
For softmax outputs p and one-hot labels y, the gradient of cross-entropy loss with respect to the logits z is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: p − y
Why: The softmax Jacobian times the CE derivative telescopes to ∇_z L = p − y. Verified: z=[1,2,3], class 2 gave [0.0900, 0.2447, −0.3348] from both hand and PyTorch autograd.
Check
The result you must know cold.
Check your understanding
For softmax outputs p and one-hot labels y, the gradient of cross-entropy loss with respect to the logits z is:
Answer: A
Why: The softmax Jacobian times the CE derivative telescopes to ∇_z L = p − y. Verified: z=[1,2,3], class 2 gave [0.0900, 0.2447, −0.3348] from both hand and PyTorch autograd.
Section
Part 9 of 11 — the classic reversal
Intuition
A p-value is P(data this extreme | H₀ true) — a statement about the data, assuming the null. It is emphatically not P(H₀ true | data) — a statement about the hypothesis. The two conditionals point opposite ways.
Confusing them is the single most common — and most-examined — stats error. p = 0.04 does not mean 'a 4% chance the null is true'. Getting P(H₀ | data) would require a prior and Bayes' rule, which the p-value never uses.
Missing information
Discussion prompt
Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
This is the tail area under the null distribution — how surprising the data is IF the null holds. It never touches the probability the null itself is true.
Worked example
Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:
from scipy import stats
z = 2.05
p_two = 2 * (1 - stats.norm.cdf(abs(z))) # two-sided
print("z =", z, " p-value =", round(p_two, 4))
print("this is P(|Z| >= 2.05 | H0), NOT P(H0 true)")p = 0.0404 = P(|Z| ≥ 2.05 | H₀)
Why: This is the tail area under the null distribution — how surprising the data is IF the null holds. It never touches the probability the null itself is true.
| expression | meaning | value |
|---|---|---|
| P(data extreme | H₀) | the p-value | 0.0404 |
| P(H₀ | data) | needs a prior + Bayes | NOT computed |
| conclusion from small p | reject H₀ | data is surprising under H₀ |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
p = 0.0404 = P(|Z| ≥ 2.05 | H₀)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Suppose a two-sided z-test gives statistic z = 2.05. The p-value is the probability, under H₀, of a statistic at least that extreme:
Anomaly
Predict first
A student writes this, and it looks reasonable:
A study reports p = 0.04. Conclusion: there's a 4% probability the null hypothesis is true.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This reverses the conditioning.
p = P(data this extreme | H₀ true) — a statement about the data, not the hypothesis.
Why: This reverses the conditioning. The p-value is P(data this extreme | H₀), a statement about the data given the null — not P(H₀ | data). Our example computed p = 0.0404 as a tail area, not a hypothesis probability.
Trap
A study reports p = 0.04. Conclusion: there's a 4% probability the null hypothesis is true.
Read p as P(H₀ true) ✗
Why: This reverses the conditioning. The p-value is P(data this extreme | H₀), a statement about the data given the null — not P(H₀ | data). Our example computed p = 0.0404 as a tail area, not a hypothesis probability.
p = P(data this extreme | H₀ true) — a statement about the data, not the hypothesis.
Small p ⇒ data is surprising under H₀ ⇒ reject H₀ ✓
Why: A small p means the observed data would be unlikely if the null were true, so we reject the null. It never gives the probability H₀ is true — that would require a prior and Bayes' rule.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
A = [[2, 1], [1, 2]] — symmetric, so the theorem applies. Let's find its eigenpairs by hand.; An eigenvector v is a direction A only scales: Av = λv. The scale factor λ is the eigenvalue. Rearrange to (A − λI)v = 0.A = [[2, 1], [1, 2]] — the diagonal is 2, 2, so the eigenvalues are 2 and 2.; Variance means the sample variance, so divide the sum of squares by n − 1 = 7.Elimination
Eliminate the wrong options
A study reports p = 0.04. Which interpretation is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The p-value is P(data this extreme | H₀ true) — a tail probability under the null. Our z = 2.05 example gave p = 0.0404 as exactly that tail area, a statement about the data, not the hypothesis.
Check
Mind the direction of the conditional.
Check your understanding
A study reports p = 0.04. Which interpretation is correct?
Answer: A
Why: The p-value is P(data this extreme | H₀ true) — a tail probability under the null. Our z = 2.05 example gave p = 0.0404 as exactly that tail area, a statement about the data, not the hypothesis.
Section
Part 10 of 11 — the project
Concept
Q5 tested metrics conceptually — now implement them. Given true labels and predictions, build the confusion matrix by hand, derive precision / recall / F1, and prove your numbers match sklearn.
| # | milestone | tool |
|---|---|---|
| 1 | Count TP, FP, FN, TN with boolean masks | np.sum((a==1)&(b==1)) |
| 2 | Derive precision, recall, F1 from the counts | arithmetic |
| 3 | Verify against sklearn | precision_score, recall_score, f1_score |
The fixed dataset for all three milestones — y_true = [1,1,1,1,0,0,0,0,0,0], y_pred = [1,1,1,0,1,1,0,0,0,0]. Build rules: type every line yourself, run after each, and if a mask misbehaves, print the boolean array and look at it.
Analogy
Discussion prompt
Explain Project: precision, recall, F1 from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Q5 tested metrics conceptually — now implement them. Given true labels and predictions, build the confusion matrix by hand, derive precision / recall / F1, and prove your numbers match sklearn.
Worked example
Your turn: count the four cells with boolean masks. Predict each count out loud from the two label arrays before you print.
Hint: a positive prediction on a true positive is (y_true==1) & (y_pred==1) — sum the mask to count. TP is a hit, FP is a false alarm, FN is a miss.
import numpy as np
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
TP = int(np.sum((y_true==1)&(y_pred==1)))
FP = int(np.sum((y_true==0)&(y_pred==1)))
FN = int(np.sum((y_true==1)&(y_pred==0)))
TN = int(np.sum((y_true==0)&(y_pred==0)))
print("TP FP FN TN:", TP, FP, FN, TN)| cell | meaning | value (verified) |
|---|---|---|
| TP | predicted +, truly + (hits) | 3 |
| FP | predicted +, truly − (false alarms) | 2 |
| FN | predicted −, truly + (misses) | 1 |
| TN | predicted −, truly − (correct rejects) | 4 |
Worked example
Your turn: from TP=3, FP=2, FN=1, compute the three metrics. Say aloud which one a spam filter should maximize before printing.
Hint: precision = TP/(TP+FP) (of what you flagged, how much was right). recall = TP/(TP+FN) (of the real positives, how many you caught). F1 is their harmonic mean 2PR/(P+R).
TP, FP, FN, TN = 3, 2, 1, 4
precision = TP / (TP + FP)
recall = TP / (TP + FN)
f1 = 2 * precision * recall / (precision + recall)
print("precision:", round(precision, 4))
print("recall: ", round(recall, 4))
print("F1: ", round(f1, 4))| metric | formula | value (verified) |
|---|---|---|
| precision | 3/(3+2) | 0.6 |
| recall | 3/(3+1) | 0.75 |
| F1 | 2·0.6·0.75/(0.6+0.75) | 0.6667 |
Worked example
Your turn: feed the same arrays to sklearn and confirm the three numbers match your from-scratch values exactly. Predict whether they'll agree to the last digit.
Hint: import precision_score, recall_score, f1_score from sklearn.metrics and call each on (y_true, y_pred).
import numpy as np
from sklearn.metrics import precision_score, recall_score, f1_score
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
print("precision:", precision_score(y_true, y_pred))
print("recall: ", recall_score(y_true, y_pred))
print("F1: ", round(f1_score(y_true, y_pred), 4))| metric | your value | sklearn (verified) |
|---|---|---|
| precision | 0.6 | 0.6 |
| recall | 0.75 | 0.75 |
| F1 | 0.6667 | 0.6667 |
Pattern
Step through it
Step through Milestone 3 — verify against sklearn one row at a time. What is driving the change, and what would the row after the last one be?
Estimation
Predict first
All three milestones in one program — build the counts, derive the metrics, and check sklearn in the same run:
Commit before you compute: what does The full evaluator, assembled come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Both lines print 0.6 0.75 0.6667
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Your hand-built confusion-matrix metrics and sklearn agree to the last digit — you built classification evaluation from the counts up.
Worked example
All three milestones in one program — build the counts, derive the metrics, and check sklearn in the same run:
import numpy as np
from sklearn.metrics import precision_score, recall_score, f1_score
y_true = np.array([1,1,1,1,0,0,0,0,0,0])
y_pred = np.array([1,1,1,0,1,1,0,0,0,0])
TP = int(np.sum((y_true==1)&(y_pred==1)))
FP = int(np.sum((y_true==0)&(y_pred==1)))
FN = int(np.sum((y_true==1)&(y_pred==0)))
prec = TP/(TP+FP); rec = TP/(TP+FN)
f1 = 2*prec*rec/(prec+rec)
print("mine: ", round(prec,4), round(rec,4), round(f1,4))
print("sklearn:", precision_score(y_true,y_pred), recall_score(y_true,y_pred), round(f1_score(y_true,y_pred),4))Both lines print 0.6 0.75 0.6667
Why: Your hand-built confusion-matrix metrics and sklearn agree to the last digit — you built classification evaluation from the counts up.
| printed line | precision | recall | F1 |
|---|---|---|---|
| mine: | 0.6 | 0.75 | 0.6667 |
| sklearn: | 0.6 | 0.75 | 0.6667 |
Trade off
Comparison matrix
From The full evaluator, assembled: every row here is a choice with a cost. Fill the F1 column, then say which row you would actually pick and what you give up for it.
| printed line | precision | recall | F1 |
|---|---|---|---|
| mine: | 0.6 | 0.75 | 0.6667 |
| sklearn: | 0.6 | 0.75 | 0.6667 |
Section
Part 11 of 11
Concept
The exam's value is in the debrief. For each wrong answer, write three sentences: the correct reasoning, the exact step where your path went wrong, and the one-line result to memorize.
Add every miss to your retrieval log with a re-quiz date. A concept is only 'done' when you reproduce it cold on a later day — not right after seeing the answer.
Counterexample
Discussion prompt
The exam's value is in the debrief. For each wrong answer, write three sentences: the correct reasoning, the exact step where your path went wrong, and the one-line result to memorize.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Rank your misses and pick the top 5 weakest areas for targeted review before Phase 2. Don't carry shaky foundations into classical ML — every classical algorithm rests on the linear algebra, probability, and optimization you just tested.
| if you missed… | the fastest fix |
|---|---|
| Q1 eigen | re-derive det(A−λI)=0 on three 2×2 matrices |
| Q2 MLE | differentiate the Gaussian log-likelihood from scratch |
| Q3 Hessian | classify 5 random symmetric 2×2 matrices by eigen-sign |
| Q4/Q8 info+grad | re-derive p − y and the CE = H + KL identity |
| Q5–Q7, Q9 | one flashcard each: AUC, Mercer, overfit gap, p-value |
Comparison
Comparison matrix
From Rank your gaps for Phase 2: refill the the fastest fix column from what you know. The rest of the table is as it appeared.
| if you missed… | the fastest fix |
|---|---|
| Q1 eigen | re-derive det(A−λI)=0 on three 2×2 matrices |
| Q2 MLE | differentiate the Gaussian log-likelihood from scratch |
| Q3 Hessian | classify 5 random symmetric 2×2 matrices by eigen-sign |
| Q4/Q8 info+grad | re-derive p − y and the CE = H + KL identity |
| Q5–Q7, Q9 | one flashcard each: AUC, Mercer, overfit gap, p-value |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Q1 — Linear Algebra · Q2 — MLE (Probability) · Q3 — Optimization · Q4 — Information Theory · Q5 — Evaluation Metrics · Q6 — SVM & Kernels. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
det(A−λI)=0, checking trace and orthogonality — [[2,1],[1,2]] → λ = {1,3}μ̂=5, σ̂²=4 and know the MLE divides by n, not n−1p − yCE = H + KL; read AUC as P(score₊ > score₋) = 8/9sklearn — precision 0.6, recall 0.75, F1 0.667| domain | the keystone to know cold |
|---|---|
| linear algebra | symmetric ⇒ real λ, orthogonal eigenvectors |
| MLE | σ̂² divides by n; μ̂ is the sample mean |
| optimization | mixed-sign Hessian eigenvalues ⇒ saddle |
| information | CE = H + KL; KL ≥ 0 and asymmetric |
| metrics | AUC = P(pos ranked above neg) |
| kernels | valid ⟺ every Gram matrix is PSD (Mercer) |
| generalization | low train / high test ⇒ overfit |
| deep learning | softmax+CE logit-gradient = p − y |
| statistics | p = P(data | H₀), never P(H₀ | data) |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.