USAAIO Lesson 13, from Week 5 on linear algebra, fully worked. It builds the quadratic form xᵀAx entry by entry, defines positive semi-definite and positive definite, and connects both to eigenvalues by completing the square. The running matrix A = [[2,1],[1,2]] is tested five independent ways - by the definition, by the completed square, by its eigenvalues, by Sylvester's minors, and by Cholesky - and contrasted against the singular M = [[4,2],[2,1]] and the indefinite N = [[1,2],[2,1]]. It then proves that XᵀX and the covariance matrix are PSD from ‖Xx‖² ≥ 0, shows ridge regression lifting every eigenvalue, and ends with a from-scratch definiteness checker that is cross-validated. Every snippet runs standalone, and every number comes from real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 114 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 13 · Week 5 (Linear Algebra)
The matrices that behave like positive numbers. We build the quadratic form xᵀAx by hand, complete the square to see why the eigenvalue test works, then test one running matrix five independent ways — and prove every XᵀX and every covariance is PSD.
Objectives
xᵀAx into its scalar polynomial and read the entries of A off itxᵀAx ≥ 0) and PD (> 0 ∀x≠0) and prove the eigenvalue link by completing the squareA = LLᵀXᵀX and the covariance are always PSD from xᵀXᵀXx = ‖Xx‖² ≥ 0, and PD iff X has full column rankis_pd (Cholesky) and classify (eigenvalues) from scratch and prove they always agreeWarm-up
Discussion prompt
Before we open Lesson 13: Positive (Semi)Definite Matrices: without looking back, what was the main idea of MAP, Priors & Regularization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
MAP as MLE plus a log-prior, the Gaussian prior → L2/ridge and Laplace prior → L1/lasso correspondences, why L1 produces sparse solutions, conjugate priors, and dropout as approximate Bayesian inference. Compare ridge vs lasso sparsity from scratch.
Section
Part 1 of 6 — what xᵀAx even is
Concept
One matrix carries the whole lesson. Meet A — symmetric, 2×2, and (we'll prove) positive definite:
\[ A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix} \]
It is symmetric (A = Aᵀ), which matters: definiteness is only defined for symmetric matrices. We'll test this exact A by five different criteria and watch them all agree.
Concept
Only the symmetric part of a matrix survives in the quadratic form. The antisymmetric part contributes exactly zero, so a non-symmetric matrix and its symmetrization give the same xᵀAx.
\[ x^\top A x = x^\top\!\left(\tfrac{A + A^\top}{2}\right)\! x, \qquad \text{so symmetrize first: } A \leftarrow \tfrac{A + A^\top}{2} \]
Because of this, 'definiteness of A' really means definiteness of its symmetric part. We always symmetrize before testing — otherwise the eigenvalues can even come out complex and the whole framework breaks.
Counterexample
Discussion prompt
Only the symmetric part of a matrix survives in the quadratic form. The antisymmetric part contributes exactly zero, so a non-symmetric matrix and its symmetrization give the same xᵀAx.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Take a non-symmetric C = [[2,4],[0,2]]. Its form must equal that of its symmetric part (C + Cᵀ)/2. Confirm on x = [1.3, −0.7]. Self-contained:
Commit before you compute: what does Symmetrization leaves the form unchanged come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Both forms print 0.72; symmetric part is [[2,2],[2,2]]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The off-diagonal 4 and 0 average to 2 and 2.
Worked example
Take a non-symmetric C = [[2,4],[0,2]]. Its form must equal that of its symmetric part (C + Cᵀ)/2. Confirm on x = [1.3, −0.7]. Self-contained:
import numpy as np
C = np.array([[2., 4.], [0., 2.]]) # non-symmetric
x = np.array([1.3, -0.7])
S = (C + C.T) / 2 # symmetric part
print(round(x @ C @ x, 4)) # form of C
print(round(x @ S @ x, 4)) # same!
print(S.tolist()) # [[2, 2], [2, 2]]Both forms print 0.72; symmetric part is [[2,2],[2,2]]
Why: The off-diagonal 4 and 0 average to 2 and 2. The antisymmetric leftover [[0,2],[−2,0]] contributes nothing to xᵀCx, so the two forms coincide. Test S, never the raw C.
| quantity | value |
|---|---|
| x @ C @ x (non-sym) | 0.72 |
| x @ S @ x (sym part) | 0.72 |
| S = (C+Cᵀ)/2 | [[2, 2], [2, 2]] |
| eigvalsh(S) | [0, 4] → PSD |
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at Symmetrization leaves the form unchanged. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
The central object is a single number built from a vector x and the matrix A: sandwich A between xᵀ on the left and x on the right.
\[ f(x) = x^\top A x \quad\in\; \mathbb{R} \]
quadratic form — A pure degree-2 polynomial in the entries of x, with the coefficients packed into a symmetric matrix A. Every term is xᵢxⱼ — no linear part, no constant. It is the matrix generalization of the single-variable a·x².
Analogy
Discussion prompt
Explain The quadratic form xᵀAx by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The central object is a single number built from a vector x and the matrix A: sandwich A between xᵀ on the left and x on the right.
Ranking
Put in order
Put the moves of Expand xᵀAx entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row i of A dotted with x. Row 0 is [2,1], row 1 is [1,2].
Worked example
Take x = [x₁, x₂]. Multiply from the inside out. First Ax, then dot with x:
Inner product first: Ax
Why: Row i of A dotted with x. Row 0 is [2,1], row 1 is [1,2].
\[ A x = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = \begin{bmatrix} 2x_1 + x_2 \\ x_1 + 2x_2 \end{bmatrix} \]
Now xᵀ(Ax): dot x with that vector
Why: xᵀ times a column is x₁·(row 0 result) + x₂·(row 1 result).
\[ x^\top A x = x_1(2x_1 + x_2) + x_2(x_1 + 2x_2) \]
Distribute and collect
Why: x₁x₂ appears twice (from the two off-diagonal 1s) and combines into 2x₁x₂. This is the general rule: the off-diagonal entry aᵢⱼ shows up on BOTH sides, so its coefficient in the form is 2aᵢⱼ.
\[ x^\top A x = 2x_1^2 + 2x_1 x_2 + 2x_2^2 \]
Notation
Annotate
From Expand xᵀAx entry by entry — read this one piece at a time. What is each part doing?
On: \( A x = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = \begin{bmatrix} 2x_1 + x_2 \\ x_1 + 2x_2 \end{bmatrix} \)
Missing information
Discussion prompt
Our hand expansion says xᵀAx = 2x₁² + 2x₁x₂ + 2x₂². Plug in a messy x = [1.3, −0.7] both ways and confirm they match. Runnable as written:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The matrix sandwich and the polynomial are the same function of x. Verified by execution — not memory.
Worked example
Our hand expansion says xᵀAx = 2x₁² + 2x₁x₂ + 2x₂². Plug in a messy x = [1.3, −0.7] both ways and confirm they match. Runnable as written:
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4)) # x^T A x
print(round(2*x[0]**2 + 2*x[0]*x[1] + 2*x[1]**2, 4))Both print 2.54 → the expansion is faithful
Why: The matrix sandwich and the polynomial are the same function of x. Verified by execution — not memory.
| expression | printed value |
|---|---|
| x @ A @ x | 2.54 |
| 2·x₁² + 2·x₁x₂ + 2·x₂² | 2.54 |
| match? | yes — identical |
Comparison
Comparison matrix
From Verify the form numerically: refill the printed value column from what you know. The rest of the table is as it appeared.
| expression | printed value |
|---|---|
| x @ A @ x | 2.54 |
| 2·x₁² + 2·x₁x₂ + 2·x₂² | 2.54 |
| match? | yes — identical |
Concept
The pattern generalizes. For a symmetric A = [[a,b],[b,c]], expanding the sandwich gives a fixed template — the diagonal entries weight the pure squares, the off-diagonal is doubled in the cross term:
\[ \begin{bmatrix} x_1 & x_2 \end{bmatrix}\begin{bmatrix} a & b \\ b & c \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = a\,x_1^2 + 2b\,x_1 x_2 + c\,x_2^2 \]
Go backwards: a form 4x₁² + 6x₁x₂ + 9x₂² comes from which A?
Why: Match coefficients: a = 4, c = 9, and 2b = 6 so b = 3. The matrix is [[4,3],[3,9]]. Always HALVE the cross-term coefficient to get the off-diagonal entry — the single most common slip.
\[ 4x_1^2 + 6x_1 x_2 + 9x_2^2 \;\longleftrightarrow\; \begin{bmatrix} 4 & 3 \\ 3 & 9 \end{bmatrix} \]
Intuition
As x ranges over the plane, f(x) = xᵀAx traces a surface sitting above (or below) that plane, always passing through 0 at the origin.
The shape of that surface is the whole story. If it curves strictly upward in every direction — a genuine bowl with one lowest point — the matrix is positive definite. If it is a bowl that flattens into a trough along some line, it is only semidefinite. If it goes up one way and down another (a saddle), it is indefinite.
So 'classify the matrix' really means 'what does its bowl do?'. The next section makes that precise.
Section
Part 2 of 6 — and why eigenvalues decide it
Concept
A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.
\[ A \succ 0 \;\iff\; x^\top A x > 0 \quad \text{for all } x \neq 0 \]
The bowl strictly rises in every direction away from the origin — a single lowest point at x = 0. PD is the matrix version of a positive number.
Explain it
Discussion prompt
Explain Positive definite (PD) to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.
Concept
Positive semidefinite relaxes the strict > to ≥: the form is never negative, but is allowed to hit zero away from the origin.
\[ A \succeq 0 \;\iff\; x^\top A x \ge 0 \quad \text{for all } x \]
A PSD-but-not-PD matrix has a flat direction — a nonzero x with xᵀAx = 0, a trough in the bowl. Every PD matrix is PSD, but not the reverse: PD ⟹ PSD, one-way.
Intuition
PSD/PD is not a curiosity — it is a recurring gate across the whole course. A loss is convex exactly when its Hessian is PSD (Lesson 11). Least squares has a unique solution exactly when XᵀX is PD (Lesson 7). A covariance matrix is meaningful — a valid Gaussian — only because it is PSD.
Kernel methods demand a PSD Gram matrix (Mercer, Week 10); ridge regularization exists precisely to force a PSD-but-singular matrix to become PD. Learn the test once, reuse it everywhere.
Concept
Testing xᵀAx ≥ 0 for all x sounds impossible — infinitely many x. The rescue is that it reduces to a finite check on the eigenvalues:
\[ A \succeq 0 \iff \text{all } \lambda_i \ge 0, \qquad A \succ 0 \iff \text{all } \lambda_i > 0 \]
This isn't magic — it drops out of the spectral theorem plus one change of variables. We derive it next, no step skipped.
Step zero
Discussion prompt
Derive the eigenvalue test — set up — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Start from the spectral theorem
Answer:
Worked example
Start from the spectral theorem
Why: Every symmetric A factors as A = QΛQᵀ, where Q holds orthonormal eigenvectors (QᵀQ = I) and Λ = diag(λ₁,…,λₙ) holds the eigenvalues. This is Lesson 10.
\[ A = Q \Lambda Q^\top, \qquad Q^\top Q = I \]
Substitute into the quadratic form
Why: Replace A by QΛQᵀ inside xᵀAx.
\[ x^\top A x = x^\top Q \Lambda Q^\top x \]
Change variables: let y = Qᵀx
Why: Then xᵀQ = (Qᵀx)ᵀ = yᵀ. Because Q is orthogonal, y ranges over ALL of ℝⁿ exactly as x does — no direction is lost. So checking all x is the same as checking all y.
\[ x^\top A x = y^\top \Lambda y \]
Blank canvas
Draw it
Draw what Derive the eigenvalue test — set up just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Step zero
Discussion prompt
Derive the eigenvalue test — finish — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Λ is diagonal, so yᵀΛy is a plain weighted sum of squares
Answer:
Worked example
Λ is diagonal, so yᵀΛy is a plain weighted sum of squares
Why: A diagonal matrix in a quadratic form has no cross terms: yᵀΛy = Σ λᵢ yᵢ². Each eigenvalue simply weights the square of one coordinate.
\[ x^\top A x = \sum_{i} \lambda_i\, y_i^2 \]
Read off the sign
Why: Every yᵢ² ≥ 0. If every λᵢ ≥ 0, the sum is ≥ 0 for all y — that is PSD. If some λⱼ < 0, pick y = eⱼ (i.e. x = the j-th eigenvector): the sum is λⱼ < 0, so the form goes negative and A is not PSD.
Conclusion
Why: PSD ⟺ all λ ≥ 0, and PD ⟺ all λ > 0 (strict, because a zero eigenvalue makes yᵀΛy = 0 along that eigenvector — a flat direction). The infinite test collapsed to n sign checks.
\[ \boxed{\,A \succeq 0 \iff \lambda_i \ge 0\ \forall i, \qquad A \succ 0 \iff \lambda_i > 0\ \forall i\,} \]
Pattern
Predict first
The table runs: np.linalg.eigvalsh(A) | [1. 3.] | all > 0 → PD · np.linalg.det(A) | 2.9999999999999996 | ≈ 3 = λ₁·λ₂ (float noise)
In Test A by eigenvalues, given the rows so far: what is the next one — the row where quantity is np.trace(A)?
Correct: np.trace(A) | 4.0 | = λ₁ + λ₂
| quantity | printed value | reads as |
|---|---|---|
| np.linalg.eigvalsh(A) | [1. 3.] | all > 0 → PD |
| np.linalg.det(A) | 2.9999999999999996 | ≈ 3 = λ₁·λ₂ (float noise) |
| np.trace(A) | 4.0 | = λ₁ + λ₂ |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Both strictly positive, so by the test A ≻ 0.
Worked example
Apply the freshly derived test to our running A. Use eigvalsh — the symmetric-matrix eigenvalue routine, which returns real eigenvalues sorted ascending. Runnable:
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
print(np.linalg.det(A)) # 3.0 (up to float error)
print(np.linalg.eigvalsh(A)) # [1. 3.]
print(np.trace(A)) # 4.0Eigenvalues [1, 3] — both > 0 → A is PD
Why: Both strictly positive, so by the test A ≻ 0. Sanity: their product 1·3 = 3 = det(A), and their sum 1+3 = 4 = trace(A). Determinant is the product of eigenvalues; trace is their sum.
| quantity | printed value | reads as |
|---|---|---|
| np.linalg.eigvalsh(A) | [1. 3.] | all > 0 → PD |
| np.linalg.det(A) | 2.9999999999999996 | ≈ 3 = λ₁·λ₂ (float noise) |
| np.trace(A) | 4.0 | = λ₁ + λ₂ |
Trade off
Comparison matrix
From Test A by eigenvalues: every row here is a choice with a cost. Fill the reads as column, then say which row you would actually pick and what you give up for it.
| quantity | printed value | reads as |
|---|---|---|
| np.linalg.eigvalsh(A) | [1. 3.] | all > 0 → PD |
| np.linalg.det(A) | 2.9999999999999996 | ≈ 3 = λ₁·λ₂ (float noise) |
| np.trace(A) | 4.0 | = λ₁ + λ₂ |
Hypothesis
Predict first
Prove A ≻ 0 by hand: complete the square is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Pull the x₁ terms together
Why: Group the two terms containing x₁: 2x₁² + 2x₁x₂ = 2(x₁² + x₁x₂). We'll complete the square inside.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
Eigenvalues via a library is fine, but the exam wants you to show it. Take the form 2x₁² + 2x₁x₂ + 2x₂² and rewrite it as a sum of squares — then non-negativity is visible.
Pull the x₁ terms together
Why: Group the two terms containing x₁: 2x₁² + 2x₁x₂ = 2(x₁² + x₁x₂). We'll complete the square inside.
\[ 2x_1^2 + 2x_1 x_2 + 2x_2^2 = 2\left(x_1^2 + x_1 x_2\right) + 2x_2^2 \]
Complete the square inside the parentheses
Why: x₁² + x₁x₂ = (x₁ + x₂/2)² − x₂²/4. Multiply back by 2: 2(x₁ + x₂/2)² − x₂²/2.
\[ = 2\left(x_1 + \tfrac{x_2}{2}\right)^2 - \tfrac{1}{2}x_2^2 + 2x_2^2 \]
Combine the leftover x₂² terms
Why: −½x₂² + 2x₂² = 3/2·x₂². Both coefficients (2 and 3/2) are POSITIVE, so the whole form is a sum of positive multiples of squares — it is ≥ 0, and = 0 only when both squares vanish (x = 0). That is exactly A ≻ 0.
\[ \boxed{\,x^\top A x = 2\left(x_1 + \tfrac{x_2}{2}\right)^2 + \tfrac{3}{2}x_2^2 \;>\; 0\ \ (x \neq 0)\,} \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Combine the leftover x₂² terms
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Eigenvalues via a library is fine, but the exam wants you to show it. Take the form 2x₁² + 2x₁x₂ + 2x₂² and rewrite it as a sum of squares — then non-negativity is visible.
Fill the middle
Fill in the blanks
From Check the completed square numerically — one line has had its right-hand side removed. Put it back.
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4)) # matrix form
print(round(2(x[0] + x[1]/2)2 + 1.5x[1]**2, 4)) # completed square
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Since both weights (2 and 1.5) are positive, the form can never be negative — an algebraic proof of PD that needs no eigenvalue solver.
Worked example
The completed-square form 2(x₁ + x₂/2)² + 1.5x₂² must equal xᵀAx for every x. Verify on the same x = [1.3, −0.7]. Runnable and self-contained:
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4)) # matrix form
print(round(2*(x[0] + x[1]/2)**2 + 1.5*x[1]**2, 4)) # completed squareBoth print 2.54 → the sum-of-squares form is exact
Why: Since both weights (2 and 1.5) are positive, the form can never be negative — an algebraic proof of PD that needs no eigenvalue solver.
| form | value at x=[1.3,−0.7] |
|---|---|
| x @ A @ x | 2.54 |
| 2(x₁+x₂/2)² + 1.5x₂² | 2.54 |
| weights (2, 1.5) | both > 0 → PD |
Anomaly
Predict first
A student writes this, and it looks reasonable:
The determinant is the product of eigenvalues, so det > 0 must mean all eigenvalues are positive — hence PD.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Two NEGATIVE eigenvalues also multiply to a positive determinant.
Check every leading principal minor (Sylvester) or all the eigenvalues — never just the full determinant.
Why: Two NEGATIVE eigenvalues also multiply to a positive determinant. det > 0 fixes only the sign of the PRODUCT, not each factor. This matrix is negative definite — its bowl points straight down.
Trap
The determinant is the product of eigenvalues, so det > 0 must mean all eigenvalues are positive — hence PD.
\[ \det\!\begin{bmatrix} -1 & 0 \\ 0 & -1 \end{bmatrix} = (-1)(-1) = 1 > 0 \]
Conclude [[−1,0],[0,−1]] is PD because det = 1 > 0 ✗
Why: Two NEGATIVE eigenvalues also multiply to a positive determinant. det > 0 fixes only the sign of the PRODUCT, not each factor. This matrix is negative definite — its bowl points straight down.
Check every leading principal minor (Sylvester) or all the eigenvalues — never just the full determinant.
\[ \text{1st leading minor} = -1 < 0 \]
First minor is −1 < 0 → NOT PD (it is negative definite) ✓
Why: Sylvester's criterion needs the 1×1, 2×2, … leading minors ALL strictly positive. Here the very first one fails. Both eigenvalues are −1, confirming it.
Translation
\( \text{1st leading minor} = -1 < 0 \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Section
Part 3 of 6 — definition, square, eigen, Sylvester, Cholesky
Concept
For a symmetric A, these five checks all decide definiteness — and all agree. Different exam questions reward different ones, so keep all five in hand:
| # | test | PSD condition | PD condition |
|---|---|---|---|
| 1 | Definition | xᵀAx ≥ 0 ∀x | xᵀAx > 0 ∀x≠0 |
| 2 | Completed square | sum of squares, coeffs ≥ 0 | coeffs all > 0 |
| 3 | Eigenvalues | all λ ≥ 0 | all λ > 0 |
| 4 | Sylvester (leading minors) | — (need ALL minors) | all leading minors > 0 |
| 5 | Cholesky A = LLᵀ | exists (diag ≥ 0) | unique, diag > 0 |
Caution on Sylvester: the leading-minor test only certifies PD. For PSD you must check all principal minors (not just the leading ones), so for PSD the eigenvalue test is cleanest. We'll hit that trap explicitly.
Ranking
Put in order
Put the moves of Test 4 — Sylvester's leading minors on A into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. det = 2·2 − 1·1 = 4 − 1 = 3 > 0.
Worked example
The leading principal minors are the determinants of the top-left 1×1, 2×2, … blocks. For our A = [[2,1],[1,2]] there are two:
First leading minor: top-left 1×1 = 2
Why: det[2] = 2 > 0. The first check passes.
\[ D_1 = \det\begin{bmatrix} 2 \end{bmatrix} = 2 > 0 \]
Second leading minor: full 2×2 determinant = 3
Why: det = 2·2 − 1·1 = 4 − 1 = 3 > 0. The second check passes too.
\[ D_2 = \det\begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix} = 4 - 1 = 3 > 0 \]
Both leading minors > 0 → A is PD
Why: Sylvester: all leading minors strictly positive certifies PD. This is the fastest by-hand test for a small matrix — no eigenvalues needed.
| leading minor | value | > 0? |
|---|---|---|
| D₁ (1×1) | 2 | yes |
| D₂ (2×2 det) | 3 | yes |
| verdict | PD | ✓ |
Pattern
Step through it
Step through Test 4 — Sylvester's leading minors on A one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Cholesky factors a PD matrix as A = LLᵀ, where L is lower-triangular with positive diagonal. It is the matrix square root that exists exactly when A ≻ 0.
\[ A = L L^\top, \qquad L = \begin{bmatrix} \ell_{11} & 0 \\ \ell_{21} & \ell_{22} \end{bmatrix},\ \ \ell_{11}, \ell_{22} > 0 \]
This is the computer's favorite definiteness test: np.linalg.cholesky(A) succeeds iff A is PD and raises LinAlgError otherwise — a fast, numerically stable yes/no with no eigenvalue solve.
Sorting
Sort into buckets
These are the pieces of Lesson 13: Positive (Semi)Definite Matrices, out of order. Put each one back under the part of the lesson it belongs to.
Step zero
Discussion prompt
Compute Cholesky of A by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Top-left: ℓ₁₁² = 2 → ℓ₁₁ = √2
Answer:
Worked example
Match LLᵀ = A entry by entry. With L = [[ℓ₁₁,0],[ℓ₂₁,ℓ₂₂]], multiply out and solve:
Top-left: ℓ₁₁² = 2 → ℓ₁₁ = √2
Why: (LLᵀ)₁₁ = ℓ₁₁². Set equal to A₁₁ = 2 and take the positive root: ℓ₁₁ = √2 ≈ 1.4142.
\[ \ell_{11}^2 = 2 \;\Rightarrow\; \ell_{11} = \sqrt{2} \]
Off-diagonal: ℓ₂₁·ℓ₁₁ = 1 → ℓ₂₁ = 1/√2
Why: (LLᵀ)₂₁ = ℓ₂₁ℓ₁₁. Set equal to A₂₁ = 1, so ℓ₂₁ = 1/√2 ≈ 0.7071.
\[ \ell_{21}\,\ell_{11} = 1 \;\Rightarrow\; \ell_{21} = \tfrac{1}{\sqrt{2}} \]
Bottom-right: ℓ₂₁² + ℓ₂₂² = 2 → ℓ₂₂ = √(3/2)
Why: (LLᵀ)₂₂ = ℓ₂₁² + ℓ₂₂² = 2. Since ℓ₂₁² = 1/2, we need ℓ₂₂² = 3/2, so ℓ₂₂ = √1.5 ≈ 1.2247. Every diagonal entry is a real positive number → the factorization exists → A is PD.
\[ L = \begin{bmatrix} \sqrt{2} & 0 \\ \tfrac{1}{\sqrt2} & \sqrt{\tfrac{3}{2}} \end{bmatrix} \approx \begin{bmatrix} 1.4142 & 0 \\ 0.7071 & 1.2247 \end{bmatrix} \]
Faded example
Fill in the blanks
Confirm Cholesky in NumPy, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
L = np.linalg.cholesky(A)
print(L.round(6))
print((L @ L.T).round(6)) # rebuilds A
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The solver returns [[1.414214, 0],[0.707107, 1.224745]] — our √2, 1/√2, √1.5 — and multiplying back reconstructs [[2,1],[1,2]].
Worked example
np.linalg.cholesky should return exactly our hand-computed L, and L @ L.T should rebuild A. Self-contained:
import numpy as np
A = np.array([[2., 1.], [1., 2.]])
L = np.linalg.cholesky(A)
print(L.round(6))
print((L @ L.T).round(6)) # rebuilds AL matches our hand answer; LLᵀ = A exactly
Why: The solver returns [[1.414214, 0],[0.707107, 1.224745]] — our √2, 1/√2, √1.5 — and multiplying back reconstructs [[2,1],[1,2]]. Cholesky succeeded, so A is certified PD.
| output | value |
|---|---|
| L row 0 | [1.414214, 0.0] |
| L row 1 | [0.707107, 1.224745] |
| L @ L.T | [[2, 1], [1, 2]] = A |
Concept
Five independent roads, one destination — A = [[2,1],[1,2]] is positive definite by every test:
| test | evidence | verdict |
|---|---|---|
| 1. Definition | xᵀAx = 2.54 > 0 at test x | PD |
| 2. Completed square | 2(x₁+x₂/2)² + 1.5x₂², coeffs > 0 | PD |
| 3. Eigenvalues | [1, 3], both > 0 | PD |
| 4. Sylvester minors | D₁=2>0, D₂=3>0 | PD |
| 5. Cholesky | L exists, positive diagonal | PD |
That redundancy is the point: if two tests disagree, you made an arithmetic error — the tests themselves never do.
Section
Part 4 of 6 — the contrast cases
Intuition
Every symmetric 2×2 produces one of three surface shapes, and the eigenvalue signs name them directly:
| shape of xᵀAx | eigenvalue signs | verdict | example |
|---|---|---|---|
| bowl — up in every direction | both > 0 | PD | [[2,1],[1,2]] |
| trough — flat along a line | one = 0 | PSD only | [[4,2],[2,1]] |
| saddle — up one way, down another | one < 0 | indefinite | [[1,2],[2,1]] |
PD is a single lowest point; PSD-only trades that point for a whole flat line at the bottom; indefinite has no bottom at all. The next three slides make each concrete.
Concept
Now a matrix that is PSD but not PD. It looks similar to A but hides a zero eigenvalue:
\[ M = \begin{bmatrix} 4 & 2 \\ 2 & 1 \end{bmatrix} \]
Its quadratic form is a perfect square, which is exactly why it can touch zero. Watch the completed square collapse to a single term.
Step zero
Discussion prompt
M's form is a perfect square — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Expand xᵀMx
Answer:
Worked example
Expand xᵀMx
Why: Off-diagonal 2 contributes 2·2 = 4 to the cross term. So the form is 4x₁² + 4x₁x₂ + x₂².
\[ x^\top M x = 4x_1^2 + 4x_1 x_2 + x_2^2 \]
Recognize the perfect square
Why: 4x₁² + 4x₁x₂ + x₂² = (2x₁ + x₂)². Only ONE square, no leftover second term — that missing second square is the zero eigenvalue.
\[ x^\top M x = (2x_1 + x_2)^2 \ge 0 \]
Find the flat direction
Why: (2x₁ + x₂)² = 0 whenever 2x₁ + x₂ = 0, e.g. x = [1, −2]. There the form is zero for a NONZERO x — so M is PSD but not PD. That line is the trough in the bowl.
\[ x = \begin{bmatrix} 1 \\ -2 \end{bmatrix} \;\Rightarrow\; x^\top M x = (2 - 2)^2 = 0 \]
Blank canvas
Draw it
Draw what M's form is a perfect square just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From Verify M is PSD-only in NumPy — one line has had its right-hand side removed. Put it back.
import numpy as np
M = np.array([[4., 2.], [2., 1.]])
print(np.linalg.det(M)) # 0.0
print(np.linalg.eigvalsh(M)) # [0. 5.]
x = np.array([1., -2.])
print(round(x @ M @ x, 6)) # 0.0 (flat direction)
Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. The zero eigenvalue makes M singular (det = 0) and gives the flat direction.
Worked example
Determinant 0, eigenvalues [0, 5], and the form vanishing at [1, −2] — all three confirm PSD-not-PD. Self-contained:
import numpy as np
M = np.array([[4., 2.], [2., 1.]])
print(np.linalg.det(M)) # 0.0
print(np.linalg.eigvalsh(M)) # [0. 5.]
x = np.array([1., -2.])
print(round(x @ M @ x, 6)) # 0.0 (flat direction)det 0, eigenvalues [0, 5], form 0 at [1,−2]
Why: The zero eigenvalue makes M singular (det = 0) and gives the flat direction. All eigenvalues ≥ 0 → PSD; the zero → not PD, not invertible.
| test | result | implies |
|---|---|---|
| np.linalg.eigvalsh(M) | [0, 5] | PSD (≥0), not PD |
| np.linalg.det(M) | 0.0 | singular → not PD |
| xᵀMx at [1,−2] | 0.0 | flat direction on a line |
Estimation
Predict first
Because Cholesky needs a positive diagonal, it raises on a merely-PSD matrix. Catch the error instead of crashing. Runnable:
Commit before you compute: what does Cholesky refuses M come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Prints: not PD: Matrix is not positive definite
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The final pivot ℓ₂₂² would be 1 − (2/2)² = 0, not positive, so no real positive diagonal exists.
Worked example
Because Cholesky needs a positive diagonal, it raises on a merely-PSD matrix. Catch the error instead of crashing. Runnable:
import numpy as np
M = np.array([[4., 2.], [2., 1.]])
try:
np.linalg.cholesky(M)
print("PD")
except np.linalg.LinAlgError as e:
print("not PD:", e)Prints: not PD: Matrix is not positive definite
Why: The final pivot ℓ₂₂² would be 1 − (2/2)² = 0, not positive, so no real positive diagonal exists. cholesky raises LinAlgError — this IS the PD test, so a raise means 'not PD'.
| matrix | cholesky | verdict |
|---|---|---|
| A = [[2,1],[1,2]] | succeeds | PD |
| M = [[4,2],[2,1]] | LinAlgError | not PD (only PSD) |
| error message | "Matrix is not positive definite" |
Comparison
Comparison matrix
From Cholesky refuses M: refill the cholesky column from what you know. The rest of the table is as it appeared.
| matrix | cholesky | verdict |
|---|---|---|
| A = [[2,1],[1,2]] | succeeds | PD |
| M = [[4,2],[2,1]] | LinAlgError | not PD (only PSD) |
| error message | "Matrix is not positive definite" |
Worked example
One more contrast — a matrix whose bowl goes up one way and down another. Its form has a genuine negative direction:
import numpy as np
N = np.array([[1., 2.], [2., 1.]])
print(np.linalg.eigvalsh(N)) # [-1. 3.]
print(round(np.array([1., -1.]) @ N @ np.array([1., -1.]), 4)) # -2
print(round(np.array([1., 1.]) @ N @ np.array([1., 1.]), 4)) # 6Eigenvalues [−1, 3] → one negative → indefinite
Why: xᵀNx = x₁² + 4x₁x₂ + x₂². At [1,−1] it is 1 − 4 + 1 = −2 (negative direction); at [1,1] it is 1 + 4 + 1 = 6 (positive direction). Up AND down → neither PSD nor negative definite.
| direction x | xᵀNx | sign |
|---|---|---|
| [1, −1] | −2 | negative (down) |
| [1, 1] | 6 | positive (up) |
| eigenvalues | [−1, 3] | mixed → indefinite |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: But xᵀBx = −x₂² < 0 for x = [0,1]!
For PSD you must check all principal minors (every diagonal block, not only the top-left ones) — or just use eigenvalues.
Why: But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.
Trap
Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.
\[ B = \begin{bmatrix} 0 & 0 \\ 0 & -1 \end{bmatrix}, \quad D_1 = 0 \ge 0,\ \ D_2 = 0 \ge 0 \]
Conclude B is PSD because both leading minors are ≥ 0 ✗
Why: But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.
For PSD you must check all principal minors (every diagonal block, not only the top-left ones) — or just use eigenvalues.
\[ \text{principal minor from row/col 2} = \det[-1] = -1 < 0 \]
That principal minor is −1 < 0 → NOT PSD ✓
Why: The leading-minor shortcut certifies PD only. For PSD, one negative principal minor (or one negative eigenvalue) rules it out. When in doubt, run eigvalsh.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The leading-minor shortcut certifies PD only. For PSD, one negative principal minor (or one negative eigenvalue) rules it out. When in doubt, run eigvalsh.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.
Section
Part 5 of 6 — XᵀX, covariance, ridge
Intuition
You rarely have to test the matrices that matter in ML — they are PSD by construction. Two families cover almost everything: any Gram matrix XᵀX, and any covariance Σ.
Both are PSD for the same one-line reason: their quadratic form is a squared length, and a squared length is never negative. Once you see that argument, you recognize PSD on sight — no eigenvalues required.
Explain it
Discussion prompt
Explain The freebies you'll meet everywhere to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
You rarely have to test the matrices that matter in ML — they are PSD by construction. Two families cover almost everything: any Gram matrix XᵀX, and any covariance Σ.
Ranking
Put in order
Put the moves of Prove XᵀX is always PSD into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Sandwich XᵀX between xᵀ and x. Group as (Xx) using associativity.
Worked example
Write the quadratic form of XᵀX
Why: Sandwich XᵀX between xᵀ and x. Group as (Xx) using associativity.
\[ x^\top (X^\top X)\, x = (x^\top X^\top)(X x) \]
Recognize xᵀXᵀ = (Xx)ᵀ
Why: Transpose reverses a product: (Xx)ᵀ = xᵀXᵀ. So the form is (Xx)ᵀ(Xx) — a vector dotted with itself.
\[ = (X x)^\top (X x) = \lVert X x \rVert^2 \]
A squared norm is never negative
Why: ‖Xx‖² = Σ(Xx)ᵢ² ≥ 0 for EVERY x, no matter what X is. Therefore XᵀX is PSD for any real matrix X — always.
\[ \boxed{\,x^\top (X^\top X) x = \lVert X x \rVert^2 \ge 0 \quad \Rightarrow \quad X^\top X \succeq 0\,} \]
Notation
Annotate
From Prove XᵀX is always PSD — read this one piece at a time. What is each part doing?
On: \( \boxed{\,x^\top (X^\top X) x = \lVert X x \rVert^2 \ge 0 \quad \Rightarrow \quad X^\top X \succeq 0\,} \)
Concept
‖Xx‖² = 0 forces Xx = 0. So the form hits zero at a nonzero x exactly when some nonzero x satisfies Xx = 0 — i.e. X has a nontrivial null space, i.e. its columns are linearly dependent.
\[ X^\top X \succ 0 \iff X \text{ has full column rank} \]
This is the same condition that made the normal equations XᵀXw = Xᵀy uniquely solvable in Lesson 7. Full rank → PD → invertible → unique fit. Collinear columns → PSD-only → singular → no unique fit.
Fill the middle
Fill in the blanks
From The ‖Xx‖² identity, numerically — one line has had its right-hand side removed. Put it back.
import numpy as np
np.random.seed(0)
X = np.random.randn(5, 2)
G = X.T @ X
v = np.array([0.8, -0.5])
print(round(v @ G @ v, 6)) # x^T (X^T X) x
print(round((X @ v) @ (X @ v), 6)) # ||Xv||^2 -- equal
print(np.linalg.eigvalsh(G).round(4))
Why: G is what everything below it consumes, so the wrong expression here fails later and somewhere else. The form is a literal squared length, so it matches ‖Xv‖² to the last digit.
Worked example
Confirm xᵀ(XᵀX)x = ‖Xx‖² on a random full-rank X (seed fixed so you see these exact numbers):
import numpy as np
np.random.seed(0)
X = np.random.randn(5, 2)
G = X.T @ X
v = np.array([0.8, -0.5])
print(round(v @ G @ v, 6)) # x^T (X^T X) x
print(round((X @ v) @ (X @ v), 6)) # ||Xv||^2 -- equal
print(np.linalg.eigvalsh(G).round(4))Both equal 6.293183; eigenvalues [6.0082, 8.791] > 0
Why: The form is a literal squared length, so it matches ‖Xv‖² to the last digit. X has rank 2 (full column rank), so both eigenvalues are strictly positive → XᵀX is PD here.
| quantity | printed value |
|---|---|
| v @ (XᵀX) @ v | 6.293183 |
| ‖Xv‖² | 6.293183 |
| eigenvalues of XᵀX | [6.0082, 8.791] |
Pattern
Predict first
The table runs: rank(Xd) | 1 (of 2 cols) | rank-deficient · eigvalsh(XᵀX) | [0, 70] | PSD, not PD
In Collinear columns → PSD, not PD, given the rows so far: what is the next one — the row where quantity is det(XᵀX)?
Correct: det(XᵀX) | 0.0 | singular
| quantity | value | implies |
|---|---|---|
| rank(Xd) | 1 (of 2 cols) | rank-deficient |
| eigvalsh(XᵀX) | [0, 70] | PSD, not PD |
| det(XᵀX) | 0.0 | singular |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The redundant column gives a null direction x with Xx = 0, so ‖Xx‖² = 0 for that nonzero x — a zero eigenvalue.
Worked example
Make column 2 a copy of 2·column 1. Now X is rank-deficient, so XᵀX picks up a zero eigenvalue — PSD but singular. Runnable:
import numpy as np
Xd = np.array([[1., 2.], [2., 4.], [3., 6.]]) # col2 = 2 * col1
G = Xd.T @ Xd
print(np.linalg.matrix_rank(Xd)) # 1
print(np.linalg.eigvalsh(G)) # [0. 70.]
print(np.linalg.det(G)) # 0.0rank 1, eigenvalues [0, 70], det 0 → PSD not PD
Why: The redundant column gives a null direction x with Xx = 0, so ‖Xx‖² = 0 for that nonzero x — a zero eigenvalue. Still ≥ 0 everywhere (PSD), but singular (not PD). Exactly the collinearity failure of Lesson 7.
| quantity | value | implies |
|---|---|---|
| rank(Xd) | 1 (of 2 cols) | rank-deficient |
| eigvalsh(XᵀX) | [0, 70] | PSD, not PD |
| det(XᵀX) | 0.0 | singular |
Pattern
Step through it
Step through Collinear columns → PSD, not PD one row at a time. What is driving the change, and what would the row after the last one be?
Missing information
Discussion prompt
The sample covariance is Σ = XᶜᵀXᶜ / n where Xᶜ is the mean-centered data. That's a Gram matrix over √(1/n)·Xᶜ, so the same ‖·‖² ≥ 0 argument makes it PSD. Verify against np.cov:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Every covariance eigenvalue is a variance along a principal direction, so it must be ≥ 0. With 100 independent samples in 3-D the data has full rank, so all three are strictly positive. Matches np.cov(D.T, bias=True) exactly.
Worked example
The sample covariance is Σ = XᶜᵀXᶜ / n where Xᶜ is the mean-centered data. That's a Gram matrix over √(1/n)·Xᶜ, so the same ‖·‖² ≥ 0 argument makes it PSD. Verify against np.cov:
import numpy as np
np.random.seed(1)
D = np.random.randn(100, 3)
Dc = D - D.mean(axis=0)
Cov = (Dc.T @ Dc) / D.shape[0]
print(Cov.round(4))
print(np.linalg.eigvalsh(Cov).round(4)) # all >= 0Eigenvalues [0.7544, 0.8614, 1.0515] — all > 0 → PSD (here PD)
Why: Every covariance eigenvalue is a variance along a principal direction, so it must be ≥ 0. With 100 independent samples in 3-D the data has full rank, so all three are strictly positive. Matches np.cov(D.T, bias=True) exactly.
| quantity | value |
|---|---|
| diag(Cov) | [0.9973, 0.861, 0.809] |
| eigvalsh(Cov) | [0.7544, 0.8614, 1.0515] |
| matches np.cov? | yes (bias=True) |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Covariance matrices are PSD and we always invert Σ⁻¹ (Mahalanobis distance, the Gaussian). So they must be positive definite.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR.
Covariance is always PSD; it is PD only with full column rank and enough independent samples.
Why: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR. Inverting it divides by ~0 and blows up (the exact failure of Lessons 7 & 10).
Trap
Covariance matrices are PSD and we always invert Σ⁻¹ (Mahalanobis distance, the Gaussian). So they must be positive definite.
Assume Σ ≻ 0 always and invert it freely ✗
Why: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR. Inverting it divides by ~0 and blows up (the exact failure of Lessons 7 & 10).
Covariance is always PSD; it is PD only with full column rank and enough independent samples.
Σ ⪰ 0 always; Σ ≻ 0 iff Xᶜ has full column rank ✓
Why: This is why the multivariate Gaussian needs Σ ≻ 0 (so Σ⁻¹ exists), and why ridge adds λI — it lifts every eigenvalue above 0 to force PD and a stable inverse. Next slide shows that lift.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x and the matrix A: sandwich A between xᵀ on the left and x on the right.; A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.; A PSD-but-not-PD matrix has a flat direction — a nonzero x with xᵀAx = 0, a trough in the bowl. Every PD matrix is PSD, but not the reverse: PD ⟹ PSD, one-way.det > 0 must mean all eigenvalues are positive — hence PD.; Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.Estimation
Predict first
Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:
Commit before you compute: what does Ridge lifts every eigenvalue by λ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Eigenvalues [0,5] → [1,6]; Cholesky now succeeds
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Why the shift works: if Av = λv then (A+λ_r I)v = (λ + λ_r)v — same eigenvectors, every eigenvalue raised by λ_r.
Worked example
Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:
import numpy as np
M = np.array([[4., 2.], [2., 1.]]) # PSD, singular
print(np.linalg.eigvalsh(M)) # [0. 5.]
print(np.linalg.eigvalsh(M + np.eye(2))) # [1. 6.] every lambda + 1
L = np.linalg.cholesky(M + np.eye(2)) # now PD -> succeeds
print(L.round(4))Eigenvalues [0,5] → [1,6]; Cholesky now succeeds
Why: Why the shift works: if Av = λv then (A+λ_r I)v = (λ + λ_r)v — same eigenvectors, every eigenvalue raised by λ_r. The zero floor is gone, so M + I is PD and Cholesky returns a real L. This is exactly why ridge regression is always solvable.
| matrix | eigenvalues | cholesky |
|---|---|---|
| M | [0, 5] | LinAlgError (PSD only) |
| M + 1·I | [1, 6] | succeeds → PD |
| L of M+I | [[2.2361, 0], [0.8944, 1.0954]] |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Eigenvalues [0,5] → [1,6]; Cholesky now succeeds
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:
Concept
One more closure fact you'll reuse (and prove in homework): if P and Q are PSD, so is P + Q, because xᵀ(P+Q)x = xᵀPx + xᵀQx ≥ 0 — a sum of two non-negatives.
Concretely, our PD A = [[2,1],[1,2]] plus PSD M = [[4,2],[2,1]] gives S = [[6,3],[3,3]] with eigenvalues [1.146, 7.854] — both positive. (Ridge is the special case Q = λI.)
\[ P \succeq 0,\ Q \succeq 0 \;\Rightarrow\; x^\top(P+Q)x = \underbrace{x^\top P x}_{\ge 0} + \underbrace{x^\top Q x}_{\ge 0} \ge 0 \]
Analogy
Discussion prompt
Explain Sum of two PSD matrices is PSD by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One more closure fact you'll reuse (and prove in homework): if P and Q are PSD, so is P + Q, because xᵀ(P+Q)x = xᵀPx + xᵀQx ≥ 0 — a sum of two non-negatives.
Ranking
Put in order
These are the steps of The definiteness toolkit, scrambled. Put them back in order before the next slide shows you.
A — symmetrize with (A + Aᵀ)/2 if needednp.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/exceptnp.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite> 0; for PSD you need all principal minors, so prefer eigenvaluesXᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PDWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
A — symmetrize with (A + Aᵀ)/2 if needednp.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/exceptnp.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite> 0; for PSD you need all principal minors, so prefer eigenvaluesXᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PDEdge cases
Discussion prompt
The definiteness toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
A — symmetrize with (A + Aᵀ)/2 if needednp.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/exceptnp.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite> 0; for PSD you need all principal minors, so prefer eigenvaluesXᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PDElimination
Eliminate the wrong options
A symmetric matrix has eigenvalues {0, 3, 5}. It is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: All eigenvalues are ≥ 0, so xᵀAx = Σλᵢyᵢ² ≥ 0 for all x — PSD. But one eigenvalue is exactly 0 (a flat direction, and the matrix is singular), so it fails the strict > 0 test for PD.
Check
Read the eigenvalues, then decide.
Check your understanding
A symmetric matrix has eigenvalues {0, 3, 5}. It is:
Answer: A
Why: All eigenvalues are ≥ 0, so xᵀAx = Σλᵢyᵢ² ≥ 0 for all x — PSD. But one eigenvalue is exactly 0 (a flat direction, and the matrix is singular), so it fails the strict > 0 test for PD.
Prediction
Predict first
For a symmetric 2×2 matrix A, det(A) > 0 guarantees that A is:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: either positive definite or negative definite — det > 0 only fixes the sign of the product
Why: det = λ₁λ₂ > 0 means the two eigenvalues have the SAME sign — both positive (PD) or both negative (ND). It cannot tell them apart. You must also check a leading minor (e.g. A₁₁ > 0 for PD) or the eigenvalues themselves.
Check
Recall what the determinant does and doesn't tell you.
Check your understanding
For a symmetric 2×2 matrix A, det(A) > 0 guarantees that A is:
Answer: A
Why: det = λ₁λ₂ > 0 means the two eigenvalues have the SAME sign — both positive (PD) or both negative (ND). It cannot tell them apart. You must also check a leading minor (e.g. A₁₁ > 0 for PD) or the eigenvalues themselves.
Prediction
Predict first
For an arbitrary real matrix X, the Gram matrix XᵀX is always…
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: PSD, and PD iff X has full column rank
Why: xᵀ(XᵀX)x = (Xx)ᵀ(Xx) = ‖Xx‖² ≥ 0 for any X, so XᵀX is always PSD. It is PD exactly when ‖Xx‖² = 0 forces x = 0, i.e. Xx = 0 has only the trivial solution — X has full column rank.
Check
What is guaranteed about XᵀX for an arbitrary real X?
Check your understanding
For an arbitrary real matrix X, the Gram matrix XᵀX is always…
Answer: A
Why: xᵀ(XᵀX)x = (Xx)ᵀ(Xx) = ‖Xx‖² ≥ 0 for any X, so XᵀX is always PSD. It is PD exactly when ‖Xx‖² = 0 forces x = 0, i.e. Xx = 0 has only the trivial solution — X has full column rank.
Elimination
Eliminate the wrong options
What happens at np.linalg.cholesky(np.array([[4.,2.],[2.,1.]]))?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Cholesky needs a strictly positive diagonal; the last pivot for M would be 1 − (2/2)² = 0, not positive. NumPy raises LinAlgError('Matrix is not positive definite'), which is precisely the PD test — a raise means 'not PD'.
Check
You call np.linalg.cholesky(A) on the PSD-only M = [[4,2],[2,1]].
Check your understanding
What happens at np.linalg.cholesky(np.array([[4.,2.],[2.,1.]]))?
Answer: A
Why: Cholesky needs a strictly positive diagonal; the last pivot for M would be 1 − (2/2)² = 0, not positive. NumPy raises LinAlgError('Matrix is not positive definite'), which is precisely the PD test — a raise means 'not PD'.
Section
Part 6 of 6 — the project
Concept
Build two functions: is_pd(A) using Cholesky, and classify(A) using eigenvalues. Then prove they agree — is_pd is True exactly when classify returns 'PD'. You've derived every piece; assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | is_pd via Cholesky (try/except) | np.linalg.cholesky |
| 2 | classify via eigenvalues (PD/PSD/indefinite) | np.linalg.eigvalsh |
| 3 | Cross-check both on A, M, and N | compare verdicts |
Build rules: type every line yourself, catch LinAlgError (never let it crash), and always include a matrix you KNOW is only PSD (M) and one that is indefinite (N) — not just easy PD cases.
Counterexample
Discussion prompt
Build rules: type every line yourself, catch LinAlgError (never let it crash), and always include a matrix you KNOW is only PSD (M) and one that is indefinite (N) — not just easy PD cases.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: write is_pd(A) returning True/False. Predict the result for A = [[2,1],[1,2]] and M = [[4,2],[2,1]] before you run.
Hint: try a np.linalg.cholesky(A) and return True; except np.linalg.LinAlgError and return False. A raise means 'not PD'.
import numpy as np
def is_pd(A):
try:
np.linalg.cholesky(A)
return True
except np.linalg.LinAlgError:
return False
print(is_pd(np.array([[2., 1.], [1., 2.]]))) # True
print(is_pd(np.array([[4., 2.], [2., 1.]]))) # False| matrix | is_pd |
|---|---|
| A = [[2,1],[1,2]] | True |
| M = [[4,2],[2,1]] | False |
| reason | M's Cholesky raises |
Worked example
Your turn: classify into 'PD', 'PSD', or 'indefinite' from the eigenvalues. Predict what A, M, and N = [[1,2],[2,1]] return.
Hint: w = np.linalg.eigvalsh(A); if w.min() > tol → PD; elif w.min() > −tol → PSD; else indefinite. Use a small tol = 1e-10 to absorb float noise near zero.
import numpy as np
def classify(A, tol=1e-10):
w = np.linalg.eigvalsh(A)
if w.min() > tol:
return 'PD'
if w.min() > -tol:
return 'PSD'
return 'indefinite'
print(classify(np.array([[2., 1.], [1., 2.]]))) # PD
print(classify(np.array([[4., 2.], [2., 1.]]))) # PSD
print(classify(np.array([[1., 2.], [2., 1.]]))) # indefinite| matrix | min eigenvalue | classify |
|---|---|---|
| A = [[2,1],[1,2]] | 1 | PD |
| M = [[4,2],[2,1]] | 0 | PSD |
| N = [[1,2],[2,1]] | −1 | indefinite |
Worked example
Your turn: confirm the two tests never disagree — is_pd(A) must equal classify(A) == 'PD' for every matrix. Predict the agree column before running.
Hint: loop over A, M, N; print classify(M), is_pd(M), and the boolean is_pd(M) == (classify(M) == 'PD').
import numpy as np
def is_pd(A):
try:
np.linalg.cholesky(A); return True
except np.linalg.LinAlgError:
return False
def classify(A, tol=1e-10):
w = np.linalg.eigvalsh(A)
if w.min() > tol: return 'PD'
if w.min() > -tol: return 'PSD'
return 'indefinite'
for Mx in [np.array([[2.,1.],[1.,2.]]), np.array([[4.,2.],[2.,1.]]), np.array([[1.,2.],[2.,1.]])]:
print(classify(Mx), '| is_pd =', is_pd(Mx), '| agree =', is_pd(Mx) == (classify(Mx) == 'PD'))| matrix | classify | is_pd | agree |
|---|---|---|---|
| A | PD | True | True |
| M | PSD | False | True |
| N | indefinite | False | True |
Trade off
Comparison matrix
From Milestone 3 — prove they agree: every row here is a choice with a cost. Fill the agree column, then say which row you would actually pick and what you give up for it.
| matrix | classify | is_pd | agree |
|---|---|---|---|
| A | PD | True | True |
| M | PSD | False | True |
| N | indefinite | False | True |
Concept
import numpy as np
def is_pd(A):
try:
np.linalg.cholesky(A); return True
except np.linalg.LinAlgError:
return False
def classify(A, tol=1e-10):
w = np.linalg.eigvalsh(A)
if w.min() > tol: return 'PD'
if w.min() > -tol: return 'PSD'
return 'indefinite'
for Mx in [np.array([[2.,1.],[1.,2.]]),
np.array([[4.,2.],[2.,1.]]),
np.array([[1.,2.],[2.,1.]])]:
print(classify(Mx), '| is_pd =', is_pd(Mx))| matrix | printed output |
|---|---|
| A = [[2,1],[1,2]] | PD | is_pd = True |
| M = [[4,2],[2,1]] | PSD | is_pd = False |
| N = [[1,2],[2,1]] | indefinite | is_pd = False |
If Cholesky and the eigenvalue test agree on all three matrices — PD, PSD-only, and indefinite — you've built a definiteness checker the exam can't trick.
Comparison
Comparison matrix
From The full program: refill the printed output column from what you know. The rest of the table is as it appeared.
| matrix | printed output |
|---|---|
| A = [[2,1],[1,2]] | PD | is_pd = True |
| M = [[4,2],[2,1]] | PSD | is_pd = False |
| N = [[1,2],[2,1]] | indefinite | is_pd = False |
Concept
Slides closed, out loud: explain (1) why xᵀAx = ‖Xx‖² forces every XᵀX to be PSD, (2) what a single zero eigenvalue does — PSD vs PD, the flat direction, singularity — and (3) why det > 0 does not prove PD.
Stretch (homework): prove XᵀX is PD iff X has full column rank, and that the sum of two PSD matrices is PSD. These feed Mercer kernels (Week 10), the multivariate Gaussian's Σ ≻ 0, and the VAE latent covariance (Week 44).
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The quadratic form · PSD and PD, defined · Five tests, one matrix · PSD-only and indefinite · Where PSD comes from · Your turn: PD checker. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
xᵀAx into its polynomial (off-diagonal aᵢⱼ contributes 2aᵢⱼ) and read A back off itxᵀAx = Σλᵢyᵢ², so PSD ⟺ λ ≥ 0, PD ⟺ λ > 0A = [[2,1],[1,2]] is PD by all fiveM = [[4,2],[2,1]] ((2x₁+x₂)², eigenvalues [0,5]) and indefinite N = [[1,2],[2,1]] ([−1,3])XᵀX and covariance are PSD from ‖Xx‖² ≥ 0, PD iff full rank, and see ridge lift [0,5] → [1,6]| idea | the one thing to remember |
|---|---|
| PSD vs PD | xᵀAx ≥ 0 (λ ≥ 0) vs > 0 (λ > 0) |
| eigenvalue test | diagonalize: xᵀAx = Σλᵢyᵢ² |
| det > 0 | NOT enough for PD — check a leading minor too |
| Cholesky | succeeds ⟺ PD; the fast computational test |
| XᵀX / covariance | always PSD (‖Xx‖²); PD iff full column rank; +λI forces PD |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.