Lesson 13: Positive (Semi)Definite Matrices

USAAIO Lesson 13, from Week 5 on linear algebra, fully worked. It builds the quadratic form xᵀAx entry by entry, defines positive semi-definite and positive definite, and connects both to eigenvalues by completing the square. The running matrix A = [[2,1],[1,2]] is tested five independent ways - by the definition, by the completed square, by its eigenvalues, by Sylvester's minors, and by Cholesky - and contrasted against the singular M = [[4,2],[2,1]] and the indefinite N = [[1,2],[2,1]]. It then proves that XᵀX and the covariance matrix are PSD from ‖Xx‖² ≥ 0, shows ridge regression lifting every eigenvalue, and ends with a from-scratch definiteness checker that is cross-validated. Every snippet runs standalone, and every number comes from real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 114 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Positive (Semi)Definite Matrices

Title

USAAIO · Lesson 13 · Week 5 (Linear Algebra)

The matrices that behave like positive numbers. We build the quadratic form xᵀAx by hand, complete the square to see why the eigenvalue test works, then test one running matrix five independent ways — and prove every XᵀX and every covariance is PSD.

2. By the end of this lesson you can

Objectives

  1. Expand the quadratic form xᵀAx into its scalar polynomial and read the entries of A off it
  2. Define PSD (xᵀAx ≥ 0) and PD (> 0 ∀x≠0) and prove the eigenvalue link by completing the square
  3. Classify a matrix five ways — definition, completed square, eigenvalues, Sylvester's leading minors, and Cholesky A = LLᵀ
  4. Prove XᵀX and the covariance are always PSD from xᵀXᵀXx = ‖Xx‖² ≥ 0, and PD iff X has full column rank
  5. Build is_pd (Cholesky) and classify (eigenvalues) from scratch and prove they always agree

3. What survived from MAP, Priors & Regularization?

Warm-up

Discussion prompt

Before we open Lesson 13: Positive (Semi)Definite Matrices: without looking back, what was the main idea of MAP, Priors & Regularization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

MAP as MLE plus a log-prior, the Gaussian prior → L2/ridge and Laplace prior → L1/lasso correspondences, why L1 produces sparse solutions, conjugate priors, and dropout as approximate Bayesian inference. Compare ridge vs lasso sparsity from scratch.

4. The quadratic form

Section

Part 1 of 6 — what xᵀAx even is

5. Our running matrix

Concept

One matrix carries the whole lesson. Meet A — symmetric, 2×2, and (we'll prove) positive definite:

\[ A = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix} \]

It is symmetric (A = Aᵀ), which matters: definiteness is only defined for symmetric matrices. We'll test this exact A by five different criteria and watch them all agree.

6. Why symmetry is required

Concept

Only the symmetric part of a matrix survives in the quadratic form. The antisymmetric part contributes exactly zero, so a non-symmetric matrix and its symmetrization give the same xᵀAx.

\[ x^\top A x = x^\top\!\left(\tfrac{A + A^\top}{2}\right)\! x, \qquad \text{so symmetrize first: } A \leftarrow \tfrac{A + A^\top}{2} \]

Because of this, 'definiteness of A' really means definiteness of its symmetric part. We always symmetrize before testing — otherwise the eigenvalues can even come out complex and the whole framework breaks.

7. Break it if you can: Why symmetry is required

Counterexample

Discussion prompt

Only the symmetric part of a matrix survives in the quadratic form. The antisymmetric part contributes exactly zero, so a non-symmetric matrix and its symmetrization give the same xᵀAx.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

8. Guess the shape of the answer: Symmetrization leaves the form unchanged

Estimation

Predict first

Take a non-symmetric C = [[2,4],[0,2]]. Its form must equal that of its symmetric part (C + Cᵀ)/2. Confirm on x = [1.3, −0.7]. Self-contained:

Commit before you compute: what does Symmetrization leaves the form unchanged come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Both forms print 0.72; symmetric part is [[2,2],[2,2]]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The off-diagonal 4 and 0 average to 2 and 2.

9. Symmetrization leaves the form unchanged

Worked example

Take a non-symmetric C = [[2,4],[0,2]]. Its form must equal that of its symmetric part (C + Cᵀ)/2. Confirm on x = [1.3, −0.7]. Self-contained:

import numpy as np
C = np.array([[2., 4.], [0., 2.]])   # non-symmetric
x = np.array([1.3, -0.7])
S = (C + C.T) / 2                     # symmetric part
print(round(x @ C @ x, 4))           # form of C
print(round(x @ S @ x, 4))           # same!
print(S.tolist())                    # [[2, 2], [2, 2]]

Both forms print 0.72; symmetric part is [[2,2],[2,2]]

Why: The off-diagonal 4 and 0 average to 2 and 2. The antisymmetric leftover [[0,2],[−2,0]] contributes nothing to xᵀCx, so the two forms coincide. Test S, never the raw C.

quantityvalue
x @ C @ x (non-sym)0.72
x @ S @ x (sym part)0.72
S = (C+Cᵀ)/2[[2, 2], [2, 2]]
eigvalsh(S)[0, 4] → PSD

10. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at Symmetrization leaves the form unchanged. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.72
x @ C @ x (non-sym); x @ S @ x (sym part)
[[2, 2], [2, 2]]
S = (C+Cᵀ)/2
[0, 4] → PSD
eigvalsh(S)
g1
value is "0.72" for x @ C @ x (non-sym), x @ S @ x (sym part) — that is what the table on "Symmetrization leaves the form unchanged" records, and it is the single property separating this group from the rest.
g2
value is "[[2, 2], [2, 2]]" for S = (C+Cᵀ)/2 — that is what the table on "Symmetrization leaves the form unchanged" records, and it is the single property separating this group from the rest.
g3
value is "[0, 4] → PSD" for eigvalsh(S) — that is what the table on "Symmetrization leaves the form unchanged" records, and it is the single property separating this group from the rest.

11. The quadratic form xᵀAx

Concept

The central object is a single number built from a vector x and the matrix A: sandwich A between xᵀ on the left and x on the right.

\[ f(x) = x^\top A x \quad\in\; \mathbb{R} \]

quadratic form — A pure degree-2 polynomial in the entries of x, with the coefficients packed into a symmetric matrix A. Every term is xᵢxⱼ — no linear part, no constant. It is the matrix generalization of the single-variable a·x².

12. By analogy: The quadratic form xᵀAx

Analogy

Discussion prompt

Explain The quadratic form xᵀAx by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The central object is a single number built from a vector x and the matrix A: sandwich A between xᵀ on the left and x on the right.

13. What has to happen first: Expand xᵀAx entry by entry

Ranking

Put in order

Put the moves of Expand xᵀAx entry by entry into the order they have to happen.

  1. Inner product first: Ax
  2. Now xᵀ(Ax): dot x with that vector
  3. Distribute and collect

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Row i of A dotted with x. Row 0 is [2,1], row 1 is [1,2].

14. Expand xᵀAx entry by entry

Worked example

Take x = [x₁, x₂]. Multiply from the inside out. First Ax, then dot with x:

Inner product first: Ax

Why: Row i of A dotted with x. Row 0 is [2,1], row 1 is [1,2].

\[ A x = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = \begin{bmatrix} 2x_1 + x_2 \\ x_1 + 2x_2 \end{bmatrix} \]

Now xᵀ(Ax): dot x with that vector

Why: xᵀ times a column is x₁·(row 0 result) + x₂·(row 1 result).

\[ x^\top A x = x_1(2x_1 + x_2) + x_2(x_1 + 2x_2) \]

Distribute and collect

Why: x₁x₂ appears twice (from the two off-diagonal 1s) and combines into 2x₁x₂. This is the general rule: the off-diagonal entry aᵢⱼ shows up on BOTH sides, so its coefficient in the form is 2aᵢⱼ.

\[ x^\top A x = 2x_1^2 + 2x_1 x_2 + 2x_2^2 \]

15. Decode the notation: Expand xᵀAx entry by entry

Notation

Annotate

From Expand xᵀAx entry by entry — read this one piece at a time. What is each part doing?

On: \( A x = \begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = \begin{bmatrix} 2x_1 + x_2 \\ x_1 + 2x_2 \end{bmatrix} \)

  • Row i of A dotted with x. Row 0 is [2,1], row 1 is [1,2].
  • xᵀ times a column is x₁·(row 0 result) + x₂·(row 1 result).
  • x₁x₂ appears twice (from the two off-diagonal 1s) and combines into 2x₁x₂. This is the general rule: the off-diagonal entry aᵢⱼ shows up on BOTH sides, so its coefficient in the form is 2aᵢⱼ.

16. What has to be given first: Verify the form numerically

Missing information

Discussion prompt

Our hand expansion says xᵀAx = 2x₁² + 2x₁x₂ + 2x₂². Plug in a messy x = [1.3, −0.7] both ways and confirm they match. Runnable as written:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The matrix sandwich and the polynomial are the same function of x. Verified by execution — not memory.

17. Verify the form numerically

Worked example

Our hand expansion says xᵀAx = 2x₁² + 2x₁x₂ + 2x₂². Plug in a messy x = [1.3, −0.7] both ways and confirm they match. Runnable as written:

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4))          # x^T A x
print(round(2*x[0]**2 + 2*x[0]*x[1] + 2*x[1]**2, 4))

Both print 2.54 → the expansion is faithful

Why: The matrix sandwich and the polynomial are the same function of x. Verified by execution — not memory.

expressionprinted value
x @ A @ x2.54
2·x₁² + 2·x₁x₂ + 2·x₂²2.54
match?yes — identical

18. Fill in: printed value for Verify the form numerically

Comparison

Comparison matrix

From Verify the form numerically: refill the printed value column from what you know. The rest of the table is as it appeared.

expressionprinted value
x @ A @ x2.54
2·x₁² + 2·x₁x₂ + 2·x₂²2.54
match?yes — identical

19. Read the matrix off any 2×2 form

Concept

The pattern generalizes. For a symmetric A = [[a,b],[b,c]], expanding the sandwich gives a fixed template — the diagonal entries weight the pure squares, the off-diagonal is doubled in the cross term:

\[ \begin{bmatrix} x_1 & x_2 \end{bmatrix}\begin{bmatrix} a & b \\ b & c \end{bmatrix}\begin{bmatrix} x_1 \\ x_2 \end{bmatrix} = a\,x_1^2 + 2b\,x_1 x_2 + c\,x_2^2 \]

Go backwards: a form 4x₁² + 6x₁x₂ + 9x₂² comes from which A?

Why: Match coefficients: a = 4, c = 9, and 2b = 6 so b = 3. The matrix is [[4,3],[3,9]]. Always HALVE the cross-term coefficient to get the off-diagonal entry — the single most common slip.

\[ 4x_1^2 + 6x_1 x_2 + 9x_2^2 \;\longleftrightarrow\; \begin{bmatrix} 4 & 3 \\ 3 & 9 \end{bmatrix} \]

20. It's a bowl over the plane

Intuition

As x ranges over the plane, f(x) = xᵀAx traces a surface sitting above (or below) that plane, always passing through 0 at the origin.

The shape of that surface is the whole story. If it curves strictly upward in every direction — a genuine bowl with one lowest point — the matrix is positive definite. If it is a bowl that flattens into a trough along some line, it is only semidefinite. If it goes up one way and down another (a saddle), it is indefinite.

So 'classify the matrix' really means 'what does its bowl do?'. The next section makes that precise.

21. PSD and PD, defined

Section

Part 2 of 6 — and why eigenvalues decide it

22. Positive definite (PD)

Concept

A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.

\[ A \succ 0 \;\iff\; x^\top A x > 0 \quad \text{for all } x \neq 0 \]

The bowl strictly rises in every direction away from the origin — a single lowest point at x = 0. PD is the matrix version of a positive number.

23. Teach it back: Positive definite (PD)

Explain it

Discussion prompt

Explain Positive definite (PD) to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.

24. Positive semidefinite (PSD)

Concept

Positive semidefinite relaxes the strict > to ≥: the form is never negative, but is allowed to hit zero away from the origin.

\[ A \succeq 0 \;\iff\; x^\top A x \ge 0 \quad \text{for all } x \]

A PSD-but-not-PD matrix has a flat direction — a nonzero x with xᵀAx = 0, a trough in the bowl. Every PD matrix is PSD, but not the reverse: PD ⟹ PSD, one-way.

25. Why we care in ML

Intuition

PSD/PD is not a curiosity — it is a recurring gate across the whole course. A loss is convex exactly when its Hessian is PSD (Lesson 11). Least squares has a unique solution exactly when XᵀX is PD (Lesson 7). A covariance matrix is meaningful — a valid Gaussian — only because it is PSD.

Kernel methods demand a PSD Gram matrix (Mercer, Week 10); ridge regularization exists precisely to force a PSD-but-singular matrix to become PD. Learn the test once, reuse it everywhere.

26. The claim: eigenvalues decide it

Concept

Testing xᵀAx ≥ 0 for all x sounds impossible — infinitely many x. The rescue is that it reduces to a finite check on the eigenvalues:

\[ A \succeq 0 \iff \text{all } \lambda_i \ge 0, \qquad A \succ 0 \iff \text{all } \lambda_i > 0 \]

This isn't magic — it drops out of the spectral theorem plus one change of variables. We derive it next, no step skipped.

27. Plan first: Derive the eigenvalue test — set up

Step zero

Discussion prompt

Derive the eigenvalue test — set up — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Start from the spectral theorem

Answer:

  1. Start from the spectral theorem
  2. Substitute into the quadratic form
  3. Change variables: let y = Qᵀx

28. Derive the eigenvalue test — set up

Worked example

Start from the spectral theorem

Why: Every symmetric A factors as A = QΛQᵀ, where Q holds orthonormal eigenvectors (QᵀQ = I) and Λ = diag(λ₁,…,λₙ) holds the eigenvalues. This is Lesson 10.

\[ A = Q \Lambda Q^\top, \qquad Q^\top Q = I \]

Substitute into the quadratic form

Why: Replace A by QΛQᵀ inside xᵀAx.

\[ x^\top A x = x^\top Q \Lambda Q^\top x \]

Change variables: let y = Qᵀx

Why: Then xᵀQ = (Qᵀx)ᵀ = yᵀ. Because Q is orthogonal, y ranges over ALL of ℝⁿ exactly as x does — no direction is lost. So checking all x is the same as checking all y.

\[ x^\top A x = y^\top \Lambda y \]

29. Draw the shape of it: Derive the eigenvalue test — set up

Blank canvas

Draw it

Draw what Derive the eigenvalue test — set up just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

30. Plan first: Derive the eigenvalue test — finish

Step zero

Discussion prompt

Derive the eigenvalue test — finish — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Λ is diagonal, so yᵀΛy is a plain weighted sum of squares

Answer:

  1. Λ is diagonal, so yᵀΛy is a plain weighted sum of squares
  2. Read off the sign
  3. Conclusion

31. Derive the eigenvalue test — finish

Worked example

Λ is diagonal, so yᵀΛy is a plain weighted sum of squares

Why: A diagonal matrix in a quadratic form has no cross terms: yᵀΛy = Σ λᵢ yᵢ². Each eigenvalue simply weights the square of one coordinate.

\[ x^\top A x = \sum_{i} \lambda_i\, y_i^2 \]

Read off the sign

Why: Every yᵢ² ≥ 0. If every λᵢ ≥ 0, the sum is ≥ 0 for all y — that is PSD. If some λⱼ < 0, pick y = eⱼ (i.e. x = the j-th eigenvector): the sum is λⱼ < 0, so the form goes negative and A is not PSD.

Conclusion

Why: PSD ⟺ all λ ≥ 0, and PD ⟺ all λ > 0 (strict, because a zero eigenvalue makes yᵀΛy = 0 along that eigenvector — a flat direction). The infinite test collapsed to n sign checks.

\[ \boxed{\,A \succeq 0 \iff \lambda_i \ge 0\ \forall i, \qquad A \succ 0 \iff \lambda_i > 0\ \forall i\,} \]

32. Predict the next row: Test A by eigenvalues

Pattern

Predict first

The table runs: np.linalg.eigvalsh(A) | [1. 3.] | all > 0 → PD · np.linalg.det(A) | 2.9999999999999996 | ≈ 3 = λ₁·λ₂ (float noise)

In Test A by eigenvalues, given the rows so far: what is the next one — the row where quantity is np.trace(A)?

Correct: np.trace(A) | 4.0 | = λ₁ + λ₂

quantityprinted valuereads as
np.linalg.eigvalsh(A)[1. 3.]all > 0 → PD
np.linalg.det(A)2.9999999999999996≈ 3 = λ₁·λ₂ (float noise)
np.trace(A)4.0= λ₁ + λ₂

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Both strictly positive, so by the test A ≻ 0.

33. Test A by eigenvalues

Worked example

Apply the freshly derived test to our running A. Use eigvalsh — the symmetric-matrix eigenvalue routine, which returns real eigenvalues sorted ascending. Runnable:

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
print(np.linalg.det(A))             # 3.0 (up to float error)
print(np.linalg.eigvalsh(A))        # [1. 3.]
print(np.trace(A))                  # 4.0

Eigenvalues [1, 3] — both > 0 → A is PD

Why: Both strictly positive, so by the test A ≻ 0. Sanity: their product 1·3 = 3 = det(A), and their sum 1+3 = 4 = trace(A). Determinant is the product of eigenvalues; trace is their sum.

quantityprinted valuereads as
np.linalg.eigvalsh(A)[1. 3.]all > 0 → PD
np.linalg.det(A)2.9999999999999996≈ 3 = λ₁·λ₂ (float noise)
np.trace(A)4.0= λ₁ + λ₂

34. What each one costs: Test A by eigenvalues

Trade off

Comparison matrix

From Test A by eigenvalues: every row here is a choice with a cost. Fill the reads as column, then say which row you would actually pick and what you give up for it.

quantityprinted valuereads as
np.linalg.eigvalsh(A)[1. 3.]all > 0 → PD
np.linalg.det(A)2.9999999999999996≈ 3 = λ₁·λ₂ (float noise)
np.trace(A)4.0= λ₁ + λ₂

35. State the rule before it runs: Prove A ≻ 0 by hand: complete the square

Hypothesis

Predict first

Prove A ≻ 0 by hand: complete the square is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.

Correct: Pull the x₁ terms together

Why: Group the two terms containing x₁: 2x₁² + 2x₁x₂ = 2(x₁² + x₁x₂). We'll complete the square inside.

A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.

36. Prove A ≻ 0 by hand: complete the square

Worked example

Eigenvalues via a library is fine, but the exam wants you to show it. Take the form 2x₁² + 2x₁x₂ + 2x₂² and rewrite it as a sum of squares — then non-negativity is visible.

Pull the x₁ terms together

Why: Group the two terms containing x₁: 2x₁² + 2x₁x₂ = 2(x₁² + x₁x₂). We'll complete the square inside.

\[ 2x_1^2 + 2x_1 x_2 + 2x_2^2 = 2\left(x_1^2 + x_1 x_2\right) + 2x_2^2 \]

Complete the square inside the parentheses

Why: x₁² + x₁x₂ = (x₁ + x₂/2)² − x₂²/4. Multiply back by 2: 2(x₁ + x₂/2)² − x₂²/2.

\[ = 2\left(x_1 + \tfrac{x_2}{2}\right)^2 - \tfrac{1}{2}x_2^2 + 2x_2^2 \]

Combine the leftover x₂² terms

Why: −½x₂² + 2x₂² = 3/2·x₂². Both coefficients (2 and 3/2) are POSITIVE, so the whole form is a sum of positive multiples of squares — it is ≥ 0, and = 0 only when both squares vanish (x = 0). That is exactly A ≻ 0.

\[ \boxed{\,x^\top A x = 2\left(x_1 + \tfrac{x_2}{2}\right)^2 + \tfrac{3}{2}x_2^2 \;>\; 0\ \ (x \neq 0)\,} \]

37. Work backwards from the answer: Prove A ≻ 0 by hand: complete the square

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Combine the leftover x₂² terms

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Eigenvalues via a library is fine, but the exam wants you to show it. Take the form 2x₁² + 2x₁x₂ + 2x₂² and rewrite it as a sum of squares — then non-negativity is visible.

38. Restore the missing line: Check the completed square numerically

Fill the middle

Fill in the blanks

From Check the completed square numerically — one line has had its right-hand side removed. Put it back.

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4)) # matrix form
print(round(2(x[0] + x[1]/2)2 + 1.5x[1]**2, 4)) # completed square

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. Since both weights (2 and 1.5) are positive, the form can never be negative — an algebraic proof of PD that needs no eigenvalue solver.

39. Check the completed square numerically

Worked example

The completed-square form 2(x₁ + x₂/2)² + 1.5x₂² must equal xᵀAx for every x. Verify on the same x = [1.3, −0.7]. Runnable and self-contained:

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
x = np.array([1.3, -0.7])
print(round(x @ A @ x, 4))                        # matrix form
print(round(2*(x[0] + x[1]/2)**2 + 1.5*x[1]**2, 4)) # completed square

Both print 2.54 → the sum-of-squares form is exact

Why: Since both weights (2 and 1.5) are positive, the form can never be negative — an algebraic proof of PD that needs no eigenvalue solver.

formvalue at x=[1.3,−0.7]
x @ A @ x2.54
2(x₁+x₂/2)² + 1.5x₂²2.54
weights (2, 1.5)both > 0 → PD

40. Something is wrong here: does a positive determinant mean PD?

Anomaly

Predict first

A student writes this, and it looks reasonable:

The determinant is the product of eigenvalues, so det > 0 must mean all eigenvalues are positive — hence PD.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Two NEGATIVE eigenvalues also multiply to a positive determinant.

Check every leading principal minor (Sylvester) or all the eigenvalues — never just the full determinant.

Why: Two NEGATIVE eigenvalues also multiply to a positive determinant. det > 0 fixes only the sign of the PRODUCT, not each factor. This matrix is negative definite — its bowl points straight down.

41. Trap: does a positive determinant mean PD?

Trap

The trap

The determinant is the product of eigenvalues, so det > 0 must mean all eigenvalues are positive — hence PD.

\[ \det\!\begin{bmatrix} -1 & 0 \\ 0 & -1 \end{bmatrix} = (-1)(-1) = 1 > 0 \]

Conclude [[−1,0],[0,−1]] is PD because det = 1 > 0 ✗

Why: Two NEGATIVE eigenvalues also multiply to a positive determinant. det > 0 fixes only the sign of the PRODUCT, not each factor. This matrix is negative definite — its bowl points straight down.

The fix

Check every leading principal minor (Sylvester) or all the eigenvalues — never just the full determinant.

\[ \text{1st leading minor} = -1 < 0 \]

First minor is −1 < 0 → NOT PD (it is negative definite) ✓

Why: Sylvester's criterion needs the 1×1, 2×2, … leading minors ALL strictly positive. Here the very first one fails. Both eigenvalues are −1, confirming it.

42. Say it in words: Trap: does a positive determinant mean PD?

Translation

\( \text{1st leading minor} = -1 < 0 \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

43. Five tests, one matrix

Section

Part 3 of 6 — definition, square, eigen, Sylvester, Cholesky

44. The five equivalent tests

Concept

For a symmetric A, these five checks all decide definiteness — and all agree. Different exam questions reward different ones, so keep all five in hand:

#testPSD conditionPD condition
1DefinitionxᵀAx ≥ 0 ∀xxᵀAx > 0 ∀x≠0
2Completed squaresum of squares, coeffs ≥ 0coeffs all > 0
3Eigenvaluesall λ ≥ 0all λ > 0
4Sylvester (leading minors)— (need ALL minors)all leading minors > 0
5Cholesky A = LLᵀexists (diag ≥ 0)unique, diag > 0

Caution on Sylvester: the leading-minor test only certifies PD. For PSD you must check all principal minors (not just the leading ones), so for PSD the eigenvalue test is cleanest. We'll hit that trap explicitly.

45. What has to happen first: Test 4 — Sylvester's leading minors on A

Ranking

Put in order

Put the moves of Test 4 — Sylvester's leading minors on A into the order they have to happen.

  1. First leading minor: top-left 1×1 = 2
  2. Second leading minor: full 2×2 determinant = 3
  3. Both leading minors > 0 → A is PD

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. det = 2·2 − 1·1 = 4 − 1 = 3 > 0.

46. Test 4 — Sylvester's leading minors on A

Worked example

The leading principal minors are the determinants of the top-left 1×1, 2×2, … blocks. For our A = [[2,1],[1,2]] there are two:

First leading minor: top-left 1×1 = 2

Why: det[2] = 2 > 0. The first check passes.

\[ D_1 = \det\begin{bmatrix} 2 \end{bmatrix} = 2 > 0 \]

Second leading minor: full 2×2 determinant = 3

Why: det = 2·2 − 1·1 = 4 − 1 = 3 > 0. The second check passes too.

\[ D_2 = \det\begin{bmatrix} 2 & 1 \\ 1 & 2 \end{bmatrix} = 4 - 1 = 3 > 0 \]

Both leading minors > 0 → A is PD

Why: Sylvester: all leading minors strictly positive certifies PD. This is the fastest by-hand test for a small matrix — no eigenvalues needed.

leading minorvalue> 0?
D₁ (1×1)2yes
D₂ (2×2 det)3yes
verdictPD✓

47. Watch it run: Test 4 — Sylvester's leading minors on A

Pattern

Step through it

Step through Test 4 — Sylvester's leading minors on A one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: leading minor is D₁ (1×1)
  2. Step 2: leading minor is D₂ (2×2 det)
  3. Step 3: leading minor is verdict

48. Test 5 — the Cholesky decomposition

Concept

Cholesky factors a PD matrix as A = LLᵀ, where L is lower-triangular with positive diagonal. It is the matrix square root that exists exactly when A ≻ 0.

\[ A = L L^\top, \qquad L = \begin{bmatrix} \ell_{11} & 0 \\ \ell_{21} & \ell_{22} \end{bmatrix},\ \ \ell_{11}, \ell_{22} > 0 \]

This is the computer's favorite definiteness test: np.linalg.cholesky(A) succeeds iff A is PD and raises LinAlgError otherwise — a fast, numerically stable yes/no with no eigenvalue solve.

49. Where does each piece belong: Lesson 13: Positive (Semi)Definite Matrices

Sorting

Sort into buckets

These are the pieces of Lesson 13: Positive (Semi)Definite Matrices, out of order. Put each one back under the part of the lesson it belongs to.

The quadratic form
Our running matrix; Why symmetry is required; Symmetrization leaves the form unchanged
PSD and PD, defined
Positive definite (PD); Positive semidefinite (PSD); Why we care in ML
Five tests, one matrix
The five equivalent tests; Test 4 — Sylvester's leading minors on A; Test 5 — the Cholesky decomposition
s1
The quadratic form is where Lesson 13: Positive (Semi)Definite Matrices puts Our running matrix, Why symmetry is required, Symmetrization leaves the form unchanged. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
PSD and PD, defined is where Lesson 13: Positive (Semi)Definite Matrices puts Positive definite (PD), Positive semidefinite (PSD), Why we care in ML. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Five tests, one matrix is where Lesson 13: Positive (Semi)Definite Matrices puts The five equivalent tests, Test 4 — Sylvester's leading minors on A, Test 5 — the Cholesky decomposition. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

50. Plan first: Compute Cholesky of A by hand

Step zero

Discussion prompt

Compute Cholesky of A by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Top-left: ℓ₁₁² = 2 → ℓ₁₁ = √2

Answer:

  1. Top-left: ℓ₁₁² = 2 → ℓ₁₁ = √2
  2. Off-diagonal: ℓ₂₁·ℓ₁₁ = 1 → ℓ₂₁ = 1/√2
  3. Bottom-right: ℓ₂₁² + ℓ₂₂² = 2 → ℓ₂₂ = √(3/2)

51. Compute Cholesky of A by hand

Worked example

Match LLᵀ = A entry by entry. With L = [[ℓ₁₁,0],[ℓ₂₁,ℓ₂₂]], multiply out and solve:

Top-left: ℓ₁₁² = 2 → ℓ₁₁ = √2

Why: (LLᵀ)₁₁ = ℓ₁₁². Set equal to A₁₁ = 2 and take the positive root: ℓ₁₁ = √2 ≈ 1.4142.

\[ \ell_{11}^2 = 2 \;\Rightarrow\; \ell_{11} = \sqrt{2} \]

Off-diagonal: ℓ₂₁·ℓ₁₁ = 1 → ℓ₂₁ = 1/√2

Why: (LLᵀ)₂₁ = ℓ₂₁ℓ₁₁. Set equal to A₂₁ = 1, so ℓ₂₁ = 1/√2 ≈ 0.7071.

\[ \ell_{21}\,\ell_{11} = 1 \;\Rightarrow\; \ell_{21} = \tfrac{1}{\sqrt{2}} \]

Bottom-right: ℓ₂₁² + ℓ₂₂² = 2 → ℓ₂₂ = √(3/2)

Why: (LLᵀ)₂₂ = ℓ₂₁² + ℓ₂₂² = 2. Since ℓ₂₁² = 1/2, we need ℓ₂₂² = 3/2, so ℓ₂₂ = √1.5 ≈ 1.2247. Every diagonal entry is a real positive number → the factorization exists → A is PD.

\[ L = \begin{bmatrix} \sqrt{2} & 0 \\ \tfrac{1}{\sqrt2} & \sqrt{\tfrac{3}{2}} \end{bmatrix} \approx \begin{bmatrix} 1.4142 & 0 \\ 0.7071 & 1.2247 \end{bmatrix} \]

52. Finish it with less help: Confirm Cholesky in NumPy

Faded example

Fill in the blanks

Confirm Cholesky in NumPy, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
L = np.linalg.cholesky(A)
print(L.round(6))
print((L @ L.T).round(6)) # rebuilds A

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The solver returns [[1.414214, 0],[0.707107, 1.224745]] — our √2, 1/√2, √1.5 — and multiplying back reconstructs [[2,1],[1,2]].

53. Confirm Cholesky in NumPy

Worked example

np.linalg.cholesky should return exactly our hand-computed L, and L @ L.T should rebuild A. Self-contained:

import numpy as np
A = np.array([[2., 1.], [1., 2.]])
L = np.linalg.cholesky(A)
print(L.round(6))
print((L @ L.T).round(6))            # rebuilds A

L matches our hand answer; LLᵀ = A exactly

Why: The solver returns [[1.414214, 0],[0.707107, 1.224745]] — our √2, 1/√2, √1.5 — and multiplying back reconstructs [[2,1],[1,2]]. Cholesky succeeded, so A is certified PD.

outputvalue
L row 0[1.414214, 0.0]
L row 1[0.707107, 1.224745]
L @ L.T[[2, 1], [1, 2]] = A

54. All five agree on A

Concept

Five independent roads, one destination — A = [[2,1],[1,2]] is positive definite by every test:

testevidenceverdict
1. DefinitionxᵀAx = 2.54 > 0 at test xPD
2. Completed square2(x₁+x₂/2)² + 1.5x₂², coeffs > 0PD
3. Eigenvalues[1, 3], both > 0PD
4. Sylvester minorsD₁=2>0, D₂=3>0PD
5. CholeskyL exists, positive diagonalPD

That redundancy is the point: if two tests disagree, you made an arithmetic error — the tests themselves never do.

55. PSD-only and indefinite

Section

Part 4 of 6 — the contrast cases

56. Three bowls, three verdicts

Intuition

Every symmetric 2×2 produces one of three surface shapes, and the eigenvalue signs name them directly:

shape of xᵀAxeigenvalue signsverdictexample
bowl — up in every directionboth > 0PD[[2,1],[1,2]]
trough — flat along a lineone = 0PSD only[[4,2],[2,1]]
saddle — up one way, down anotherone < 0indefinite[[1,2],[2,1]]

PD is a single lowest point; PSD-only trades that point for a whole flat line at the bottom; indefinite has no bottom at all. The next three slides make each concrete.

57. The exam favorite: M = [[4,2],[2,1]]

Concept

Now a matrix that is PSD but not PD. It looks similar to A but hides a zero eigenvalue:

\[ M = \begin{bmatrix} 4 & 2 \\ 2 & 1 \end{bmatrix} \]

Its quadratic form is a perfect square, which is exactly why it can touch zero. Watch the completed square collapse to a single term.

58. Plan first: M's form is a perfect square

Step zero

Discussion prompt

M's form is a perfect square — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Expand xᵀMx

Answer:

  1. Expand xᵀMx
  2. Recognize the perfect square
  3. Find the flat direction

59. M's form is a perfect square

Worked example

Expand xᵀMx

Why: Off-diagonal 2 contributes 2·2 = 4 to the cross term. So the form is 4x₁² + 4x₁x₂ + x₂².

\[ x^\top M x = 4x_1^2 + 4x_1 x_2 + x_2^2 \]

Recognize the perfect square

Why: 4x₁² + 4x₁x₂ + x₂² = (2x₁ + x₂)². Only ONE square, no leftover second term — that missing second square is the zero eigenvalue.

\[ x^\top M x = (2x_1 + x_2)^2 \ge 0 \]

Find the flat direction

Why: (2x₁ + x₂)² = 0 whenever 2x₁ + x₂ = 0, e.g. x = [1, −2]. There the form is zero for a NONZERO x — so M is PSD but not PD. That line is the trough in the bowl.

\[ x = \begin{bmatrix} 1 \\ -2 \end{bmatrix} \;\Rightarrow\; x^\top M x = (2 - 2)^2 = 0 \]

60. Draw the shape of it: M's form is a perfect square

Blank canvas

Draw it

Draw what M's form is a perfect square just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

61. Restore the missing line: Verify M is PSD-only in NumPy

Fill the middle

Fill in the blanks

From Verify M is PSD-only in NumPy — one line has had its right-hand side removed. Put it back.

import numpy as np
M = np.array([[4., 2.], [2., 1.]])
print(np.linalg.det(M)) # 0.0
print(np.linalg.eigvalsh(M)) # [0. 5.]
x = np.array([1., -2.])
print(round(x @ M @ x, 6)) # 0.0 (flat direction)

Why: x is what everything below it consumes, so the wrong expression here fails later and somewhere else. The zero eigenvalue makes M singular (det = 0) and gives the flat direction.

62. Verify M is PSD-only in NumPy

Worked example

Determinant 0, eigenvalues [0, 5], and the form vanishing at [1, −2] — all three confirm PSD-not-PD. Self-contained:

import numpy as np
M = np.array([[4., 2.], [2., 1.]])
print(np.linalg.det(M))             # 0.0
print(np.linalg.eigvalsh(M))        # [0. 5.]
x = np.array([1., -2.])
print(round(x @ M @ x, 6))          # 0.0  (flat direction)

det 0, eigenvalues [0, 5], form 0 at [1,−2]

Why: The zero eigenvalue makes M singular (det = 0) and gives the flat direction. All eigenvalues ≥ 0 → PSD; the zero → not PD, not invertible.

testresultimplies
np.linalg.eigvalsh(M)[0, 5]PSD (≥0), not PD
np.linalg.det(M)0.0singular → not PD
xᵀMx at [1,−2]0.0flat direction on a line

63. Guess the shape of the answer: Cholesky refuses M

Estimation

Predict first

Because Cholesky needs a positive diagonal, it raises on a merely-PSD matrix. Catch the error instead of crashing. Runnable:

Commit before you compute: what does Cholesky refuses M come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Prints: not PD: Matrix is not positive definite

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The final pivot ℓ₂₂² would be 1 − (2/2)² = 0, not positive, so no real positive diagonal exists.

64. Cholesky refuses M

Worked example

Because Cholesky needs a positive diagonal, it raises on a merely-PSD matrix. Catch the error instead of crashing. Runnable:

import numpy as np
M = np.array([[4., 2.], [2., 1.]])
try:
    np.linalg.cholesky(M)
    print("PD")
except np.linalg.LinAlgError as e:
    print("not PD:", e)

Prints: not PD: Matrix is not positive definite

Why: The final pivot ℓ₂₂² would be 1 − (2/2)² = 0, not positive, so no real positive diagonal exists. cholesky raises LinAlgError — this IS the PD test, so a raise means 'not PD'.

matrixcholeskyverdict
A = [[2,1],[1,2]]succeedsPD
M = [[4,2],[2,1]]LinAlgErrornot PD (only PSD)
error message"Matrix is not positive definite"

65. Fill in: cholesky for Cholesky refuses M

Comparison

Comparison matrix

From Cholesky refuses M: refill the cholesky column from what you know. The rest of the table is as it appeared.

matrixcholeskyverdict
A = [[2,1],[1,2]]succeedsPD
M = [[4,2],[2,1]]LinAlgErrornot PD (only PSD)
error message"Matrix is not positive definite"

66. Indefinite: N = [[1,2],[2,1]]

Worked example

One more contrast — a matrix whose bowl goes up one way and down another. Its form has a genuine negative direction:

import numpy as np
N = np.array([[1., 2.], [2., 1.]])
print(np.linalg.eigvalsh(N))        # [-1. 3.]
print(round(np.array([1., -1.]) @ N @ np.array([1., -1.]), 4))  # -2
print(round(np.array([1.,  1.]) @ N @ np.array([1.,  1.]), 4))  #  6

Eigenvalues [−1, 3] → one negative → indefinite

Why: xᵀNx = x₁² + 4x₁x₂ + x₂². At [1,−1] it is 1 − 4 + 1 = −2 (negative direction); at [1,1] it is 1 + 4 + 1 = 6 (positive direction). Up AND down → neither PSD nor negative definite.

direction xxᵀNxsign
[1, −1]−2negative (down)
[1, 1]6positive (up)
eigenvalues[−1, 3]mixed → indefinite

67. Something is wrong here: Sylvester's leading minors for PSD

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: But xᵀBx = −x₂² < 0 for x = [0,1]!

For PSD you must check all principal minors (every diagonal block, not only the top-left ones) — or just use eigenvalues.

Why: But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.

68. Trap: Sylvester's leading minors for PSD

Trap

The trap

Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.

\[ B = \begin{bmatrix} 0 & 0 \\ 0 & -1 \end{bmatrix}, \quad D_1 = 0 \ge 0,\ \ D_2 = 0 \ge 0 \]

Conclude B is PSD because both leading minors are ≥ 0 ✗

Why: But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.

The fix

For PSD you must check all principal minors (every diagonal block, not only the top-left ones) — or just use eigenvalues.

\[ \text{principal minor from row/col 2} = \det[-1] = -1 < 0 \]

That principal minor is −1 < 0 → NOT PSD ✓

Why: The leading-minor shortcut certifies PD only. For PSD, one negative principal minor (or one negative eigenvalue) rules it out. When in doubt, run eigvalsh.

69. Break it on purpose: Sylvester's leading minors for PSD

Break the constraint

Discussion prompt

The rule this trap just fixed:

The leading-minor shortcut certifies PD only. For PSD, one negative principal minor (or one negative eigenvalue) rules it out. When in doubt, run eigvalsh.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

But xᵀBx = −x₂² < 0 for x = [0,1]! B has eigenvalue −1 — it is NOT PSD. Leading minors ≥ 0 is not enough: they miss the bottom-right block.

70. Where PSD comes from

Section

Part 5 of 6 — XᵀX, covariance, ridge

71. The freebies you'll meet everywhere

Intuition

You rarely have to test the matrices that matter in ML — they are PSD by construction. Two families cover almost everything: any Gram matrix XᵀX, and any covariance Σ.

Both are PSD for the same one-line reason: their quadratic form is a squared length, and a squared length is never negative. Once you see that argument, you recognize PSD on sight — no eigenvalues required.

72. Teach it back: The freebies you'll meet everywhere

Explain it

Discussion prompt

Explain The freebies you'll meet everywhere to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

You rarely have to test the matrices that matter in ML — they are PSD by construction. Two families cover almost everything: any Gram matrix XᵀX, and any covariance Σ.

73. What has to happen first: Prove XᵀX is always PSD

Ranking

Put in order

Put the moves of Prove XᵀX is always PSD into the order they have to happen.

  1. Write the quadratic form of XᵀX
  2. Recognize xᵀXᵀ = (Xx)ᵀ
  3. A squared norm is never negative

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Sandwich XᵀX between xᵀ and x. Group as (Xx) using associativity.

74. Prove XᵀX is always PSD

Worked example

Write the quadratic form of XᵀX

Why: Sandwich XᵀX between xᵀ and x. Group as (Xx) using associativity.

\[ x^\top (X^\top X)\, x = (x^\top X^\top)(X x) \]

Recognize xᵀXᵀ = (Xx)ᵀ

Why: Transpose reverses a product: (Xx)ᵀ = xᵀXᵀ. So the form is (Xx)ᵀ(Xx) — a vector dotted with itself.

\[ = (X x)^\top (X x) = \lVert X x \rVert^2 \]

A squared norm is never negative

Why: ‖Xx‖² = Σ(Xx)ᵢ² ≥ 0 for EVERY x, no matter what X is. Therefore XᵀX is PSD for any real matrix X — always.

\[ \boxed{\,x^\top (X^\top X) x = \lVert X x \rVert^2 \ge 0 \quad \Rightarrow \quad X^\top X \succeq 0\,} \]

75. Decode the notation: Prove XᵀX is always PSD

Notation

Annotate

From Prove XᵀX is always PSD — read this one piece at a time. What is each part doing?

On: \( \boxed{\,x^\top (X^\top X) x = \lVert X x \rVert^2 \ge 0 \quad \Rightarrow \quad X^\top X \succeq 0\,} \)

  • Sandwich XᵀX between xᵀ and x. Group as (Xx) using associativity.
  • Transpose reverses a product: (Xx)ᵀ = xᵀXᵀ. So the form is (Xx)ᵀ(Xx) — a vector dotted with itself.
  • ‖Xx‖² = Σ(Xx)ᵢ² ≥ 0 for EVERY x, no matter what X is. Therefore XᵀX is PSD for any real matrix X — always.

76. When is XᵀX strictly PD?

Concept

‖Xx‖² = 0 forces Xx = 0. So the form hits zero at a nonzero x exactly when some nonzero x satisfies Xx = 0 — i.e. X has a nontrivial null space, i.e. its columns are linearly dependent.

\[ X^\top X \succ 0 \iff X \text{ has full column rank} \]

This is the same condition that made the normal equations XᵀXw = Xᵀy uniquely solvable in Lesson 7. Full rank → PD → invertible → unique fit. Collinear columns → PSD-only → singular → no unique fit.

77. Restore the missing line: The ‖Xx‖² identity, numerically

Fill the middle

Fill in the blanks

From The ‖Xx‖² identity, numerically — one line has had its right-hand side removed. Put it back.

import numpy as np
np.random.seed(0)
X = np.random.randn(5, 2)
G = X.T @ X
v = np.array([0.8, -0.5])
print(round(v @ G @ v, 6)) # x^T (X^T X) x
print(round((X @ v) @ (X @ v), 6)) # ||Xv||^2 -- equal
print(np.linalg.eigvalsh(G).round(4))

Why: G is what everything below it consumes, so the wrong expression here fails later and somewhere else. The form is a literal squared length, so it matches ‖Xv‖² to the last digit.

78. The ‖Xx‖² identity, numerically

Worked example

Confirm xᵀ(XᵀX)x = ‖Xx‖² on a random full-rank X (seed fixed so you see these exact numbers):

import numpy as np
np.random.seed(0)
X = np.random.randn(5, 2)
G = X.T @ X
v = np.array([0.8, -0.5])
print(round(v @ G @ v, 6))          # x^T (X^T X) x
print(round((X @ v) @ (X @ v), 6))  # ||Xv||^2  -- equal
print(np.linalg.eigvalsh(G).round(4))

Both equal 6.293183; eigenvalues [6.0082, 8.791] > 0

Why: The form is a literal squared length, so it matches ‖Xv‖² to the last digit. X has rank 2 (full column rank), so both eigenvalues are strictly positive → XᵀX is PD here.

quantityprinted value
v @ (XᵀX) @ v6.293183
‖Xv‖²6.293183
eigenvalues of XᵀX[6.0082, 8.791]

79. Predict the next row: Collinear columns → PSD, not PD

Pattern

Predict first

The table runs: rank(Xd) | 1 (of 2 cols) | rank-deficient · eigvalsh(XᵀX) | [0, 70] | PSD, not PD

In Collinear columns → PSD, not PD, given the rows so far: what is the next one — the row where quantity is det(XᵀX)?

Correct: det(XᵀX) | 0.0 | singular

quantityvalueimplies
rank(Xd)1 (of 2 cols)rank-deficient
eigvalsh(XᵀX)[0, 70]PSD, not PD
det(XᵀX)0.0singular

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The redundant column gives a null direction x with Xx = 0, so ‖Xx‖² = 0 for that nonzero x — a zero eigenvalue.

80. Collinear columns → PSD, not PD

Worked example

Make column 2 a copy of 2·column 1. Now X is rank-deficient, so XᵀX picks up a zero eigenvalue — PSD but singular. Runnable:

import numpy as np
Xd = np.array([[1., 2.], [2., 4.], [3., 6.]])  # col2 = 2 * col1
G = Xd.T @ Xd
print(np.linalg.matrix_rank(Xd))    # 1
print(np.linalg.eigvalsh(G))        # [0. 70.]
print(np.linalg.det(G))             # 0.0

rank 1, eigenvalues [0, 70], det 0 → PSD not PD

Why: The redundant column gives a null direction x with Xx = 0, so ‖Xx‖² = 0 for that nonzero x — a zero eigenvalue. Still ≥ 0 everywhere (PSD), but singular (not PD). Exactly the collinearity failure of Lesson 7.

quantityvalueimplies
rank(Xd)1 (of 2 cols)rank-deficient
eigvalsh(XᵀX)[0, 70]PSD, not PD
det(XᵀX)0.0singular

81. Watch it run: Collinear columns → PSD, not PD

Pattern

Step through it

Step through Collinear columns → PSD, not PD one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is rank(Xd)
  2. Step 2: quantity is eigvalsh(XᵀX)
  3. Step 3: quantity is det(XᵀX)

82. What has to be given first: Covariance is a scaled Gram matrix

Missing information

Discussion prompt

The sample covariance is Σ = XᶜᵀXᶜ / n where Xᶜ is the mean-centered data. That's a Gram matrix over √(1/n)·Xᶜ, so the same ‖·‖² ≥ 0 argument makes it PSD. Verify against np.cov:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Every covariance eigenvalue is a variance along a principal direction, so it must be ≥ 0. With 100 independent samples in 3-D the data has full rank, so all three are strictly positive. Matches np.cov(D.T, bias=True) exactly.

83. Covariance is a scaled Gram matrix

Worked example

The sample covariance is Σ = XᶜᵀXᶜ / n where Xᶜ is the mean-centered data. That's a Gram matrix over √(1/n)·Xᶜ, so the same ‖·‖² ≥ 0 argument makes it PSD. Verify against np.cov:

import numpy as np
np.random.seed(1)
D = np.random.randn(100, 3)
Dc = D - D.mean(axis=0)
Cov = (Dc.T @ Dc) / D.shape[0]
print(Cov.round(4))
print(np.linalg.eigvalsh(Cov).round(4))   # all >= 0

Eigenvalues [0.7544, 0.8614, 1.0515] — all > 0 → PSD (here PD)

Why: Every covariance eigenvalue is a variance along a principal direction, so it must be ≥ 0. With 100 independent samples in 3-D the data has full rank, so all three are strictly positive. Matches np.cov(D.T, bias=True) exactly.

quantityvalue
diag(Cov)[0.9973, 0.861, 0.809]
eigvalsh(Cov)[0.7544, 0.8614, 1.0515]
matches np.cov?yes (bias=True)

84. Something is wrong here: is a covariance matrix always PD?

Anomaly

Predict first

A student writes this, and it looks reasonable:

Covariance matrices are PSD and we always invert Σ⁻¹ (Mahalanobis distance, the Gaussian). So they must be positive definite.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR.

Covariance is always PSD; it is PD only with full column rank and enough independent samples.

Why: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR. Inverting it divides by ~0 and blows up (the exact failure of Lessons 7 & 10).

85. Trap: is a covariance matrix always PD?

Trap

The trap

Covariance matrices are PSD and we always invert Σ⁻¹ (Mahalanobis distance, the Gaussian). So they must be positive definite.

Assume Σ ≻ 0 always and invert it freely ✗

Why: If features are collinear, or you have fewer samples than dimensions (n < d), Σ has a zero eigenvalue — PSD but SINGULAR. Inverting it divides by ~0 and blows up (the exact failure of Lessons 7 & 10).

The fix

Covariance is always PSD; it is PD only with full column rank and enough independent samples.

Σ ⪰ 0 always; Σ ≻ 0 iff Xᶜ has full column rank ✓

Why: This is why the multivariate Gaussian needs Σ ≻ 0 (so Σ⁻¹ exists), and why ridge adds λI — it lifts every eigenvalue above 0 to force PD and a stable inverse. Next slide shows that lift.

86. Which of these survive contact with Lesson 13: Positive (Semi)Definite Matrices?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
The central object is a single number built from a vector x and the matrix A: sandwich A between xᵀ on the left and x on the right.; A symmetric matrix A is positive definite if its quadratic form is strictly positive for every nonzero direction x.; A PSD-but-not-PD matrix has a flat direction — a nonzero x with xᵀAx = 0, a trough in the bowl. Every PD matrix is PSD, but not the reverse: PD ⟹ PSD, one-way.
Breaks
The determinant is the product of eigenvalues, so det > 0 must mean all eigenvalues are positive — hence PD.; Sylvester says 'all leading minors > 0 ⟹ PD'. So to test PSD, just relax to 'all leading minors ≥ 0'.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 13: Positive (Semi)Definite Matrices puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

87. Guess the shape of the answer: Ridge lifts every eigenvalue by λ

Estimation

Predict first

Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:

Commit before you compute: what does Ridge lifts every eigenvalue by λ come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Eigenvalues [0,5] → [1,6]; Cholesky now succeeds

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Why the shift works: if Av = λv then (A+λ_r I)v = (λ + λ_r)v — same eigenvectors, every eigenvalue raised by λ_r.

88. Ridge lifts every eigenvalue by λ

Worked example

Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:

import numpy as np
M = np.array([[4., 2.], [2., 1.]])       # PSD, singular
print(np.linalg.eigvalsh(M))              # [0. 5.]
print(np.linalg.eigvalsh(M + np.eye(2)))  # [1. 6.]  every lambda + 1
L = np.linalg.cholesky(M + np.eye(2))     # now PD -> succeeds
print(L.round(4))

Eigenvalues [0,5] → [1,6]; Cholesky now succeeds

Why: Why the shift works: if Av = λv then (A+λ_r I)v = (λ + λ_r)v — same eigenvectors, every eigenvalue raised by λ_r. The zero floor is gone, so M + I is PD and Cholesky returns a real L. This is exactly why ridge regression is always solvable.

matrixeigenvaluescholesky
M[0, 5]LinAlgError (PSD only)
M + 1·I[1, 6]succeeds → PD
L of M+I[[2.2361, 0], [0.8944, 1.0954]]

89. Work backwards from the answer: Ridge lifts every eigenvalue by λ

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Eigenvalues [0,5] → [1,6]; Cholesky now succeeds

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Take the singular M = [[4,2],[2,1]] (eigenvalues [0, 5]). Adding λI shifts every eigenvalue up by λ, so λ = 1 turns the deadly 0 into 1 and the matrix becomes PD:

90. Sum of two PSD matrices is PSD

Concept

One more closure fact you'll reuse (and prove in homework): if P and Q are PSD, so is P + Q, because xᵀ(P+Q)x = xᵀPx + xᵀQx ≥ 0 — a sum of two non-negatives.

Concretely, our PD A = [[2,1],[1,2]] plus PSD M = [[4,2],[2,1]] gives S = [[6,3],[3,3]] with eigenvalues [1.146, 7.854] — both positive. (Ridge is the special case Q = λI.)

\[ P \succeq 0,\ Q \succeq 0 \;\Rightarrow\; x^\top(P+Q)x = \underbrace{x^\top P x}_{\ge 0} + \underbrace{x^\top Q x}_{\ge 0} \ge 0 \]

91. By analogy: Sum of two PSD matrices is PSD

Analogy

Discussion prompt

Explain Sum of two PSD matrices is PSD by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

One more closure fact you'll reuse (and prove in homework): if P and Q are PSD, so is P + Q, because xᵀ(P+Q)x = xᵀPx + xᵀQx ≥ 0 — a sum of two non-negatives.

92. Rebuild the recipe: The definiteness toolkit

Ranking

Put in order

These are the steps of The definiteness toolkit, scrambled. Put them back in order before the next slide shows you.

  1. Symmetric first. PSD/PD are only defined for symmetric A — symmetrize with (A + Aᵀ)/2 if needed
  2. Fast yes/no (PD): np.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/except
  3. Full picture: np.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite
  4. By hand (PD): Sylvester — all leading minors > 0; for PSD you need all principal minors, so prefer eigenvalues
  5. Recognize the freebies: any XᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PD

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

93. The definiteness toolkit

Pattern

  1. Symmetric first. PSD/PD are only defined for symmetric A — symmetrize with (A + Aᵀ)/2 if needed
  2. Fast yes/no (PD): np.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/except
  3. Full picture: np.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite
  4. By hand (PD): Sylvester — all leading minors > 0; for PSD you need all principal minors, so prefer eigenvalues
  5. Recognize the freebies: any XᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PD

94. Where does it stop working: The definiteness toolkit

Edge cases

Discussion prompt

The definiteness toolkit works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Symmetric first. PSD/PD are only defined for symmetric A — symmetrize with (A + Aᵀ)/2 if needed
  2. Fast yes/no (PD): np.linalg.cholesky(A) succeeds ⟺ PD; wrap it in try/except
  3. Full picture: np.linalg.eigvalsh(A) — all ≥ 0 ⟹ PSD, all > 0 ⟹ PD, any < 0 ⟹ indefinite
  4. By hand (PD): Sylvester — all leading minors > 0; for PSD you need all principal minors, so prefer eigenvalues
  5. Recognize the freebies: any XᵀX and any covariance is PSD (‖Xx‖² ≥ 0); PD iff full column rank; ridge +λI forces PD

95. Rule out three: Check yourself — PSD vs PD

Elimination

Eliminate the wrong options

A symmetric matrix has eigenvalues {0, 3, 5}. It is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. PSD but not PD
  • B. PD
  • C. Neither PSD nor PD
  • D. Negative definite

Survives elimination: A

Why: All eigenvalues are ≥ 0, so xᵀAx = Σλᵢyᵢ² ≥ 0 for all x — PSD. But one eigenvalue is exactly 0 (a flat direction, and the matrix is singular), so it fails the strict > 0 test for PD.

96. Check yourself — PSD vs PD

Check

Read the eigenvalues, then decide.

Check your understanding

A symmetric matrix has eigenvalues {0, 3, 5}. It is:

  • A. PSD but not PD (correct)
  • B. PD
  • C. Neither PSD nor PD
  • D. Negative definite

Answer: A

Why: All eigenvalues are ≥ 0, so xᵀAx = Σλᵢyᵢ² ≥ 0 for all x — PSD. But one eigenvalue is exactly 0 (a flat direction, and the matrix is singular), so it fails the strict > 0 test for PD.

Why B tempts people
PD requires ALL eigenvalues strictly positive. The zero eigenvalue gives a nonzero x with xᵀAx = 0 (the corresponding eigenvector), which violates PD and makes the matrix singular.
Why C tempts people
It IS PSD — every eigenvalue is non-negative, so the form is never negative. 'Neither' would require at least one negative eigenvalue.
Why D tempts people
Negative definite needs all eigenvalues < 0. Here they are 0, 3, 5 — the opposite sign, so the bowl opens upward, not downward.

97. Answer it before you see the options: Check yourself — the det > 0 trap

Prediction

Predict first

For a symmetric 2×2 matrix A, det(A) > 0 guarantees that A is:

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: either positive definite or negative definite — det > 0 only fixes the sign of the product

Why: det = λ₁λ₂ > 0 means the two eigenvalues have the SAME sign — both positive (PD) or both negative (ND). It cannot tell them apart. You must also check a leading minor (e.g. A₁₁ > 0 for PD) or the eigenvalues themselves.

98. Check yourself — the det > 0 trap

Check

Recall what the determinant does and doesn't tell you.

Check your understanding

For a symmetric 2×2 matrix A, det(A) > 0 guarantees that A is:

  • A. either positive definite or negative definite — det > 0 only fixes the sign of the product (correct)
  • B. positive definite
  • C. positive semidefinite but not definite
  • D. indefinite

Answer: A

Why: det = λ₁λ₂ > 0 means the two eigenvalues have the SAME sign — both positive (PD) or both negative (ND). It cannot tell them apart. You must also check a leading minor (e.g. A₁₁ > 0 for PD) or the eigenvalues themselves.

Why B tempts people
[[−1,0],[0,−1]] has det = 1 > 0 but is negative definite. det > 0 alone never certifies PD — it's the classic trap.
Why C tempts people
A singular (PSD-not-PD) matrix has a zero eigenvalue, so det = λ₁λ₂ = 0, not > 0. det > 0 actually rules out the semidefinite-only case.
Why D tempts people
An indefinite 2×2 has eigenvalues of opposite sign, so det = λ₁λ₂ < 0. det > 0 rules indefinite OUT, not in.

99. Answer it before you see the options: Check yourself — where PSD comes from

Prediction

Predict first

For an arbitrary real matrix X, the Gram matrix XᵀX is always…

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: PSD, and PD iff X has full column rank

Why: xᵀ(XᵀX)x = (Xx)ᵀ(Xx) = ‖Xx‖² ≥ 0 for any X, so XᵀX is always PSD. It is PD exactly when ‖Xx‖² = 0 forces x = 0, i.e. Xx = 0 has only the trivial solution — X has full column rank.

100. Check yourself — where PSD comes from

Check

What is guaranteed about XᵀX for an arbitrary real X?

Check your understanding

For an arbitrary real matrix X, the Gram matrix XᵀX is always…

  • A. PSD, and PD iff X has full column rank (correct)
  • B. always PD
  • C. PSD only when X is square
  • D. indefinite in general

Answer: A

Why: xᵀ(XᵀX)x = (Xx)ᵀ(Xx) = ‖Xx‖² ≥ 0 for any X, so XᵀX is always PSD. It is PD exactly when ‖Xx‖² = 0 forces x = 0, i.e. Xx = 0 has only the trivial solution — X has full column rank.

Why B tempts people
Not always PD: if X has collinear (linearly dependent) columns, some nonzero x gives Xx = 0, hence ‖Xx‖² = 0 — a zero eigenvalue, so PSD but not PD.
Why C tempts people
The ‖Xx‖² argument holds for ANY shape of X, tall, wide, or square. Squareness is irrelevant to the sign of a squared norm.
Why D tempts people
A squared norm is never negative, so xᵀ(XᵀX)x ≥ 0 always — XᵀX can never be indefinite; it is at least PSD.

101. Rule out three: Check yourself — Cholesky as a test

Elimination

Eliminate the wrong options

What happens at np.linalg.cholesky(np.array([[4.,2.],[2.,1.]]))?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It raises LinAlgError — M is PSD but not positive definite
  • B. It returns L = [[2,0],[1,0]] with a zero on the diagonal
  • C. It returns the eigenvalues [0, 5]
  • D. It succeeds — M is symmetric, which is all Cholesky needs

Survives elimination: A

Why: Cholesky needs a strictly positive diagonal; the last pivot for M would be 1 − (2/2)² = 0, not positive. NumPy raises LinAlgError('Matrix is not positive definite'), which is precisely the PD test — a raise means 'not PD'.

102. Check yourself — Cholesky as a test

Check

You call np.linalg.cholesky(A) on the PSD-only M = [[4,2],[2,1]].

Check your understanding

What happens at np.linalg.cholesky(np.array([[4.,2.],[2.,1.]]))?

  • A. It raises LinAlgError — M is PSD but not positive definite (correct)
  • B. It returns L = [[2,0],[1,0]] with a zero on the diagonal
  • C. It returns the eigenvalues [0, 5]
  • D. It succeeds — M is symmetric, which is all Cholesky needs

Answer: A

Why: Cholesky needs a strictly positive diagonal; the last pivot for M would be 1 − (2/2)² = 0, not positive. NumPy raises LinAlgError('Matrix is not positive definite'), which is precisely the PD test — a raise means 'not PD'.

Why B tempts people
NumPy does not return a degenerate L with a zero pivot; it raises before producing one. A zero diagonal would make the factor non-unique and the routine rejects it.
Why C tempts people
cholesky computes a triangular factor, not eigenvalues. Eigenvalues come from eigvalsh; cholesky either returns L or raises.
Why D tempts people
Symmetry is necessary but not sufficient. M is symmetric yet only PSD, so Cholesky still fails. It succeeds only for strictly positive-definite matrices.

103. Your turn: PD checker

Section

Part 6 of 6 — the project

104. Project: a definiteness checker

Concept

Build two functions: is_pd(A) using Cholesky, and classify(A) using eigenvalues. Then prove they agree — is_pd is True exactly when classify returns 'PD'. You've derived every piece; assemble it yourself.

#requirementtool
1is_pd via Cholesky (try/except)np.linalg.cholesky
2classify via eigenvalues (PD/PSD/indefinite)np.linalg.eigvalsh
3Cross-check both on A, M, and Ncompare verdicts

Build rules: type every line yourself, catch LinAlgError (never let it crash), and always include a matrix you KNOW is only PSD (M) and one that is indefinite (N) — not just easy PD cases.

105. Break it if you can: Project: a definiteness checker

Counterexample

Discussion prompt

Build rules: type every line yourself, catch LinAlgError (never let it crash), and always include a matrix you KNOW is only PSD (M) and one that is indefinite (N) — not just easy PD cases.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

106. Milestone 1 — is_pd via Cholesky

Worked example

Your turn: write is_pd(A) returning True/False. Predict the result for A = [[2,1],[1,2]] and M = [[4,2],[2,1]] before you run.

Hint: try a np.linalg.cholesky(A) and return True; except np.linalg.LinAlgError and return False. A raise means 'not PD'.

import numpy as np
def is_pd(A):
    try:
        np.linalg.cholesky(A)
        return True
    except np.linalg.LinAlgError:
        return False
print(is_pd(np.array([[2., 1.], [1., 2.]])))  # True
print(is_pd(np.array([[4., 2.], [2., 1.]])))  # False
matrixis_pd
A = [[2,1],[1,2]]True
M = [[4,2],[2,1]]False
reasonM's Cholesky raises

107. Milestone 2 — classify via eigenvalues

Worked example

Your turn: classify into 'PD', 'PSD', or 'indefinite' from the eigenvalues. Predict what A, M, and N = [[1,2],[2,1]] return.

Hint: w = np.linalg.eigvalsh(A); if w.min() > tol → PD; elif w.min() > −tol → PSD; else indefinite. Use a small tol = 1e-10 to absorb float noise near zero.

import numpy as np
def classify(A, tol=1e-10):
    w = np.linalg.eigvalsh(A)
    if w.min() > tol:
        return 'PD'
    if w.min() > -tol:
        return 'PSD'
    return 'indefinite'
print(classify(np.array([[2., 1.], [1., 2.]])))   # PD
print(classify(np.array([[4., 2.], [2., 1.]])))   # PSD
print(classify(np.array([[1., 2.], [2., 1.]])))   # indefinite
matrixmin eigenvalueclassify
A = [[2,1],[1,2]]1PD
M = [[4,2],[2,1]]0PSD
N = [[1,2],[2,1]]−1indefinite

108. Milestone 3 — prove they agree

Worked example

Your turn: confirm the two tests never disagree — is_pd(A) must equal classify(A) == 'PD' for every matrix. Predict the agree column before running.

Hint: loop over A, M, N; print classify(M), is_pd(M), and the boolean is_pd(M) == (classify(M) == 'PD').

import numpy as np
def is_pd(A):
    try:
        np.linalg.cholesky(A); return True
    except np.linalg.LinAlgError:
        return False
def classify(A, tol=1e-10):
    w = np.linalg.eigvalsh(A)
    if w.min() > tol:  return 'PD'
    if w.min() > -tol: return 'PSD'
    return 'indefinite'
for Mx in [np.array([[2.,1.],[1.,2.]]), np.array([[4.,2.],[2.,1.]]), np.array([[1.,2.],[2.,1.]])]:
    print(classify(Mx), '| is_pd =', is_pd(Mx), '| agree =', is_pd(Mx) == (classify(Mx) == 'PD'))
matrixclassifyis_pdagree
APDTrueTrue
MPSDFalseTrue
NindefiniteFalseTrue

109. What each one costs: Milestone 3 — prove they agree

Trade off

Comparison matrix

From Milestone 3 — prove they agree: every row here is a choice with a cost. Fill the agree column, then say which row you would actually pick and what you give up for it.

matrixclassifyis_pdagree
APDTrueTrue
MPSDFalseTrue
NindefiniteFalseTrue

110. The full program

Concept

import numpy as np

def is_pd(A):
    try:
        np.linalg.cholesky(A); return True
    except np.linalg.LinAlgError:
        return False

def classify(A, tol=1e-10):
    w = np.linalg.eigvalsh(A)
    if w.min() > tol:  return 'PD'
    if w.min() > -tol: return 'PSD'
    return 'indefinite'

for Mx in [np.array([[2.,1.],[1.,2.]]),
           np.array([[4.,2.],[2.,1.]]),
           np.array([[1.,2.],[2.,1.]])]:
    print(classify(Mx), '| is_pd =', is_pd(Mx))
matrixprinted output
A = [[2,1],[1,2]]PD | is_pd = True
M = [[4,2],[2,1]]PSD | is_pd = False
N = [[1,2],[2,1]]indefinite | is_pd = False

If Cholesky and the eigenvalue test agree on all three matrices — PD, PSD-only, and indefinite — you've built a definiteness checker the exam can't trick.

111. Fill in: printed output for The full program

Comparison

Comparison matrix

From The full program: refill the printed output column from what you know. The rest of the table is as it appeared.

matrixprinted output
A = [[2,1],[1,2]]PD | is_pd = True
M = [[4,2],[2,1]]PSD | is_pd = False
N = [[1,2],[2,1]]indefinite | is_pd = False

112. Show it off

Concept

Slides closed, out loud: explain (1) why xᵀAx = ‖Xx‖² forces every XᵀX to be PSD, (2) what a single zero eigenvalue does — PSD vs PD, the flat direction, singularity — and (3) why det > 0 does not prove PD.

Stretch (homework): prove XᵀX is PD iff X has full column rank, and that the sum of two PSD matrices is PSD. These feed Mercer kernels (Week 10), the multivariate Gaussian's Σ ≻ 0, and the VAE latent covariance (Week 44).

113. Connect it up: Lesson 13: Positive (Semi)Definite Matrices

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The quadratic form · PSD and PD, defined · Five tests, one matrix · PSD-only and indefinite · Where PSD comes from · Your turn: PD checker. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

114. What you can do now

Recap

ideathe one thing to remember
PSD vs PDxᵀAx ≥ 0 (λ ≥ 0) vs > 0 (λ > 0)
eigenvalue testdiagonalize: xᵀAx = Σλᵢyᵢ²
det > 0NOT enough for PD — check a leading minor too
Choleskysucceeds ⟺ PD; the fast computational test
XᵀX / covariancealways PSD (‖Xx‖²); PD iff full column rank; +λI forces PD

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 13 (Week 5 — PSD/PD Matrices) — Barron · USAAIO Round 2 Preparation, 2026
  2. NumPy linalg.cholesky / eigvalsh
  3. Gilbert Strang, Introduction to Linear Algebra, Ch. 6 (Positive Definite Matrices) — Wellesley-Cambridge Press, 5th ed.
  4. Every eigenvalue, Cholesky factor, determinant, and quadratic-form value produced by real execution — numpy 2.2.6, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108