Lesson 7: Linear Systems & the Normal Equations

USAAIO Lesson 7, from Week 3 on linear algebra, fully worked. It writes out overdetermined systems equation by equation, then derives the normal equations XᵀXw = Xᵀy twice - once by projection and once by calculus - with no steps skipped. It solves the 5-student 2×2 system by hand, entry by entry, covers invertibility and rank, derives ridge regression from the penalized loss, and covers R². It ends with a from-scratch OLS implementation verified against sklearn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 66 slides.

Subject: Machine Learning · 113 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Linear Systems & The Normal Equations

Title

USAAIO · Lesson 7 · Week 3 (Linear Algebra)

When there is no exact answer, find the best one. We derive XᵀXw = Xᵀy twice — once by geometry, once by calculus — skipping no step, then solve it by hand and build the OLS that becomes every regression you'll ever fit.

2. By the end of this lesson you can

Objectives

  1. Write an overdetermined system Xw = y out in full and say exactly why it has no exact solution
  2. Derive the normal equations XᵀXw = Xᵀy from the perpendicular residual and independently from ∇‖Xw − y‖² = 0
  3. Solve the 5-student 2×2 system by hand, entry by entry, and verify the residual is perpendicular to every column
  4. State when XᵀX is invertible (full column rank) and derive ridge (XᵀX + λI)w = Xᵀy when it is not
  5. Implement OLS in NumPy with solve (never inv), score R², and prove it matches sklearn.LinearRegression

3. What survived from Gradient Intuition & PyTorch Autograd?

Warm-up

Discussion prompt

Before we open Lesson 7: Linear Systems & the Normal Equations: without looking back, what was the main idea of Gradient Intuition & PyTorch Autograd, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

partial derivatives, the gradient as steepest ascent, directional derivatives, the three gradient identities, and PyTorch autograd — ending with a build-it-yourself gradient descent project.

4. The problem: no exact answer

Section

Part 1 of 8

5. The data: five students

Concept

We have five students. For each we know hours studied x and their score y, and we want to fit a straight line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁.

studenthours xscore y
112
224
335
444
555

Two unknowns (w₀, w₁), but five data points. Keep that mismatch in mind — it is the whole story of this lesson.

6. Fill in: hours x for The data: five students

Comparison

Comparison matrix

From The data: five students: refill the hours x column from what you know. The rest of the table is as it appeared.

studenthours xscore y
112
224
335
444
555

7. Guess the shape of the answer: Write the system out in full

Estimation

Predict first

If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:

Commit before you compute: what does Write the system out in full come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Count: 5 equations, 2 unknowns

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. This is an OVERDETERMINED system — more constraints than degrees of freedom.

8. Write the system out in full

Worked example

If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:

\[ \begin{aligned} w_0 + 1\,w_1 &= 2 \\ w_0 + 2\,w_1 &= 4 \\ w_0 + 3\,w_1 &= 5 \\ w_0 + 4\,w_1 &= 4 \\ w_0 + 5\,w_1 &= 5 \end{aligned} \]

Count: 5 equations, 2 unknowns

Why: This is an OVERDETERMINED system — more constraints than degrees of freedom. That is the defining feature of a regression problem.

9. Work backwards from the answer: Write the system out in full

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Count: 5 equations, 2 unknowns

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:

10. What has to be given first: Stack it into matrix form Xw = y

Missing information

Discussion prompt

Every one of those five equations is w₀·(1) + w₁·(xᵢ). Collect the coefficients into a matrix X, the unknowns into w, the targets into y:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Column 0 multiplies w₀ (the intercept), so it is a column of ones. Column 1 holds the feature x. Miss this column and your line is forced through the origin — we'll hit that trap head-on.

11. Stack it into matrix form Xw = y

Worked example

Every one of those five equations is w₀·(1) + w₁·(xᵢ). Collect the coefficients into a matrix X, the unknowns into w, the targets into y:

\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix}, \qquad w = \begin{bmatrix} w_0 \\ w_1 \end{bmatrix}, \qquad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \]

The first column of X is all 1s — the bias column

Why: Column 0 multiplies w₀ (the intercept), so it is a column of ones. Column 1 holds the feature x. Miss this column and your line is forced through the origin — we'll hit that trap head-on.

12. Check: Xw reproduces the five equations

Worked example

Matrix–vector multiply Xw = each row of X dotted with w. Row i is [1, xᵢ], so row i of Xw is 1·w₀ + xᵢ·w₁:

\[ Xw = \begin{bmatrix} w_0 + 1\,w_1 \\ w_0 + 2\,w_1 \\ w_0 + 3\,w_1 \\ w_0 + 4\,w_1 \\ w_0 + 5\,w_1 \end{bmatrix} \;\overset{?}{=}\; \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} = y \]

Xw = y is exactly the five equations, packed

Why: The matrix equation IS the system — no new content, just compact notation. Asking 'does Xw = y have a solution?' asks 'do all five hold at once?'

13. Why there is no exact solution

Concept

Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.

For example the line through students 1 and 2 (w₀ = 0, w₁ = 2) predicts 6 for student 3, who actually scored 5. No single (w₀, w₁) makes all five equalities true at once.

So Xw = y is inconsistent — it has no exact solution. We stop demanding equality.

14. Break it if you can: Why there is no exact solution

Counterexample

Discussion prompt

Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

For example the line through students 1 and 2 (w₀ = 0, w₁ = 2) predicts 6 for student 3, who actually scored 5. No single (w₀, w₁) makes all five equalities true at once.

15. Real data is noisy — that's the point

Intuition

The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.

So exactness is the wrong goal. We don't want the line that hits points — we want the line that is closest to all of them at once.

'Closest' needs a definition of distance. That definition is what turns a hopeless system into a solvable one.

16. By analogy: Real data is noisy — that's the point

Analogy

Discussion prompt

Explain Real data is noisy — that's the point by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.

17. The residual: how wrong each prediction is

Concept

For a candidate w, the prediction is ŷ = Xw and the residual is the miss on each point, rᵢ = yᵢ − ŷᵢ. As a vector:

\[ r = y - Xw \quad\in\; \mathbb{R}^5 \]

r = 0 would mean a perfect fit — impossible here. So instead we make r as small as we can. But 'small' for a 5-vector needs a single number to minimize.

18. Teach it back: The residual: how wrong each prediction is

Explain it

Discussion prompt

Explain The residual: how wrong each prediction is to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

For a candidate w, the prediction is ŷ = Xw and the residual is the miss on each point, rᵢ = yᵢ − ŷᵢ. As a vector:

19. Measuring size: the squared norm

Concept

Collapse the residual vector into one number with the squared Euclidean norm — the sum of squared misses:

\[ \lVert r \rVert^2 = \sum_{i=1}^{5} r_i^2 = r^\top r \]

Squaring makes every miss positive (so overshoots and undershoots both cost), punishes big misses hardest, and — crucially — is smooth, so calculus can find its minimum.

\[ w^\star = \arg\min_{w} \; \lVert Xw - y \rVert^2 \]

20. The objective, written out for our data

Worked example

Substitute the five residuals into ‖Xw − y‖². This is the actual function of (w₀, w₁) we are minimizing — nothing hidden:

\[ L(w_0,w_1) = \sum_{i=1}^{5}\big(w_0 + w_1 x_i - y_i\big)^2 \]

\[ \begin{aligned} =\;& (w_0 + w_1 - 2)^2 + (w_0 + 2w_1 - 4)^2 + (w_0 + 3w_1 - 5)^2 \\ &+\, (w_0 + 4w_1 - 4)^2 + (w_0 + 5w_1 - 5)^2 \end{aligned} \]

L is a sum of squares — a smooth bowl over the (w₀, w₁) plane

Why: Each term is a parabola; a sum of parabolas opens upward with a single lowest point. Our whole job is to find the (w₀, w₁) at the bottom of that bowl.

21. Derivation 1: projection

Section

Part 2 of 8 — the geometry

22. The column space: what Xw can reach

Concept

Watch what Xw is: Xw = w₀·(column 0) + w₁·(column 1). It is a linear combination of the columns of X.

\[ Xw = w_0 \begin{bmatrix}1\\1\\1\\1\\1\end{bmatrix} + w_1 \begin{bmatrix}1\\2\\3\\4\\5\end{bmatrix} \]

column space of X — The set of ALL vectors you can build as combinations of X's columns — every prediction the model can possibly produce. With 2 columns in ℝ⁵, it is a 2-dimensional plane living inside 5-dimensional space.

23. y lives off the plane

Intuition

The target y is a point in ℝ⁵. The column space is a flat 2-D plane through the origin inside that ℝ⁵. Because the system is inconsistent, y does not lie on the plane — it floats above it.

No choice of w can reach y. The best we can do is reach the point on the plane that sits nearest to y.

24. Picture it first: Nearest point = perpendicular foot

Picture it

Figure (svg): A point y above a horizontal plane, with a vertical dashed perpendicular dropping to its foot Xw-star on the plane, and a slanted longer line to a different point on the plane.

The perpendicular drop is the shortest — any other point on the plane is farther from y.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

The closest point on a plane to an outside point is its perpendicular shadow — the foot of the perpendicular dropped from y onto the plane. That shadow is Xw★, the best prediction.

25. Nearest point = perpendicular foot

Concept

The closest point on a plane to an outside point is its perpendicular shadow — the foot of the perpendicular dropped from y onto the plane. That shadow is Xw★, the best prediction.

Figure (svg): A point y above a horizontal plane, with a vertical dashed perpendicular dropping to its foot Xw-star on the plane, and a slanted longer line to a different point on the plane.

The perpendicular drop is the shortest — any other point on the plane is farther from y.

26. Minimizing distance ⇔ perpendicular residual

Intuition

‖Xw − y‖ is the length of the arrow from the plane-point Xw up to y. Minimizing that length pins Xw to the perpendicular foot.

At that foot, the leftover arrow r = y − Xw★ points straight out of the plane — it is perpendicular to the plane, hence perpendicular to both columns of X.

That single geometric fact — residual ⟂ every column — is enough to pin down w★ algebraically. Next slide turns it into equations.

27. Perpendicular means dot product zero

Concept

Two vectors are perpendicular exactly when their dot product is 0. So 'the residual is perpendicular to each column of X' means each column, dotted with r, gives 0.

The dot of a column with r is that column written as a row (its transpose) times r. So the perpendicularity conditions are (columnⱼ)ᵀ r = 0 for every column j.

28. The two perpendicularity conditions

Worked example

Column 0 is the ones vector; column 1 is x. Dot each with the residual r = y − Xw and set to zero:

\[ \text{(col 0)}\quad \sum_{i} 1\cdot(y_i - \hat y_i) = 0 \;\Longrightarrow\; \sum_i r_i = 0 \]

\[ \text{(col 1)}\quad \sum_{i} x_i\,(y_i - \hat y_i) = 0 \;\Longrightarrow\; \sum_i x_i r_i = 0 \]

Two columns → two scalar equations → two unknowns

Why: The mismatch is fixed: we traded five impossible equalities for two perpendicularity conditions, exactly matching the two unknowns w₀, w₁.

29. Stack the conditions: Xᵀr = 0

Worked example

Those two dot products are precisely the two entries of the matrix–vector product Xᵀr. Row 0 of Xᵀ is the ones vector, row 1 is x:

\[ X^\top r = \begin{bmatrix} 1 & 1 & 1 & 1 & 1 \\ 1 & 2 & 3 & 4 & 5 \end{bmatrix} r = \begin{bmatrix} \sum_i r_i \\ \sum_i x_i r_i \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} \]

All perpendicularity conditions collapse to Xᵀr = 0

Why: Multiplying by Xᵀ stacks 'dot with column j = 0' for every column into ONE clean matrix equation. This is the geometric heart of least squares.

30. What has to happen first: Substitute r and distribute → the normal equations

Ranking

Put in order

Put the moves of Substitute r and distribute → the normal equations into the order they have to happen.

  1. Start from Xᵀr = 0
  2. Distribute Xᵀ across the subtraction
  3. Move the XᵀXw term to the other side

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The perpendicular-residual condition we just derived.

31. Substitute r and distribute → the normal equations

Worked example

Now put r = y − Xw★ back in and expand. Two lines of algebra, both shown:

Start from Xᵀr = 0

Why: The perpendicular-residual condition we just derived.

\[ X^\top (y - X w^\star) = 0 \]

Distribute Xᵀ across the subtraction

Why: Matrix multiplication distributes over addition, just like numbers: Xᵀ(y − Xw) = Xᵀy − XᵀXw.

\[ X^\top y - X^\top X w^\star = 0 \]

Move the XᵀXw term to the other side

Why: Add XᵀXw★ to both sides. What remains is THE normal equations — a square, solvable linear system for w★.

\[ \boxed{\,X^\top X \, w^\star = X^\top y\,} \]

32. Decode the notation: Substitute r and distribute → the normal equations

Notation

Annotate

From Substitute r and distribute → the normal equations — read this one piece at a time. What is each part doing?

On: \( X^\top (y - X w^\star) = 0 \)

  • The perpendicular-residual condition we just derived.
  • Matrix multiplication distributes over addition, just like numbers: Xᵀ(y − Xw) = Xᵀy − XᵀXw.
  • Add XᵀXw★ to both sides. What remains is THE normal equations — a square, solvable linear system for w★.

33. Something is wrong here: which transpose, which order?

Anomaly

Predict first

A student writes this, and it looks reasonable:

We want a square system from Xw = y, so multiply both sides by X on the left: X·Xw = X·y.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions.

Multiply on the left by Xᵀ, which is (2×5), so the inner dimensions match and the product is square.

Why: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions. The multiply is illegal before it means anything.

34. Trap: which transpose, which order?

Trap

The trap

We want a square system from Xw = y, so multiply both sides by X on the left: X·Xw = X·y.

X X w = X y

Why: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions. The multiply is illegal before it means anything.

The fix

Multiply on the left by Xᵀ, which is (2×5), so the inner dimensions match and the product is square.

Xᵀ X w = Xᵀ y

Why: XᵀX is (2×5)(5×2) = (2×2), square and (here) invertible; Xᵀy is (2×5)(5×1) = (2×1). Memorize the order: Xᵀ goes on the LEFT of BOTH sides.

35. Derivation 2: calculus

Section

Part 3 of 8 — every step

36. Same answer, no geometry needed

Concept

The projection picture is beautiful but you can also get XᵀXw = Xᵀy with pure calculus: L(w) = ‖Xw − y‖² is a smooth scalar function of w, so its minimum sits where the gradient is zero.

To differentiate cleanly, first expand L(w) into matrix terms. That is the next slide — done term by term.

37. Guess the shape of the answer: Expand L(w) = ‖Xw − y‖²

Estimation

Predict first

A squared norm is a vector dotted with itself: ‖v‖² = vᵀv. Set v = Xw − y and FOIL it out, treating (Xw)ᵀ = wᵀXᵀ:

Commit before you compute: what does Expand L(w) = ‖Xw − y‖² come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Sanity check at w = [2.2, 0.6]

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Verified numerically: wᵀXᵀXw = 83.6, 2yᵀXw = 167.2, yᵀy = 86 → L = 83.6 − 167.2 + 86 = 2.4, exactly ‖Xw − y‖².

38. Expand L(w) = ‖Xw − y‖²

Worked example

A squared norm is a vector dotted with itself: ‖v‖² = vᵀv. Set v = Xw − y and FOIL it out, treating (Xw)ᵀ = wᵀXᵀ:

\[ L = (Xw - y)^\top (Xw - y) = w^\top X^\top X w - w^\top X^\top y - y^\top X w + y^\top y \]

The two middle terms are equal scalars

Why: wᵀXᵀy is a 1×1 number, and a number equals its own transpose: (wᵀXᵀy)ᵀ = yᵀXw. So the two cross terms combine into −2 yᵀXw.

\[ L(w) = w^\top X^\top X w \;-\; 2\,y^\top X w \;+\; y^\top y \]

Sanity check at w = [2.2, 0.6]

Why: Verified numerically: wᵀXᵀXw = 83.6, 2yᵀXw = 167.2, yᵀy = 86 → L = 83.6 − 167.2 + 86 = 2.4, exactly ‖Xw − y‖². The expansion is faithful.

39. Draw the shape of it: Expand L(w) = ‖Xw − y‖²

Blank canvas

Draw it

Draw what Expand L(w) = ‖Xw − y‖² just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

40. The two gradient identities we need

Concept

From Lesson 6's matrix-calculus rules, two facts handle every term. For a symmetric A and constant vector b:

\[ \nabla_w\,(w^\top A\, w) = 2 A w, \qquad \nabla_w\,(b^\top w) = b \]

The first is the vector version of d(ax²)/dx = 2ax; the second is the vector version of d(bx)/dx = b. XᵀX is symmetric, so the first applies to it directly.

41. Where does each piece belong: Lesson 7: Linear Systems & the Normal…

Sorting

Sort into buckets

These are the pieces of Lesson 7: Linear Systems & the Normal Equations, out of order. Put each one back under the part of the lesson it belongs to.

The problem: no exact answer
The data: five students; Write the system out in full; Stack it into matrix form Xw = y
Derivation 1: projection
The column space: what Xw can reach; y lives off the plane; Nearest point = perpendicular foot
Derivation 2: calculus
Same answer, no geometry needed; Expand L(w) = ‖Xw − y‖²; The two gradient identities we need
s1
The problem: no exact answer is where Lesson 7: Linear Systems & the Normal Equations puts The data: five students, Write the system out in full, Stack it into matrix form Xw = y. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Derivation 1: projection is where Lesson 7: Linear Systems & the Normal Equations puts The column space: what Xw can reach, y lives off the plane, Nearest point = perpendicular foot. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Derivation 2: calculus is where Lesson 7: Linear Systems & the Normal Equations puts Same answer, no geometry needed, Expand L(w) = ‖Xw − y‖², The two gradient identities we need. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

42. Plan first: Differentiate L term by term

Step zero

Discussion prompt

Differentiate L term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Term 1: ∇(wᵀXᵀX w) = 2XᵀX w

Answer:

  1. Term 1: ∇(wᵀXᵀX w) = 2XᵀX w
  2. Term 2: ∇(−2 yᵀX w) = −2 Xᵀy
  3. Term 3: ∇(yᵀy) = 0
  4. Add them up

43. Differentiate L term by term

Worked example

Term 1: ∇(wᵀXᵀX w) = 2XᵀX w

Why: Apply ∇(wᵀAw) = 2Aw with A = XᵀX (symmetric, since (XᵀX)ᵀ = XᵀX).

\[ \nabla_w\,(w^\top X^\top X w) = 2 X^\top X w \]

Term 2: ∇(−2 yᵀX w) = −2 Xᵀy

Why: Write yᵀXw = (Xᵀy)ᵀw, a bᵀw form with b = Xᵀy. So ∇(bᵀw) = b = Xᵀy, times the −2.

\[ \nabla_w\,(-2\,y^\top X w) = -2 X^\top y \]

Term 3: ∇(yᵀy) = 0

Why: yᵀy has no w in it — it is constant, so its gradient vanishes.

Add them up

Why: Collect the three gradients into the full gradient of L.

\[ \nabla L(w) = 2 X^\top X w - 2 X^\top y = 2\,X^\top(Xw - y) \]

44. Say it in words: Differentiate L term by term

Translation

\( \nabla_w\,(-2\,y^\top X w) = -2 X^\top y \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

45. Complete the line: Set the gradient to zero

Fill the middle

Fill in the blanks

From Set the gradient to zero — finish the line. Write what belongs on the right of the equals sign before you look.

\boxedX^\top y\,}}

Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. The bowl is smooth and opens upward, so its lowest point is the unique place where the gradient vanishes.

46. Set the gradient to zero

Worked example

At the minimum, ∇L = 0

Why: The bowl is smooth and opens upward, so its lowest point is the unique place where the gradient vanishes.

\[ 2\,X^\top(Xw^\star - y) = 0 \]

Divide out the harmless 2

Why: A nonzero scalar factor never changes where an expression equals zero.

\[ X^\top X w^\star - X^\top y = 0 \]

Rearrange

Why: The SAME normal equations we got from geometry — the algebra and the perpendicular-residual picture agree exactly. Notice ∇L = 0 literally says Xᵀ(y − Xw★) = 0, i.e. Xᵀr = 0.

\[ \boxed{\,X^\top X w^\star = X^\top y\,} \]

47. Why it's a minimum, not a saddle

Concept

Zero gradient alone only says 'flat'. We need the second derivative. The Hessian of L is ∇²L = 2XᵀX.

XᵀX is positive semidefinite: for any vector z, zᵀXᵀXz = ‖Xz‖² ≥ 0. A PSD Hessian means L is convex — the surface is a genuine bowl, so the stationary point is the global minimum.

When X has full column rank, XᵀX is strictly positive definite (> 0), the bowl has a unique bottom, and the minimizer is unique.

48. Something is wrong here: the transpose in the matrix derivative

Anomaly

Predict first

A student writes this, and it looks reasonable:

Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform.

Use A = XᵀX (which is what actually appears), and apply ∇(wᵀAw) = 2Aw.

Why: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.

49. Trap: the transpose in the matrix derivative

Trap

The trap

Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.

∇ = 2 X Xᵀ w ✗

Why: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.

The fix

Use A = XᵀX (which is what actually appears), and apply ∇(wᵀAw) = 2Aw.

∇ = 2 XᵀX w ✓

Why: XᵀX is (2×2) and multiplies the 2-vector w correctly. Shapes conform, and it lands on the normal equations. Always read off A from the expression, don't guess it.

50. Break it on purpose: the transpose in the matrix derivative

Break the constraint

Discussion prompt

The rule this trap just fixed:

XᵀX is (2×2) and multiplies the 2-vector w correctly. Shapes conform, and it lands on the normal equations. Always read off A from the expression, don't guess it.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.

51. Solve it by hand

Section

Part 4 of 8 — the 2×2, entry by entry

52. What has to happen first: Build XᵀX, entry by entry

Ranking

Put in order

Put the moves of Build XᵀX, entry by entry into the order they have to happen.

  1. Top-left = ones·ones = Σ1 = 5
  2. Off-diagonal = ones·x = Σxᵢ = 1+2+3+4+5 = 15
  3. Bottom-right = x·x = Σxᵢ² = 1+4+9+16+25 = 55

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Five 1s dotted with five 1s: 1+1+1+1+1 = 5.

53. Build XᵀX, entry by entry

Worked example

(XᵀX)ⱼₖ is column j of X dotted with column k. With column 0 = ones and column 1 = x, there are only three distinct entries (it's symmetric):

Top-left = ones·ones = Σ1 = 5

Why: Five 1s dotted with five 1s: 1+1+1+1+1 = 5. This is just n, the number of data points.

Off-diagonal = ones·x = Σxᵢ = 1+2+3+4+5 = 15

Why: The ones vector dotted with x just sums x. Appears in both off-diagonal slots by symmetry.

Bottom-right = x·x = Σxᵢ² = 1+4+9+16+25 = 55

Why: Sum of squared features.

\[ X^\top X = \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} \]

54. Build Xᵀy, entry by entry

Worked example

Top = ones·y = Σyᵢ = 2+4+5+4+5 = 20

Why: The ones row of Xᵀ dotted with y sums the targets.

Bottom = x·y = Σxᵢyᵢ

Why: The x row of Xᵀ dotted with y. Multiply pairwise then add.

\[ \sum_i x_i y_i = 1\cdot2 + 2\cdot4 + 3\cdot5 + 4\cdot4 + 5\cdot5 = 2+8+15+16+25 = 66 \]

\[ X^\top y = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]

55. Invert the 2×2 — the determinant first

Worked example

We must solve XᵀX w = Xᵀy. For a 2×2, the inverse is 1/det · [[d, −b], [−c, a]] for [[a, b], [c, d]]. First the determinant det = ad − bc:

\[ \det\!\begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} = 5\cdot 55 - 15\cdot 15 = 275 - 225 = 50 \]

det = 50 ≠ 0 → invertible, unique solution

Why: A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique. (If det were 0, we'd be stuck — that's Part 5.)

\[ (X^\top X)^{-1} = \frac{1}{50}\begin{bmatrix} 55 & -15 \\ -15 & 5 \end{bmatrix} = \begin{bmatrix} 1.1 & -0.3 \\ -0.3 & 0.1 \end{bmatrix} \]

56. Multiply through to get w★

Worked example

w★ = (XᵀX)⁻¹ Xᵀy. Multiply the inverse by [20, 66], one entry at a time:

w₀ = 1.1·20 + (−0.3)·66 = 22 − 19.8 = 2.2

Why: First row of the inverse dotted with Xᵀy. Equivalently (55·20 − 15·66)/50 = (1100 − 990)/50 = 110/50 = 2.2.

w₁ = (−0.3)·20 + 0.1·66 = −6 + 6.6 = 0.6

Why: Second row dotted with Xᵀy. Equivalently (−15·20 + 5·66)/50 = (−300 + 330)/50 = 30/50 = 0.6.

\[ w^\star = \begin{bmatrix} 2.2 \\ 0.6 \end{bmatrix} \;\Longrightarrow\; \widehat{\text{score}} = 2.2 + 0.6\,\text{hours} \]

57. Predict the next row: Predictions and residuals — full trace

Pattern

Predict first

The table runs: 1 | 2 | 2.8 | −0.8 | 0.64 · 2 | 4 | 3.4 | +0.6 | 0.36 · 3 | 5 | 4.0 | +1.0 | 1.00 · 4 | 4 | 4.6 | −0.6 | 0.36

In Predictions and residuals — full trace, given the rows so far: what is the next one — the row where x is 5?

Correct: 5 | 5 | 5.2 | −0.2 | 0.04

xyŷ = 2.2 + 0.6xr = y − ŷr²
122.8−0.80.64
243.4+0.60.36
354.0+1.01.00
444.6−0.60.36
555.2−0.20.04

Why: The relationship between the columns, not the individual numbers, is what generates the next row. This is the minimized ‖Xw − y‖² = 2.4 — the smallest total squared error any line can achieve on this data.

58. Predictions and residuals — full trace

Worked example

Plug each x into ŷ = 2.2 + 0.6x, then subtract from the true y to get every residual:

xyŷ = 2.2 + 0.6xr = y − ŷr²
122.8−0.80.64
243.4+0.60.36
354.0+1.01.00
444.6−0.60.36
555.2−0.20.04

Σr² = 0.64 + 0.36 + 1.00 + 0.36 + 0.04 = 2.4

Why: This is the minimized ‖Xw − y‖² = 2.4 — the smallest total squared error any line can achieve on this data. We'll reuse it for R².

59. What each one costs: Predictions and residuals — full trace

Trade off

Comparison matrix

From Predictions and residuals — full trace: every row here is a choice with a cost. Fill the r = y − ŷ column, then say which row you would actually pick and what you give up for it.

xyŷ = 2.2 + 0.6xr = y − ŷr²
122.8−0.80.64
243.4+0.60.36
354.0+1.01.00
444.6−0.60.36
555.2−0.20.04

60. Plan first: Verify the residual is perpendicular

Step zero

Discussion prompt

Verify the residual is perpendicular — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Row 0 (ones·r): −0.8 + 0.6 + 1.0 − 0.6 − 0.2 = 0

Answer:

  1. Row 0 (ones·r): −0.8 + 0.6 + 1.0 − 0.6 − 0.2 = 0
  2. Row 1 (x·r): 1(−0.8) + 2(0.6) + 3(1.0) + 4(−0.6) + 5(−0.2)
  3. Xᵀr = [0, 0] ✓

61. Verify the residual is perpendicular

Worked example

The projection derivation promised Xᵀr = 0. Let's confirm it on the actual residuals [−0.8, 0.6, 1.0, −0.6, −0.2]:

Row 0 (ones·r): −0.8 + 0.6 + 1.0 − 0.6 − 0.2 = 0

Why: The residuals sum to exactly zero — a direct consequence of including the bias column. The fit neither systematically over- nor under-predicts.

Row 1 (x·r): 1(−0.8) + 2(0.6) + 3(1.0) + 4(−0.6) + 5(−0.2)

Why: = −0.8 + 1.2 + 3.0 − 2.4 − 1.0 = 0. The residual is perpendicular to the feature column too.

Xᵀr = [0, 0] ✓

Why: The geometry holds exactly: the error sticks straight out of the column space, confirming w★ is the projection. This is your always-available correctness check.

62. Draw the shape of it: Verify the residual is perpendicular

Blank canvas

Draw it

Draw what Verify the residual is perpendicular just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

63. Something is wrong here: forgetting the bias column

Anomaly

Predict first

A student writes this, and it looks reasonable:

Just regress y on the raw feature x — one column, one weight, ŷ = w₁·x.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With no ones column the line is forced through the ORIGIN (0,0).

Prepend a column of ones so the model can learn an intercept.

Why: With no ones column the line is forced through the ORIGIN (0,0). The best slope-only fit is exactly 1.2, and its total squared error is 6.8 — nearly 3× worse than 2.4, because the data clearly wants a +2.2 intercept.

64. Trap: forgetting the bias column

Trap

The trap

Just regress y on the raw feature x — one column, one weight, ŷ = w₁·x.

X = x as (5,1) → w₁ = Σxy / Σx² = 66/55 = 1.2

Why: With no ones column the line is forced through the ORIGIN (0,0). The best slope-only fit is exactly 1.2, and its total squared error is 6.8 — nearly 3× worse than 2.4, because the data clearly wants a +2.2 intercept.

The fix

Prepend a column of ones so the model can learn an intercept.

X = [ones, x] → ŷ = 2.2 + 0.6x, error 2.4

Why: The bias column lets the line float off the origin to where the data actually sits. The ones column is the single most-forgotten step in a from-scratch OLS.

65. When is XᵀX invertible?

Section

Part 5 of 8 — rank & collinearity

66. Full column rank ⟺ invertible

Concept

XᵀX is invertible exactly when X has full column rank — every column is linearly independent (no column is a combination of the others), which needs at least as many rows as columns.

\[ \operatorname{rank}(X) = d \;\Longleftrightarrow\; X^\top X \text{ is invertible} \]

Reason: XᵀX z = 0 ⇒ zᵀXᵀXz = ‖Xz‖² = 0 ⇒ Xz = 0. If the columns are independent the only such z is 0, so XᵀX has trivial null space — invertible.

67. A collinear column adds no new direction

Intuition

If one feature is a copy or a multiple of another — say income and income measured in cents — the second column points the same way as the first. It adds no new direction to the column space.

Then infinitely many weight vectors produce the identical prediction (shift weight from one twin column to the other), so there is no unique w. det(XᵀX) = 0, and any solver that needs an inverse fails.

68. Predict the next row: A collinear feature breaks solve()

Pattern

Predict first

The table runs: np.linalg.matrix_rank(Xc) | 2 (of 3 columns) · np.linalg.det(Xc.T @ Xc) | 0.0 · np.linalg.cond(Xc.T @ Xc) | ≈ 9.9e33 (effectively infinite)

In A collinear feature breaks solve(), given the rows so far: what is the next one — the row where quantity is np.linalg.solve(...)?

Correct: np.linalg.solve(...) | LinAlgError: Singular matrix

quantityvalue (verified)
np.linalg.matrix_rank(Xc)2 (of 3 columns)
np.linalg.det(Xc.T @ Xc)0.0
np.linalg.cond(Xc.T @ Xc)≈ 9.9e33 (effectively infinite)
np.linalg.solve(...)LinAlgError: Singular matrix

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The redundant column adds no new direction, so the column space is 2-D and XᵀX has a zero eigenvalue.

69. A collinear feature breaks solve()

Worked example

Add a third column equal to 2·x — pure redundancy. Watch XᵀX go singular. This snippet is complete and runnable:

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x]        # col 3 = 2 * col 2
print(np.linalg.matrix_rank(Xc))      # 2  (not 3)
print(np.linalg.det(Xc.T @ Xc))       # 0.0
np.linalg.solve(Xc.T @ Xc, Xc.T @ y)  # LinAlgError: Singular matrix

rank 2 of 3 → det(XᵀX) = 0 → solve raises

Why: The redundant column adds no new direction, so the column space is 2-D and XᵀX has a zero eigenvalue. np.linalg.solve requires a non-singular matrix and raises instead of guessing.

quantityvalue (verified)
np.linalg.matrix_rank(Xc)2 (of 3 columns)
np.linalg.det(Xc.T @ Xc)0.0
np.linalg.cond(Xc.T @ Xc)≈ 9.9e33 (effectively infinite)
np.linalg.solve(...)LinAlgError: Singular matrix

70. Two ways out

Concept

When XᵀX is singular you have two honest options:

  1. Drop the redundant feature(s) so the remaining columns are independent — restores full rank and a unique fit.
  2. Regularize with ridge: add λI to XᵀX, which makes it invertible no matter what — and is the standard move when you can't tell which feature to drop.

Ridge is important enough — and appears often enough on the exam — that it gets its own derivation next.

71. Ridge regression

Section

Part 6 of 8 — derived from the penalty

72. Add an L2 penalty

Concept

Ridge adds a penalty on the size of the weights to the least-squares loss — an L2 term λ‖w‖², where λ ≥ 0 controls the strength:

\[ L_{\text{ridge}}(w) = \lVert Xw - y \rVert^2 + \lambda \lVert w \rVert^2 \]

Big weights now cost something, so the fit is discouraged from blowing up — which is exactly what happens when columns are (nearly) collinear. Let's minimize it the same way we did OLS.

73. What has to happen first: Derive the ridge normal equations

Ranking

Put in order

Put the moves of Derive the ridge normal equations into the order they have to happen.

  1. Gradient of the fit term (from Part 3)
  2. Gradient of the penalty: ∇(λ‖w‖²) = 2λw
  3. Add, set to zero
  4. Divide by 2, collect the w terms

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We already found ∇‖Xw − y‖² = 2Xᵀ(Xw − y).

74. Derive the ridge normal equations

Worked example

Gradient of the fit term (from Part 3)

Why: We already found ∇‖Xw − y‖² = 2Xᵀ(Xw − y). Nothing new here.

\[ \nabla \lVert Xw - y\rVert^2 = 2 X^\top(Xw - y) \]

Gradient of the penalty: ∇(λ‖w‖²) = 2λw

Why: ‖w‖² = wᵀw, and ∇(wᵀw) = 2w by the ∇(wᵀAw) = 2Aw rule with A = I. Times λ gives 2λw.

Add, set to zero

Why: Sum the two gradients and set the total to 0 at the minimum.

\[ 2 X^\top(Xw - y) + 2\lambda w = 0 \]

Divide by 2, collect the w terms

Why: XᵀXw + λw = Xᵀy, and factor w out of the left side using λw = λI w.

\[ \boxed{\,(X^\top X + \lambda I)\, w^\star = X^\top y\,} \]

75. Decode the notation: Derive the ridge normal equations

Notation

Annotate

From Derive the ridge normal equations — read this one piece at a time. What is each part doing?

On: \( \boxed{\,(X^\top X + \lambda I)\, w^\star = X^\top y\,} \)

  • We already found ∇‖Xw − y‖² = 2Xᵀ(Xw − y). Nothing new here.
  • ‖w‖² = wᵀw, and ∇(wᵀw) = 2w by the ∇(wᵀAw) = 2Aw rule with A = I. Times λ gives 2λw.
  • Sum the two gradients and set the total to 0 at the minimum.

76. Why XᵀX + λI is always invertible

Concept

XᵀX is symmetric PSD, so all its eigenvalues are ≥ 0. Adding λI shifts every eigenvalue up by exactly λ, so each becomes ≥ λ > 0. A symmetric matrix with all-positive eigenvalues is invertible — always.

Concretely, our singular collinear XᵀX has eigenvalues {0, 0.90, 279.1}. That zero is what killed solve. Add λ = 1:

\[ \text{eig}(X^\top X + I) = \{\,1,\; 1.90,\; 280.1\,\} \]

The dangerous 0 became 1. Nothing is singular anymore, so the ridge system has a unique solution even though the plain one didn't.

77. What has to be given first: Ridge rescues the collinear fit

Missing information

Discussion prompt

Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The exact same solve that raised LinAlgError now returns a unique weight vector — because λI lifted the zero eigenvalue off the floor.

78. Ridge rescues the collinear fit

Worked example

Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x]
lam = 1.0
wr = np.linalg.solve(Xc.T @ Xc + lam*np.eye(3), Xc.T @ y)
print(wr)   # [1.073446 0.180791 0.361582]

Adding λI = np.eye(3) makes the matrix non-singular

Why: The exact same solve that raised LinAlgError now returns a unique weight vector — because λI lifted the zero eigenvalue off the floor.

matrix passed to solvesolvable?result (verified)
XcᵀXc (collinear)no — singularLinAlgError: Singular matrix
XcᵀXc + 1·I (ridge)yesw ≈ [1.0734, 0.1808, 0.3616]

79. Work backwards from the answer: Ridge rescues the collinear fit

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Adding λI = np.eye(3) makes the matrix non-singular

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:

80. Ridge also shrinks the weights

Concept

Because the penalty charges for weight size, ridge pulls the solution toward 0 — trading a little training fit for stability. On our clean 2-column fit:

λw = [w₀, w₁]‖w‖
0 (plain OLS)[2.2, 0.6]2.28
1 (ridge)[1.171, 0.865]1.46

λ = 1 shrinks ‖w‖ from 2.28 to 1.46. Bigger λ shrinks more; λ → 0 recovers plain OLS. Choosing λ is a bias–variance trade-off you tune on validation data.

81. Something is wrong here: inverting XᵀX directly

Anomaly

Predict first

A student writes this, and it looks reasonable:

The math reads w = (XᵀX)⁻¹Xᵀy, so translate it literally with inv.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior.

Never build the inverse. Solve the system directly, or hand X to the least-squares routine.

Why: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior. And building XᵀX at all already SQUARES the conditioning: cond(X) ≈ 8.4 becomes cond(XᵀX) ≈ 70 — you've amplified sensitivity before inv even runs.

82. Trap: inverting XᵀX directly

Trap

The trap

The math reads w = (XᵀX)⁻¹Xᵀy, so translate it literally with inv.

w = np.linalg.inv(X.T @ X) @ X.T @ y

Why: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior. And building XᵀX at all already SQUARES the conditioning: cond(X) ≈ 8.4 becomes cond(XᵀX) ≈ 70 — you've amplified sensitivity before inv even runs.

The fix

Never build the inverse. Solve the system directly, or hand X to the least-squares routine.

w = np.linalg.solve(X.T @ X, X.T @ y) — or np.linalg.lstsq(X, y)

Why: solve factors the system (LU) with good stability and no explicit inverse. Better still, lstsq works on X via QR/SVD, so it never forms XᵀX and never squares the condition number — and it tolerates rank-deficient X. Same answer [2.2, 0.6], far safer.

83. Which of these survive contact with Lesson 7: Linear Systems & the Normal…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
We have five students. For each we know hours studied x and their score y, and we want to fit a straight line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁.; Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.; The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.
Breaks
We want a square system from Xw = y, so multiply both sides by X on the left: X·Xw = X·y.; Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 7: Linear Systems & the Normal Equations puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

84. Score the fit: R²

Section

Part 7 of 8

85. What R² measures

Concept

R² is the fraction of the target's variance the model explains: 1 is a perfect fit, 0 is no better than always predicting the mean ȳ, and negative means worse than the mean.

\[ R^2 = 1 - \frac{SS_{\text{res}}}{SS_{\text{tot}}} = 1 - \frac{\sum_i (y_i - \hat y_i)^2}{\sum_i (y_i - \bar y)^2} \]

SS_res is the error our line leaves (we already know it: 2.4). SS_tot is the error of the dumb mean-only baseline. The ratio is how much of the baseline error we removed.

86. Teach it back: What R² measures

Explain it

Discussion prompt

Explain What R² measures to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

R² is the fraction of the target's variance the model explains: 1 is a perfect fit, 0 is no better than always predicting the mean ȳ, and negative means worse than the mean.

87. Guess the shape of the answer: R² on our fit, from scratch

Estimation

Predict first

ȳ = 20/5 = 4. Compute both sums of squares, then combine. Runnable:

Commit before you compute: what does R² on our fit, from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: R² = 1 − 2.4/6.0 = 0.60

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The line explains 60% of the variance in scores; the remaining 40% is noise a 2-parameter line can't capture.

88. R² on our fit, from scratch

Worked example

ȳ = 20/5 = 4. Compute both sums of squares, then combine. Runnable:

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y)   # [2.2, 0.6]
yhat = X @ w
ss_res = np.sum((y - yhat)**2)
ss_tot = np.sum((y - y.mean())**2)
print(ss_res, ss_tot, 1 - ss_res/ss_tot)  # 2.4 6.0 0.6

SS_tot = Σ(y − 4)² = 4 + 0 + 1 + 0 + 1 = 6.0

Why: Deviations of y = [2,4,5,4,5] from ȳ = 4 are [−2,0,1,0,1]; squared and summed = 6.0.

R² = 1 − 2.4/6.0 = 0.60

Why: The line explains 60% of the variance in scores; the remaining 40% is noise a 2-parameter line can't capture. Verified by execution.

quantityvalue
SS_res = Σ(y − ŷ)²2.4
SS_tot = Σ(y − ȳ)²6.0
R² = 1 − SS_res/SS_tot0.60

89. Fill in: value for R² on our fit, from scratch

Comparison

Comparison matrix

From R² on our fit, from scratch: refill the value column from what you know. The rest of the table is as it appeared.

quantityvalue
SS_res = Σ(y − ŷ)²2.4
SS_tot = Σ(y − ȳ)²6.0
R² = 1 − SS_res/SS_tot0.60

90. Rebuild the recipe: The OLS recipe

Ranking

Put in order

These are the steps of The OLS recipe, scrambled. Put them back in order before the next slide shows you.

  1. Augment: prepend a ones column → X = [1, features] so the model has an intercept
  2. Assemble: compute XᵀX and Xᵀy
  3. Solve (never invert): w = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)
  4. If XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λI
  5. Score: R² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearn

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

91. The OLS recipe

Pattern

  1. Augment: prepend a ones column → X = [1, features] so the model has an intercept
  2. Assemble: compute XᵀX and Xᵀy
  3. Solve (never invert): w = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)
  4. If XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λI
  5. Score: R² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearn

92. Where does it stop working: The OLS recipe

Edge cases

Discussion prompt

The OLS recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Augment: prepend a ones column → X = [1, features] so the model has an intercept
  2. Assemble: compute XᵀX and Xᵀy
  3. Solve (never invert): w = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)
  4. If XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λI
  5. Score: R² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearn

93. Rule out three: Check yourself — the perpendicular residual

Elimination

Eliminate the wrong options

At the least-squares solution w★, what is true of the residual r = y − Xw★?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. It is perpendicular to every column of X, so Xᵀr = 0
  • B. It is the zero vector — the fit is exact
  • C. It is parallel to y
  • D. It lies inside the column space of X

Survives elimination: A

Why: w★ makes Xw★ the perpendicular projection of y onto the column space, so the leftover r sticks straight out of that space — perpendicular to each column, i.e. Xᵀr = 0. Distributing gives exactly XᵀXw★ = Xᵀy.

94. Check yourself — the perpendicular residual

Check

This is the geometric heart of the whole lesson. Picture the plane.

Check your understanding

At the least-squares solution w★, what is true of the residual r = y − Xw★?

  • A. It is perpendicular to every column of X, so Xᵀr = 0 (correct)
  • B. It is the zero vector — the fit is exact
  • C. It is parallel to y
  • D. It lies inside the column space of X

Answer: A

Why: w★ makes Xw★ the perpendicular projection of y onto the column space, so the leftover r sticks straight out of that space — perpendicular to each column, i.e. Xᵀr = 0. Distributing gives exactly XᵀXw★ = Xᵀy.

Why B tempts people
r = 0 only if y already lies in the column space (a consistent system). Here it doesn't — the minimum error is 2.4, not 0. Least squares handles precisely the case where no exact fit exists.
Why C tempts people
There's no reason for the error to align with the target. r is perpendicular to the column space, not parallel to y.
Why D tempts people
The opposite: r is perpendicular to the column space. If r were inside it you could reduce the error further by moving Xw along r — so you wouldn't be at the minimum.

95. Answer it before you see the options: Check yourself — the normal equations

Prediction

Predict first

Which system gives the least-squares solution w?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: XᵀX w = Xᵀy

Why: Left-multiply Xw = y by Xᵀ: XᵀX is (d×d) and square, Xᵀy is (d×1). Solving XᵀX w = Xᵀy projects y onto X's column space — the normal equations.

96. Check yourself — the normal equations

Check

Track the dimensions before you answer. X is (n×d), n > d.

Check your understanding

Which system gives the least-squares solution w?

  • A. XᵀX w = Xᵀy (correct)
  • B. X Xᵀ w = X y
  • C. X w = y
  • D. XᵀX w = y

Answer: A

Why: Left-multiply Xw = y by Xᵀ: XᵀX is (d×d) and square, Xᵀy is (d×1). Solving XᵀX w = Xᵀy projects y onto X's column space — the normal equations.

Why B tempts people
XXᵀ is (n×n), not a system in the d-dimensional w — wrong shape. (XXᵀ is the Gram matrix of the ROWS, used in kernel methods, not OLS for w.)
Why C tempts people
The original overdetermined system, which has no exact solution when n > d — the entire reason we move to least squares.
Why D tempts people
Mixed shapes: the left side XᵀX w is (d×1) but the right side y is (n×1). You must transform y by Xᵀ too, giving Xᵀy.

97. Answer it before you see the options: Check yourself — invertibility

Prediction

Predict first

Your feature matrix has columns [age, income, 2·income]. What happens at np.linalg.solve(XᵀX, Xᵀy)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: LinAlgError — XᵀX is singular (rank-deficient)

Why: The third column is 2× the second, so rank < number of columns, det(XᵀX) = 0, and solve raises LinAlgError. Fix it with ridge (XᵀX + λI) or by dropping the redundant feature.

98. Check yourself — invertibility

Check

Think about rank.

Check your understanding

Your feature matrix has columns [age, income, 2·income]. What happens at np.linalg.solve(XᵀX, Xᵀy)?

  • A. LinAlgError — XᵀX is singular (rank-deficient) (correct)
  • B. It returns the unique correct weights
  • C. It returns weights, but R² comes out negative
  • D. It silently drops the redundant column

Answer: A

Why: The third column is 2× the second, so rank < number of columns, det(XᵀX) = 0, and solve raises LinAlgError. Fix it with ridge (XᵀX + λI) or by dropping the redundant feature.

Why B tempts people
There is no unique solution when columns are collinear — infinitely many w give the identical fit, so 'the unique weights' don't exist.
Why C tempts people
R² isn't the issue; solve never returns at all because the matrix can't be factored. (Negative R² is real, but it comes from a model worse than the mean, not from singularity.)
Why D tempts people
np.linalg.solve does no feature selection — it requires a non-singular matrix and raises. np.linalg.lstsq tolerates rank deficiency; solve does not.

99. Rule out three: Check yourself — solve vs inv

Elimination

Eliminate the wrong options

Why prefer np.linalg.solve(XᵀX, Xᵀy) over np.linalg.inv(XᵀX) @ Xᵀy?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Forming the explicit inverse is slower and less numerically stable than solving the system directly
  • B. inv gives a different, mathematically wrong answer
  • C. solve adds L2 regularization automatically
  • D. inv only works on symmetric matrices

Survives elimination: A

Why: solve factors the system (LU) once — fewer operations and better backward stability than building the full inverse and multiplying. Best of all is lstsq(X, y), which works on X directly via QR/SVD and never even forms XᵀX (so it never squares the condition number).

100. Check yourself — solve vs inv

Check

The coding section rewards numerically sound habits.

Check your understanding

Why prefer np.linalg.solve(XᵀX, Xᵀy) over np.linalg.inv(XᵀX) @ Xᵀy?

  • A. Forming the explicit inverse is slower and less numerically stable than solving the system directly (correct)
  • B. inv gives a different, mathematically wrong answer
  • C. solve adds L2 regularization automatically
  • D. inv only works on symmetric matrices

Answer: A

Why: solve factors the system (LU) once — fewer operations and better backward stability than building the full inverse and multiplying. Best of all is lstsq(X, y), which works on X directly via QR/SVD and never even forms XᵀX (so it never squares the condition number).

Why B tempts people
In exact arithmetic both give the same w; the difference is floating-point stability and speed, not the underlying math.
Why C tempts people
solve does no regularization — that's ridge's (XᵀX + λI). solve just factors whatever matrix you pass it.
Why D tempts people
inv works on any square non-singular matrix, symmetric or not. XᵀX happens to be symmetric, but that isn't the reason to avoid inv.

101. Your turn: build OLS

Section

Part 8 of 8 — the project

102. Project: OLS from scratch

Concept

Fit score = w₀ + w₁·hours on the 5-student data using only NumPy, then prove your coefficients match sklearn. You've derived every piece — now assemble it yourself.

#requirementtool
1Build X with a bias columnnp.c_[ones, x]
2Solve the normal equations for wnp.linalg.solve
3Score R² and verify vs sklearnLinearRegression

Build rules: type every line yourself, run after each line, and when something errors, check the array shapes — read the error, don't delete it.

103. By analogy: Project: OLS from scratch

Analogy

Discussion prompt

Explain Project: OLS from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Fit score = w₀ + w₁·hours on the 5-student data using only NumPy, then prove your coefficients match sklearn. You've derived every piece — now assemble it yourself.

104. Milestone 1 — the design matrix

Worked example

Your turn: build X from x = [1,2,3,4,5] with a leading ones column. Predict its shape out loud before you print it.

Hint: np.c_[np.ones_like(x), x] stacks columns side by side. The shape should be (5, 2) — 5 rows, 2 columns (bias + feature).

import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
print(X.shape)
print(X)
checkvalue
X.shape(5, 2)
X[:, 0][1, 1, 1, 1, 1] (bias)
X[:, 1][1, 2, 3, 4, 5] (hours)

105. Milestone 2 — solve for w

Worked example

Your turn: assemble XᵀX and Xᵀy, then solve. Say aloud what intercept and slope you expect from the by-hand result before you print.

Hint: X.T @ X and X.T @ y, then np.linalg.solve(XtX, Xty). Do not use inv.

XtX = X.T @ X
Xty = X.T @ y
w = np.linalg.solve(XtX, Xty)
print(XtX)
print(Xty)
print(w)
objectvalue
XᵀX[[5, 15], [15, 55]]
Xᵀy[20, 66]
w = [w₀, w₁][2.2, 0.6]

106. Milestone 3 — score and verify

Worked example

Your turn: compute R² from scratch, then fit sklearn and confirm the coefficients agree. Predict whether they'll match exactly.

Hint: ŷ = X @ w; R² = 1 − Σ(y−ŷ)² / Σ(y−ȳ)². For sklearn, fit on x.reshape(-1, 1) and compare intercept_ and coef_.

yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print(round(r2, 4))

from sklearn.linear_model import LinearRegression
m = LinearRegression().fit(x.reshape(-1, 1), y)
print(round(m.intercept_, 4), round(m.coef_[0], 4))
sourceinterceptslopeR²
your OLS2.20.60.6
sklearn2.20.60.6

107. What each one costs: Milestone 3 — score and verify

Trade off

Comparison matrix

From Milestone 3 — score and verify: every row here is a choice with a cost. Fill the slope column, then say which row you would actually pick and what you give up for it.

sourceinterceptslopeR²
your OLS2.20.60.6
sklearn2.20.60.6

108. The full program

Concept

import numpy as np
from sklearn.linear_model import LinearRegression

x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]           # 1. bias column

w = np.linalg.solve(X.T @ X, X.T @ y)   # 2. solve, never invert
yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print('w =', w, ' R2 =', round(r2, 4)) # 3. score

m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))
printed linevalue
w = [2.2 0.6] R2 =0.6
sklearn:2.2 0.6

If your w reads [2.2, 0.6], R² is 0.6, and sklearn prints the same two numbers — you built linear regression from the math up.

109. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
w = [2.2 0.6] R2 =0.6
sklearn:2.2 0.6

110. Show it off

Concept

Slides closed, out loud: explain (1) why the residual must be perpendicular to X's columns, (2) the two ways to reach XᵀXw = Xᵀy, and (3) what λI does to the eigenvalues of XᵀX.

Stretch: swap in np.linalg.lstsq(X, y) and confirm the same [2.2, 0.6]. Then make a feature collinear and watch solve raise while lstsq and ridge survive — the exact robustness gap the exam likes to probe.

111. Break it if you can: Show it off

Counterexample

Discussion prompt

Slides closed, out loud: explain (1) why the residual must be perpendicular to X's columns, (2) the two ways to reach XᵀXw = Xᵀy, and (3) what λI does to the eigenvalues of XᵀX.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

112. Connect it up: Lesson 7: Linear Systems & the Normal Equations

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The problem: no exact answer · Derivation 1: projection · Derivation 2: calculus · Solve it by hand · When is XᵀX invertible? · Ridge regression. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

113. What you can do now

Recap

movethe one thing to remember
normal equationsXᵀ on the LEFT — XᵀX w = Xᵀy
derivationperpendicular residual = zero gradient
biasprepend a ones column or you fit through the origin
invertiblefull column rank; collinear ⇒ singular ⇒ ridge
solvingsolve / lstsq, never inv; ridge adds λI

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 7 (Week 3 — Linear Systems & Least Squares) — Barron · USAAIO Round 2 Preparation, 2026
  2. NumPy linalg.lstsq / solve
  3. Gilbert Strang, Introduction to Linear Algebra, Ch. 4 (Orthogonality & Projections / Least Squares) — Wellesley-Cambridge Press, 5th ed.
  4. Every trace table, eigenvalue, and coefficient produced by real execution — numpy 2.x + scikit-learn, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108