USAAIO Lesson 7, from Week 3 on linear algebra, fully worked. It writes out overdetermined systems equation by equation, then derives the normal equations XᵀXw = Xᵀy twice - once by projection and once by calculus - with no steps skipped. It solves the 5-student 2×2 system by hand, entry by entry, covers invertibility and rank, derives ridge regression from the penalized loss, and covers R². It ends with a from-scratch OLS implementation verified against sklearn. Every snippet runs as written, and every number was produced by real execution. The lesson runs to 66 slides.
Subject: Machine Learning · 113 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 7 · Week 3 (Linear Algebra)
When there is no exact answer, find the best one. We derive XᵀXw = Xᵀy twice — once by geometry, once by calculus — skipping no step, then solve it by hand and build the OLS that becomes every regression you'll ever fit.
Objectives
Xw = y out in full and say exactly why it has no exact solutionXᵀXw = Xᵀy from the perpendicular residual and independently from ∇‖Xw − y‖² = 02×2 system by hand, entry by entry, and verify the residual is perpendicular to every columnXᵀX is invertible (full column rank) and derive ridge (XᵀX + λI)w = Xᵀy when it is notsolve (never inv), score R², and prove it matches sklearn.LinearRegressionWarm-up
Discussion prompt
Before we open Lesson 7: Linear Systems & the Normal Equations: without looking back, what was the main idea of Gradient Intuition & PyTorch Autograd, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
partial derivatives, the gradient as steepest ascent, directional derivatives, the three gradient identities, and PyTorch autograd — ending with a build-it-yourself gradient descent project.
Section
Part 1 of 8
Concept
We have five students. For each we know hours studied x and their score y, and we want to fit a straight line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁.
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Two unknowns (w₀, w₁), but five data points. Keep that mismatch in mind — it is the whole story of this lesson.
Comparison
Comparison matrix
From The data: five students: refill the hours x column from what you know. The rest of the table is as it appeared.
| student | hours x | score y |
|---|---|---|
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Estimation
Predict first
If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:
Commit before you compute: what does Write the system out in full come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Count: 5 equations, 2 unknowns
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. This is an OVERDETERMINED system — more constraints than degrees of freedom.
Worked example
If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:
\[ \begin{aligned} w_0 + 1\,w_1 &= 2 \\ w_0 + 2\,w_1 &= 4 \\ w_0 + 3\,w_1 &= 5 \\ w_0 + 4\,w_1 &= 4 \\ w_0 + 5\,w_1 &= 5 \end{aligned} \]
Count: 5 equations, 2 unknowns
Why: This is an OVERDETERMINED system — more constraints than degrees of freedom. That is the defining feature of a regression problem.
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Count: 5 equations, 2 unknowns
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
If the line passed exactly through all five points, then w₀ + w₁·xᵢ = yᵢ would hold for every student. Write all five equations — no shorthand:
Missing information
Discussion prompt
Every one of those five equations is w₀·(1) + w₁·(xᵢ). Collect the coefficients into a matrix X, the unknowns into w, the targets into y:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Column 0 multiplies w₀ (the intercept), so it is a column of ones. Column 1 holds the feature x. Miss this column and your line is forced through the origin — we'll hit that trap head-on.
Worked example
Every one of those five equations is w₀·(1) + w₁·(xᵢ). Collect the coefficients into a matrix X, the unknowns into w, the targets into y:
\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \\ 1 & 5 \end{bmatrix}, \qquad w = \begin{bmatrix} w_0 \\ w_1 \end{bmatrix}, \qquad y = \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} \]
The first column of X is all 1s — the bias column
Why: Column 0 multiplies w₀ (the intercept), so it is a column of ones. Column 1 holds the feature x. Miss this column and your line is forced through the origin — we'll hit that trap head-on.
Worked example
Matrix–vector multiply Xw = each row of X dotted with w. Row i is [1, xᵢ], so row i of Xw is 1·w₀ + xᵢ·w₁:
\[ Xw = \begin{bmatrix} w_0 + 1\,w_1 \\ w_0 + 2\,w_1 \\ w_0 + 3\,w_1 \\ w_0 + 4\,w_1 \\ w_0 + 5\,w_1 \end{bmatrix} \;\overset{?}{=}\; \begin{bmatrix} 2 \\ 4 \\ 5 \\ 4 \\ 5 \end{bmatrix} = y \]
Xw = y is exactly the five equations, packed
Why: The matrix equation IS the system — no new content, just compact notation. Asking 'does Xw = y have a solution?' asks 'do all five hold at once?'
Concept
Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.
For example the line through students 1 and 2 (w₀ = 0, w₁ = 2) predicts 6 for student 3, who actually scored 5. No single (w₀, w₁) makes all five equalities true at once.
So Xw = y is inconsistent — it has no exact solution. We stop demanding equality.
Counterexample
Discussion prompt
Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
For example the line through students 1 and 2 (w₀ = 0, w₁ = 2) predicts 6 for student 3, who actually scored 5. No single (w₀, w₁) makes all five equalities true at once.
Intuition
The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.
So exactness is the wrong goal. We don't want the line that hits points — we want the line that is closest to all of them at once.
'Closest' needs a definition of distance. That definition is what turns a hopeless system into a solvable one.
Analogy
Discussion prompt
Explain Real data is noisy — that's the point by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.
Concept
For a candidate w, the prediction is ŷ = Xw and the residual is the miss on each point, rᵢ = yᵢ − ŷᵢ. As a vector:
\[ r = y - Xw \quad\in\; \mathbb{R}^5 \]
r = 0 would mean a perfect fit — impossible here. So instead we make r as small as we can. But 'small' for a 5-vector needs a single number to minimize.
Explain it
Discussion prompt
Explain The residual: how wrong each prediction is to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
For a candidate w, the prediction is ŷ = Xw and the residual is the miss on each point, rᵢ = yᵢ − ŷᵢ. As a vector:
Concept
Collapse the residual vector into one number with the squared Euclidean norm — the sum of squared misses:
\[ \lVert r \rVert^2 = \sum_{i=1}^{5} r_i^2 = r^\top r \]
Squaring makes every miss positive (so overshoots and undershoots both cost), punishes big misses hardest, and — crucially — is smooth, so calculus can find its minimum.
\[ w^\star = \arg\min_{w} \; \lVert Xw - y \rVert^2 \]
Worked example
Substitute the five residuals into ‖Xw − y‖². This is the actual function of (w₀, w₁) we are minimizing — nothing hidden:
\[ L(w_0,w_1) = \sum_{i=1}^{5}\big(w_0 + w_1 x_i - y_i\big)^2 \]
\[ \begin{aligned} =\;& (w_0 + w_1 - 2)^2 + (w_0 + 2w_1 - 4)^2 + (w_0 + 3w_1 - 5)^2 \\ &+\, (w_0 + 4w_1 - 4)^2 + (w_0 + 5w_1 - 5)^2 \end{aligned} \]
L is a sum of squares — a smooth bowl over the (w₀, w₁) plane
Why: Each term is a parabola; a sum of parabolas opens upward with a single lowest point. Our whole job is to find the (w₀, w₁) at the bottom of that bowl.
Section
Part 2 of 8 — the geometry
Concept
Watch what Xw is: Xw = w₀·(column 0) + w₁·(column 1). It is a linear combination of the columns of X.
\[ Xw = w_0 \begin{bmatrix}1\\1\\1\\1\\1\end{bmatrix} + w_1 \begin{bmatrix}1\\2\\3\\4\\5\end{bmatrix} \]
column space of X — The set of ALL vectors you can build as combinations of X's columns — every prediction the model can possibly produce. With 2 columns in ℝ⁵, it is a 2-dimensional plane living inside 5-dimensional space.
Intuition
The target y is a point in ℝ⁵. The column space is a flat 2-D plane through the origin inside that ℝ⁵. Because the system is inconsistent, y does not lie on the plane — it floats above it.
No choice of w can reach y. The best we can do is reach the point on the plane that sits nearest to y.
Picture it
Figure (svg): A point y above a horizontal plane, with a vertical dashed perpendicular dropping to its foot Xw-star on the plane, and a slanted longer line to a different point on the plane.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
The closest point on a plane to an outside point is its perpendicular shadow — the foot of the perpendicular dropped from y onto the plane. That shadow is Xw★, the best prediction.
Concept
The closest point on a plane to an outside point is its perpendicular shadow — the foot of the perpendicular dropped from y onto the plane. That shadow is Xw★, the best prediction.
Figure (svg): A point y above a horizontal plane, with a vertical dashed perpendicular dropping to its foot Xw-star on the plane, and a slanted longer line to a different point on the plane.
Intuition
‖Xw − y‖ is the length of the arrow from the plane-point Xw up to y. Minimizing that length pins Xw to the perpendicular foot.
At that foot, the leftover arrow r = y − Xw★ points straight out of the plane — it is perpendicular to the plane, hence perpendicular to both columns of X.
That single geometric fact — residual ⟂ every column — is enough to pin down w★ algebraically. Next slide turns it into equations.
Concept
Two vectors are perpendicular exactly when their dot product is 0. So 'the residual is perpendicular to each column of X' means each column, dotted with r, gives 0.
The dot of a column with r is that column written as a row (its transpose) times r. So the perpendicularity conditions are (columnⱼ)ᵀ r = 0 for every column j.
Worked example
Column 0 is the ones vector; column 1 is x. Dot each with the residual r = y − Xw and set to zero:
\[ \text{(col 0)}\quad \sum_{i} 1\cdot(y_i - \hat y_i) = 0 \;\Longrightarrow\; \sum_i r_i = 0 \]
\[ \text{(col 1)}\quad \sum_{i} x_i\,(y_i - \hat y_i) = 0 \;\Longrightarrow\; \sum_i x_i r_i = 0 \]
Two columns → two scalar equations → two unknowns
Why: The mismatch is fixed: we traded five impossible equalities for two perpendicularity conditions, exactly matching the two unknowns w₀, w₁.
Worked example
Those two dot products are precisely the two entries of the matrix–vector product Xᵀr. Row 0 of Xᵀ is the ones vector, row 1 is x:
\[ X^\top r = \begin{bmatrix} 1 & 1 & 1 & 1 & 1 \\ 1 & 2 & 3 & 4 & 5 \end{bmatrix} r = \begin{bmatrix} \sum_i r_i \\ \sum_i x_i r_i \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} \]
All perpendicularity conditions collapse to Xᵀr = 0
Why: Multiplying by Xᵀ stacks 'dot with column j = 0' for every column into ONE clean matrix equation. This is the geometric heart of least squares.
Ranking
Put in order
Put the moves of Substitute r and distribute → the normal equations into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The perpendicular-residual condition we just derived.
Worked example
Now put r = y − Xw★ back in and expand. Two lines of algebra, both shown:
Start from Xᵀr = 0
Why: The perpendicular-residual condition we just derived.
\[ X^\top (y - X w^\star) = 0 \]
Distribute Xᵀ across the subtraction
Why: Matrix multiplication distributes over addition, just like numbers: Xᵀ(y − Xw) = Xᵀy − XᵀXw.
\[ X^\top y - X^\top X w^\star = 0 \]
Move the XᵀXw term to the other side
Why: Add XᵀXw★ to both sides. What remains is THE normal equations — a square, solvable linear system for w★.
\[ \boxed{\,X^\top X \, w^\star = X^\top y\,} \]
Notation
Annotate
From Substitute r and distribute → the normal equations — read this one piece at a time. What is each part doing?
On: \( X^\top (y - X w^\star) = 0 \)
Anomaly
Predict first
A student writes this, and it looks reasonable:
We want a square system from Xw = y, so multiply both sides by X on the left: X·Xw = X·y.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions.
Multiply on the left by Xᵀ, which is (2×5), so the inner dimensions match and the product is square.
Why: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions. The multiply is illegal before it means anything.
Trap
We want a square system from Xw = y, so multiply both sides by X on the left: X·Xw = X·y.
X X w = X y
Why: X is (5×2) and non-square. X·X is not even defined — (5×2)(5×2) has mismatched inner dimensions. The multiply is illegal before it means anything.
Multiply on the left by Xᵀ, which is (2×5), so the inner dimensions match and the product is square.
Xᵀ X w = Xᵀ y
Why: XᵀX is (2×5)(5×2) = (2×2), square and (here) invertible; Xᵀy is (2×5)(5×1) = (2×1). Memorize the order: Xᵀ goes on the LEFT of BOTH sides.
Section
Part 3 of 8 — every step
Concept
The projection picture is beautiful but you can also get XᵀXw = Xᵀy with pure calculus: L(w) = ‖Xw − y‖² is a smooth scalar function of w, so its minimum sits where the gradient is zero.
To differentiate cleanly, first expand L(w) into matrix terms. That is the next slide — done term by term.
Estimation
Predict first
A squared norm is a vector dotted with itself: ‖v‖² = vᵀv. Set v = Xw − y and FOIL it out, treating (Xw)ᵀ = wᵀXᵀ:
Commit before you compute: what does Expand L(w) = ‖Xw − y‖² come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Sanity check at w = [2.2, 0.6]
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Verified numerically: wᵀXᵀXw = 83.6, 2yᵀXw = 167.2, yᵀy = 86 → L = 83.6 − 167.2 + 86 = 2.4, exactly ‖Xw − y‖².
Worked example
A squared norm is a vector dotted with itself: ‖v‖² = vᵀv. Set v = Xw − y and FOIL it out, treating (Xw)ᵀ = wᵀXᵀ:
\[ L = (Xw - y)^\top (Xw - y) = w^\top X^\top X w - w^\top X^\top y - y^\top X w + y^\top y \]
The two middle terms are equal scalars
Why: wᵀXᵀy is a 1×1 number, and a number equals its own transpose: (wᵀXᵀy)ᵀ = yᵀXw. So the two cross terms combine into −2 yᵀXw.
\[ L(w) = w^\top X^\top X w \;-\; 2\,y^\top X w \;+\; y^\top y \]
Sanity check at w = [2.2, 0.6]
Why: Verified numerically: wᵀXᵀXw = 83.6, 2yᵀXw = 167.2, yᵀy = 86 → L = 83.6 − 167.2 + 86 = 2.4, exactly ‖Xw − y‖². The expansion is faithful.
Blank canvas
Draw it
Draw what Expand L(w) = ‖Xw − y‖² just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Concept
From Lesson 6's matrix-calculus rules, two facts handle every term. For a symmetric A and constant vector b:
\[ \nabla_w\,(w^\top A\, w) = 2 A w, \qquad \nabla_w\,(b^\top w) = b \]
The first is the vector version of d(ax²)/dx = 2ax; the second is the vector version of d(bx)/dx = b. XᵀX is symmetric, so the first applies to it directly.
Sorting
Sort into buckets
These are the pieces of Lesson 7: Linear Systems & the Normal Equations, out of order. Put each one back under the part of the lesson it belongs to.
Step zero
Discussion prompt
Differentiate L term by term — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Term 1: ∇(wᵀXᵀX w) = 2XᵀX w
Answer:
Worked example
Term 1: ∇(wᵀXᵀX w) = 2XᵀX w
Why: Apply ∇(wᵀAw) = 2Aw with A = XᵀX (symmetric, since (XᵀX)ᵀ = XᵀX).
\[ \nabla_w\,(w^\top X^\top X w) = 2 X^\top X w \]
Term 2: ∇(−2 yᵀX w) = −2 Xᵀy
Why: Write yᵀXw = (Xᵀy)ᵀw, a bᵀw form with b = Xᵀy. So ∇(bᵀw) = b = Xᵀy, times the −2.
\[ \nabla_w\,(-2\,y^\top X w) = -2 X^\top y \]
Term 3: ∇(yᵀy) = 0
Why: yᵀy has no w in it — it is constant, so its gradient vanishes.
Add them up
Why: Collect the three gradients into the full gradient of L.
\[ \nabla L(w) = 2 X^\top X w - 2 X^\top y = 2\,X^\top(Xw - y) \]
Translation
\( \nabla_w\,(-2\,y^\top X w) = -2 X^\top y \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Fill the middle
Fill in the blanks
From Set the gradient to zero — finish the line. Write what belongs on the right of the equals sign before you look.
\boxedX^\top y\,}}
Why: Producing the right-hand side unprompted is the difference between recognising this line and being able to use it. The bowl is smooth and opens upward, so its lowest point is the unique place where the gradient vanishes.
Worked example
At the minimum, ∇L = 0
Why: The bowl is smooth and opens upward, so its lowest point is the unique place where the gradient vanishes.
\[ 2\,X^\top(Xw^\star - y) = 0 \]
Divide out the harmless 2
Why: A nonzero scalar factor never changes where an expression equals zero.
\[ X^\top X w^\star - X^\top y = 0 \]
Rearrange
Why: The SAME normal equations we got from geometry — the algebra and the perpendicular-residual picture agree exactly. Notice ∇L = 0 literally says Xᵀ(y − Xw★) = 0, i.e. Xᵀr = 0.
\[ \boxed{\,X^\top X w^\star = X^\top y\,} \]
Concept
Zero gradient alone only says 'flat'. We need the second derivative. The Hessian of L is ∇²L = 2XᵀX.
XᵀX is positive semidefinite: for any vector z, zᵀXᵀXz = ‖Xz‖² ≥ 0. A PSD Hessian means L is convex — the surface is a genuine bowl, so the stationary point is the global minimum.
When X has full column rank, XᵀX is strictly positive definite (> 0), the bowl has a unique bottom, and the minimizer is unique.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform.
Use A = XᵀX (which is what actually appears), and apply ∇(wᵀAw) = 2Aw.
Why: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.
Trap
Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.
∇ = 2 X Xᵀ w ✗
Why: XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.
Use A = XᵀX (which is what actually appears), and apply ∇(wᵀAw) = 2Aw.
∇ = 2 XᵀX w ✓
Why: XᵀX is (2×2) and multiplies the 2-vector w correctly. Shapes conform, and it lands on the normal equations. Always read off A from the expression, don't guess it.
Break the constraint
Discussion prompt
The rule this trap just fixed:
XᵀX is (2×2) and multiplies the 2-vector w correctly. Shapes conform, and it lands on the normal equations. Always read off A from the expression, don't guess it.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
XXᵀ is (5×5) — it acts on 5-vectors, not on the 2-vector w, so the shapes don't even conform. The identity is ∇(wᵀAw) = 2Aw with A = XᵀX, not XXᵀ.
Section
Part 4 of 8 — the 2×2, entry by entry
Ranking
Put in order
Put the moves of Build XᵀX, entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Five 1s dotted with five 1s: 1+1+1+1+1 = 5.
Worked example
(XᵀX)ⱼₖ is column j of X dotted with column k. With column 0 = ones and column 1 = x, there are only three distinct entries (it's symmetric):
Top-left = ones·ones = Σ1 = 5
Why: Five 1s dotted with five 1s: 1+1+1+1+1 = 5. This is just n, the number of data points.
Off-diagonal = ones·x = Σxᵢ = 1+2+3+4+5 = 15
Why: The ones vector dotted with x just sums x. Appears in both off-diagonal slots by symmetry.
Bottom-right = x·x = Σxᵢ² = 1+4+9+16+25 = 55
Why: Sum of squared features.
\[ X^\top X = \begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} \]
Worked example
Top = ones·y = Σyᵢ = 2+4+5+4+5 = 20
Why: The ones row of Xᵀ dotted with y sums the targets.
Bottom = x·y = Σxᵢyᵢ
Why: The x row of Xᵀ dotted with y. Multiply pairwise then add.
\[ \sum_i x_i y_i = 1\cdot2 + 2\cdot4 + 3\cdot5 + 4\cdot4 + 5\cdot5 = 2+8+15+16+25 = 66 \]
\[ X^\top y = \begin{bmatrix} 20 \\ 66 \end{bmatrix} \]
Worked example
We must solve XᵀX w = Xᵀy. For a 2×2, the inverse is 1/det · [[d, −b], [−c, a]] for [[a, b], [c, d]]. First the determinant det = ad − bc:
\[ \det\!\begin{bmatrix} 5 & 15 \\ 15 & 55 \end{bmatrix} = 5\cdot 55 - 15\cdot 15 = 275 - 225 = 50 \]
det = 50 ≠ 0 → invertible, unique solution
Why: A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique. (If det were 0, we'd be stuck — that's Part 5.)
\[ (X^\top X)^{-1} = \frac{1}{50}\begin{bmatrix} 55 & -15 \\ -15 & 5 \end{bmatrix} = \begin{bmatrix} 1.1 & -0.3 \\ -0.3 & 0.1 \end{bmatrix} \]
Worked example
w★ = (XᵀX)⁻¹ Xᵀy. Multiply the inverse by [20, 66], one entry at a time:
w₀ = 1.1·20 + (−0.3)·66 = 22 − 19.8 = 2.2
Why: First row of the inverse dotted with Xᵀy. Equivalently (55·20 − 15·66)/50 = (1100 − 990)/50 = 110/50 = 2.2.
w₁ = (−0.3)·20 + 0.1·66 = −6 + 6.6 = 0.6
Why: Second row dotted with Xᵀy. Equivalently (−15·20 + 5·66)/50 = (−300 + 330)/50 = 30/50 = 0.6.
\[ w^\star = \begin{bmatrix} 2.2 \\ 0.6 \end{bmatrix} \;\Longrightarrow\; \widehat{\text{score}} = 2.2 + 0.6\,\text{hours} \]
Pattern
Predict first
The table runs: 1 | 2 | 2.8 | −0.8 | 0.64 · 2 | 4 | 3.4 | +0.6 | 0.36 · 3 | 5 | 4.0 | +1.0 | 1.00 · 4 | 4 | 4.6 | −0.6 | 0.36
In Predictions and residuals — full trace, given the rows so far: what is the next one — the row where x is 5?
Correct: 5 | 5 | 5.2 | −0.2 | 0.04
| x | y | ŷ = 2.2 + 0.6x | r = y − ŷ | r² |
|---|---|---|---|---|
| 1 | 2 | 2.8 | −0.8 | 0.64 |
| 2 | 4 | 3.4 | +0.6 | 0.36 |
| 3 | 5 | 4.0 | +1.0 | 1.00 |
| 4 | 4 | 4.6 | −0.6 | 0.36 |
| 5 | 5 | 5.2 | −0.2 | 0.04 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. This is the minimized ‖Xw − y‖² = 2.4 — the smallest total squared error any line can achieve on this data.
Worked example
Plug each x into ŷ = 2.2 + 0.6x, then subtract from the true y to get every residual:
| x | y | ŷ = 2.2 + 0.6x | r = y − ŷ | r² |
|---|---|---|---|---|
| 1 | 2 | 2.8 | −0.8 | 0.64 |
| 2 | 4 | 3.4 | +0.6 | 0.36 |
| 3 | 5 | 4.0 | +1.0 | 1.00 |
| 4 | 4 | 4.6 | −0.6 | 0.36 |
| 5 | 5 | 5.2 | −0.2 | 0.04 |
Σr² = 0.64 + 0.36 + 1.00 + 0.36 + 0.04 = 2.4
Why: This is the minimized ‖Xw − y‖² = 2.4 — the smallest total squared error any line can achieve on this data. We'll reuse it for R².
Trade off
Comparison matrix
From Predictions and residuals — full trace: every row here is a choice with a cost. Fill the r = y − ŷ column, then say which row you would actually pick and what you give up for it.
| x | y | ŷ = 2.2 + 0.6x | r = y − ŷ | r² |
|---|---|---|---|---|
| 1 | 2 | 2.8 | −0.8 | 0.64 |
| 2 | 4 | 3.4 | +0.6 | 0.36 |
| 3 | 5 | 4.0 | +1.0 | 1.00 |
| 4 | 4 | 4.6 | −0.6 | 0.36 |
| 5 | 5 | 5.2 | −0.2 | 0.04 |
Step zero
Discussion prompt
Verify the residual is perpendicular — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Row 0 (ones·r): −0.8 + 0.6 + 1.0 − 0.6 − 0.2 = 0
Answer:
Worked example
The projection derivation promised Xᵀr = 0. Let's confirm it on the actual residuals [−0.8, 0.6, 1.0, −0.6, −0.2]:
Row 0 (ones·r): −0.8 + 0.6 + 1.0 − 0.6 − 0.2 = 0
Why: The residuals sum to exactly zero — a direct consequence of including the bias column. The fit neither systematically over- nor under-predicts.
Row 1 (x·r): 1(−0.8) + 2(0.6) + 3(1.0) + 4(−0.6) + 5(−0.2)
Why: = −0.8 + 1.2 + 3.0 − 2.4 − 1.0 = 0. The residual is perpendicular to the feature column too.
Xᵀr = [0, 0] ✓
Why: The geometry holds exactly: the error sticks straight out of the column space, confirming w★ is the projection. This is your always-available correctness check.
Blank canvas
Draw it
Draw what Verify the residual is perpendicular just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Just regress y on the raw feature x — one column, one weight, ŷ = w₁·x.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With no ones column the line is forced through the ORIGIN (0,0).
Prepend a column of ones so the model can learn an intercept.
Why: With no ones column the line is forced through the ORIGIN (0,0). The best slope-only fit is exactly 1.2, and its total squared error is 6.8 — nearly 3× worse than 2.4, because the data clearly wants a +2.2 intercept.
Trap
Just regress y on the raw feature x — one column, one weight, ŷ = w₁·x.
X = x as (5,1) → w₁ = Σxy / Σx² = 66/55 = 1.2
Why: With no ones column the line is forced through the ORIGIN (0,0). The best slope-only fit is exactly 1.2, and its total squared error is 6.8 — nearly 3× worse than 2.4, because the data clearly wants a +2.2 intercept.
Prepend a column of ones so the model can learn an intercept.
X = [ones, x] → ŷ = 2.2 + 0.6x, error 2.4
Why: The bias column lets the line float off the origin to where the data actually sits. The ones column is the single most-forgotten step in a from-scratch OLS.
Section
Part 5 of 8 — rank & collinearity
Concept
XᵀX is invertible exactly when X has full column rank — every column is linearly independent (no column is a combination of the others), which needs at least as many rows as columns.
\[ \operatorname{rank}(X) = d \;\Longleftrightarrow\; X^\top X \text{ is invertible} \]
Reason: XᵀX z = 0 ⇒ zᵀXᵀXz = ‖Xz‖² = 0 ⇒ Xz = 0. If the columns are independent the only such z is 0, so XᵀX has trivial null space — invertible.
Intuition
If one feature is a copy or a multiple of another — say income and income measured in cents — the second column points the same way as the first. It adds no new direction to the column space.
Then infinitely many weight vectors produce the identical prediction (shift weight from one twin column to the other), so there is no unique w. det(XᵀX) = 0, and any solver that needs an inverse fails.
Pattern
Predict first
The table runs: np.linalg.matrix_rank(Xc) | 2 (of 3 columns) · np.linalg.det(Xc.T @ Xc) | 0.0 · np.linalg.cond(Xc.T @ Xc) | ≈ 9.9e33 (effectively infinite)
In A collinear feature breaks solve(), given the rows so far: what is the next one — the row where quantity is np.linalg.solve(...)?
Correct: np.linalg.solve(...) | LinAlgError: Singular matrix
| quantity | value (verified) |
|---|---|
| np.linalg.matrix_rank(Xc) | 2 (of 3 columns) |
| np.linalg.det(Xc.T @ Xc) | 0.0 |
| np.linalg.cond(Xc.T @ Xc) | ≈ 9.9e33 (effectively infinite) |
| np.linalg.solve(...) | LinAlgError: Singular matrix |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The redundant column adds no new direction, so the column space is 2-D and XᵀX has a zero eigenvalue.
Worked example
Add a third column equal to 2·x — pure redundancy. Watch XᵀX go singular. This snippet is complete and runnable:
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x] # col 3 = 2 * col 2
print(np.linalg.matrix_rank(Xc)) # 2 (not 3)
print(np.linalg.det(Xc.T @ Xc)) # 0.0
np.linalg.solve(Xc.T @ Xc, Xc.T @ y) # LinAlgError: Singular matrixrank 2 of 3 → det(XᵀX) = 0 → solve raises
Why: The redundant column adds no new direction, so the column space is 2-D and XᵀX has a zero eigenvalue. np.linalg.solve requires a non-singular matrix and raises instead of guessing.
| quantity | value (verified) |
|---|---|
| np.linalg.matrix_rank(Xc) | 2 (of 3 columns) |
| np.linalg.det(Xc.T @ Xc) | 0.0 |
| np.linalg.cond(Xc.T @ Xc) | ≈ 9.9e33 (effectively infinite) |
| np.linalg.solve(...) | LinAlgError: Singular matrix |
Concept
When XᵀX is singular you have two honest options:
λI to XᵀX, which makes it invertible no matter what — and is the standard move when you can't tell which feature to drop.Ridge is important enough — and appears often enough on the exam — that it gets its own derivation next.
Section
Part 6 of 8 — derived from the penalty
Concept
Ridge adds a penalty on the size of the weights to the least-squares loss — an L2 term λ‖w‖², where λ ≥ 0 controls the strength:
\[ L_{\text{ridge}}(w) = \lVert Xw - y \rVert^2 + \lambda \lVert w \rVert^2 \]
Big weights now cost something, so the fit is discouraged from blowing up — which is exactly what happens when columns are (nearly) collinear. Let's minimize it the same way we did OLS.
Ranking
Put in order
Put the moves of Derive the ridge normal equations into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. We already found ∇‖Xw − y‖² = 2Xᵀ(Xw − y).
Worked example
Gradient of the fit term (from Part 3)
Why: We already found ∇‖Xw − y‖² = 2Xᵀ(Xw − y). Nothing new here.
\[ \nabla \lVert Xw - y\rVert^2 = 2 X^\top(Xw - y) \]
Gradient of the penalty: ∇(λ‖w‖²) = 2λw
Why: ‖w‖² = wᵀw, and ∇(wᵀw) = 2w by the ∇(wᵀAw) = 2Aw rule with A = I. Times λ gives 2λw.
Add, set to zero
Why: Sum the two gradients and set the total to 0 at the minimum.
\[ 2 X^\top(Xw - y) + 2\lambda w = 0 \]
Divide by 2, collect the w terms
Why: XᵀXw + λw = Xᵀy, and factor w out of the left side using λw = λI w.
\[ \boxed{\,(X^\top X + \lambda I)\, w^\star = X^\top y\,} \]
Notation
Annotate
From Derive the ridge normal equations — read this one piece at a time. What is each part doing?
On: \( \boxed{\,(X^\top X + \lambda I)\, w^\star = X^\top y\,} \)
Concept
XᵀX is symmetric PSD, so all its eigenvalues are ≥ 0. Adding λI shifts every eigenvalue up by exactly λ, so each becomes ≥ λ > 0. A symmetric matrix with all-positive eigenvalues is invertible — always.
Concretely, our singular collinear XᵀX has eigenvalues {0, 0.90, 279.1}. That zero is what killed solve. Add λ = 1:
\[ \text{eig}(X^\top X + I) = \{\,1,\; 1.90,\; 280.1\,\} \]
The dangerous 0 became 1. Nothing is singular anymore, so the ridge system has a unique solution even though the plain one didn't.
Missing information
Discussion prompt
Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The exact same solve that raised LinAlgError now returns a unique weight vector — because λI lifted the zero eigenvalue off the floor.
Worked example
Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
Xc = np.c_[np.ones(5), x, 2*x]
lam = 1.0
wr = np.linalg.solve(Xc.T @ Xc + lam*np.eye(3), Xc.T @ y)
print(wr) # [1.073446 0.180791 0.361582]Adding λI = np.eye(3) makes the matrix non-singular
Why: The exact same solve that raised LinAlgError now returns a unique weight vector — because λI lifted the zero eigenvalue off the floor.
| matrix passed to solve | solvable? | result (verified) |
|---|---|---|
| XcᵀXc (collinear) | no — singular | LinAlgError: Singular matrix |
| XcᵀXc + 1·I (ridge) | yes | w ≈ [1.0734, 0.1808, 0.3616] |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Adding λI = np.eye(3) makes the matrix non-singular
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same collinear Xc that made solve raise. Add λI and it solves cleanly — runnable as written:
Concept
Because the penalty charges for weight size, ridge pulls the solution toward 0 — trading a little training fit for stability. On our clean 2-column fit:
| λ | w = [w₀, w₁] | ‖w‖ |
|---|---|---|
| 0 (plain OLS) | [2.2, 0.6] | 2.28 |
| 1 (ridge) | [1.171, 0.865] | 1.46 |
λ = 1 shrinks ‖w‖ from 2.28 to 1.46. Bigger λ shrinks more; λ → 0 recovers plain OLS. Choosing λ is a bias–variance trade-off you tune on validation data.
Anomaly
Predict first
A student writes this, and it looks reasonable:
The math reads w = (XᵀX)⁻¹Xᵀy, so translate it literally with inv.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior.
Never build the inverse. Solve the system directly, or hand X to the least-squares routine.
Why: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior. And building XᵀX at all already SQUARES the conditioning: cond(X) ≈ 8.4 becomes cond(XᵀX) ≈ 70 — you've amplified sensitivity before inv even runs.
Trap
The math reads w = (XᵀX)⁻¹Xᵀy, so translate it literally with inv.
w = np.linalg.inv(X.T @ X) @ X.T @ y
Why: Forming the explicit inverse is ~3× the work of solving and has worse round-off behavior. And building XᵀX at all already SQUARES the conditioning: cond(X) ≈ 8.4 becomes cond(XᵀX) ≈ 70 — you've amplified sensitivity before inv even runs.
Never build the inverse. Solve the system directly, or hand X to the least-squares routine.
w = np.linalg.solve(X.T @ X, X.T @ y) — or np.linalg.lstsq(X, y)
Why: solve factors the system (LU) with good stability and no explicit inverse. Better still, lstsq works on X via QR/SVD, so it never forms XᵀX and never squares the condition number — and it tolerates rank-deficient X. Same answer [2.2, 0.6], far safer.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x and their score y, and we want to fit a straight line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁.; Two unknowns can satisfy at most two independent equations. Pick any two students and a unique line fits them — but it will generally miss the other three.; The scores don't lie on a line because real outcomes carry noise: mood, luck, a hard question. The 'true' relationship is roughly linear, but each point is nudged off it.Xw = y, so multiply both sides by X on the left: X·Xw = X·y.; Differentiating wᵀXᵀX w, treat it like 2·(XᵀX)·w... but drop the transpose and write the gradient as 2X Xᵀ w.Section
Part 7 of 8
Concept
R² is the fraction of the target's variance the model explains: 1 is a perfect fit, 0 is no better than always predicting the mean ȳ, and negative means worse than the mean.
\[ R^2 = 1 - \frac{SS_{\text{res}}}{SS_{\text{tot}}} = 1 - \frac{\sum_i (y_i - \hat y_i)^2}{\sum_i (y_i - \bar y)^2} \]
SS_res is the error our line leaves (we already know it: 2.4). SS_tot is the error of the dumb mean-only baseline. The ratio is how much of the baseline error we removed.
Explain it
Discussion prompt
Explain What R² measures to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
R² is the fraction of the target's variance the model explains: 1 is a perfect fit, 0 is no better than always predicting the mean ȳ, and negative means worse than the mean.
Estimation
Predict first
ȳ = 20/5 = 4. Compute both sums of squares, then combine. Runnable:
Commit before you compute: what does R² on our fit, from scratch come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: R² = 1 − 2.4/6.0 = 0.60
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The line explains 60% of the variance in scores; the remaining 40% is noise a 2-parameter line can't capture.
Worked example
ȳ = 20/5 = 4. Compute both sums of squares, then combine. Runnable:
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones(5), x]
w = np.linalg.solve(X.T @ X, X.T @ y) # [2.2, 0.6]
yhat = X @ w
ss_res = np.sum((y - yhat)**2)
ss_tot = np.sum((y - y.mean())**2)
print(ss_res, ss_tot, 1 - ss_res/ss_tot) # 2.4 6.0 0.6SS_tot = Σ(y − 4)² = 4 + 0 + 1 + 0 + 1 = 6.0
Why: Deviations of y = [2,4,5,4,5] from ȳ = 4 are [−2,0,1,0,1]; squared and summed = 6.0.
R² = 1 − 2.4/6.0 = 0.60
Why: The line explains 60% of the variance in scores; the remaining 40% is noise a 2-parameter line can't capture. Verified by execution.
| quantity | value |
|---|---|
| SS_res = Σ(y − ŷ)² | 2.4 |
| SS_tot = Σ(y − ȳ)² | 6.0 |
| R² = 1 − SS_res/SS_tot | 0.60 |
Comparison
Comparison matrix
From R² on our fit, from scratch: refill the value column from what you know. The rest of the table is as it appeared.
| quantity | value |
|---|---|
| SS_res = Σ(y − ŷ)² | 2.4 |
| SS_tot = Σ(y − ȳ)² | 6.0 |
| R² = 1 − SS_res/SS_tot | 0.60 |
Ranking
Put in order
These are the steps of The OLS recipe, scrambled. Put them back in order before the next slide shows you.
X = [1, features] so the model has an interceptXᵀX and Xᵀyw = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λIR² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearnWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
X = [1, features] so the model has an interceptXᵀX and Xᵀyw = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λIR² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearnEdge cases
Discussion prompt
The OLS recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
X = [1, features] so the model has an interceptXᵀX and Xᵀyw = np.linalg.solve(XᵀX, Xᵀy) — or np.linalg.lstsq(X, y)XᵀX is singular (collinear features): drop a feature, or switch to ridge XᵀX + λIR² = 1 − SS_res/SS_tot, and sanity-check the coefficients against sklearnElimination
Eliminate the wrong options
At the least-squares solution w★, what is true of the residual r = y − Xw★?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: w★ makes Xw★ the perpendicular projection of y onto the column space, so the leftover r sticks straight out of that space — perpendicular to each column, i.e. Xᵀr = 0. Distributing gives exactly XᵀXw★ = Xᵀy.
Check
This is the geometric heart of the whole lesson. Picture the plane.
Check your understanding
At the least-squares solution w★, what is true of the residual r = y − Xw★?
Answer: A
Why: w★ makes Xw★ the perpendicular projection of y onto the column space, so the leftover r sticks straight out of that space — perpendicular to each column, i.e. Xᵀr = 0. Distributing gives exactly XᵀXw★ = Xᵀy.
Prediction
Predict first
Which system gives the least-squares solution w?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: XᵀX w = Xᵀy
Why: Left-multiply Xw = y by Xᵀ: XᵀX is (d×d) and square, Xᵀy is (d×1). Solving XᵀX w = Xᵀy projects y onto X's column space — the normal equations.
Check
Track the dimensions before you answer. X is (n×d), n > d.
Check your understanding
Which system gives the least-squares solution w?
Answer: A
Why: Left-multiply Xw = y by Xᵀ: XᵀX is (d×d) and square, Xᵀy is (d×1). Solving XᵀX w = Xᵀy projects y onto X's column space — the normal equations.
Prediction
Predict first
Your feature matrix has columns [age, income, 2·income]. What happens at np.linalg.solve(XᵀX, Xᵀy)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: LinAlgError — XᵀX is singular (rank-deficient)
Why: The third column is 2× the second, so rank < number of columns, det(XᵀX) = 0, and solve raises LinAlgError. Fix it with ridge (XᵀX + λI) or by dropping the redundant feature.
Check
Think about rank.
Check your understanding
Your feature matrix has columns [age, income, 2·income]. What happens at np.linalg.solve(XᵀX, Xᵀy)?
Answer: A
Why: The third column is 2× the second, so rank < number of columns, det(XᵀX) = 0, and solve raises LinAlgError. Fix it with ridge (XᵀX + λI) or by dropping the redundant feature.
Elimination
Eliminate the wrong options
Why prefer np.linalg.solve(XᵀX, Xᵀy) over np.linalg.inv(XᵀX) @ Xᵀy?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: solve factors the system (LU) once — fewer operations and better backward stability than building the full inverse and multiplying. Best of all is lstsq(X, y), which works on X directly via QR/SVD and never even forms XᵀX (so it never squares the condition number).
Check
The coding section rewards numerically sound habits.
Check your understanding
Why prefer np.linalg.solve(XᵀX, Xᵀy) over np.linalg.inv(XᵀX) @ Xᵀy?
Answer: A
Why: solve factors the system (LU) once — fewer operations and better backward stability than building the full inverse and multiplying. Best of all is lstsq(X, y), which works on X directly via QR/SVD and never even forms XᵀX (so it never squares the condition number).
Section
Part 8 of 8 — the project
Concept
Fit score = w₀ + w₁·hours on the 5-student data using only NumPy, then prove your coefficients match sklearn. You've derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | Build X with a bias column | np.c_[ones, x] |
| 2 | Solve the normal equations for w | np.linalg.solve |
| 3 | Score R² and verify vs sklearn | LinearRegression |
Build rules: type every line yourself, run after each line, and when something errors, check the array shapes — read the error, don't delete it.
Analogy
Discussion prompt
Explain Project: OLS from scratch by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Fit score = w₀ + w₁·hours on the 5-student data using only NumPy, then prove your coefficients match sklearn. You've derived every piece — now assemble it yourself.
Worked example
Your turn: build X from x = [1,2,3,4,5] with a leading ones column. Predict its shape out loud before you print it.
Hint: np.c_[np.ones_like(x), x] stacks columns side by side. The shape should be (5, 2) — 5 rows, 2 columns (bias + feature).
import numpy as np
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x]
print(X.shape)
print(X)| check | value |
|---|---|
| X.shape | (5, 2) |
| X[:, 0] | [1, 1, 1, 1, 1] (bias) |
| X[:, 1] | [1, 2, 3, 4, 5] (hours) |
Worked example
Your turn: assemble XᵀX and Xᵀy, then solve. Say aloud what intercept and slope you expect from the by-hand result before you print.
Hint: X.T @ X and X.T @ y, then np.linalg.solve(XtX, Xty). Do not use inv.
XtX = X.T @ X
Xty = X.T @ y
w = np.linalg.solve(XtX, Xty)
print(XtX)
print(Xty)
print(w)| object | value |
|---|---|
| XᵀX | [[5, 15], [15, 55]] |
| Xᵀy | [20, 66] |
| w = [w₀, w₁] | [2.2, 0.6] |
Worked example
Your turn: compute R² from scratch, then fit sklearn and confirm the coefficients agree. Predict whether they'll match exactly.
Hint: ŷ = X @ w; R² = 1 − Σ(y−ŷ)² / Σ(y−ȳ)². For sklearn, fit on x.reshape(-1, 1) and compare intercept_ and coef_.
yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print(round(r2, 4))
from sklearn.linear_model import LinearRegression
m = LinearRegression().fit(x.reshape(-1, 1), y)
print(round(m.intercept_, 4), round(m.coef_[0], 4))| source | intercept | slope | R² |
|---|---|---|---|
| your OLS | 2.2 | 0.6 | 0.6 |
| sklearn | 2.2 | 0.6 | 0.6 |
Trade off
Comparison matrix
From Milestone 3 — score and verify: every row here is a choice with a cost. Fill the slope column, then say which row you would actually pick and what you give up for it.
| source | intercept | slope | R² |
|---|---|---|---|
| your OLS | 2.2 | 0.6 | 0.6 |
| sklearn | 2.2 | 0.6 | 0.6 |
Concept
import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4., 5.])
y = np.array([2., 4., 5., 4., 5.])
X = np.c_[np.ones_like(x), x] # 1. bias column
w = np.linalg.solve(X.T @ X, X.T @ y) # 2. solve, never invert
yhat = X @ w
r2 = 1 - np.sum((y - yhat)**2) / np.sum((y - y.mean())**2)
print('w =', w, ' R2 =', round(r2, 4)) # 3. score
m = LinearRegression().fit(x.reshape(-1, 1), y)
print('sklearn:', round(m.intercept_, 4), round(m.coef_[0], 4))| printed line | value |
|---|---|
| w = [2.2 0.6] R2 = | 0.6 |
| sklearn: | 2.2 0.6 |
If your w reads [2.2, 0.6], R² is 0.6, and sklearn prints the same two numbers — you built linear regression from the math up.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| w = [2.2 0.6] R2 = | 0.6 |
| sklearn: | 2.2 0.6 |
Concept
Slides closed, out loud: explain (1) why the residual must be perpendicular to X's columns, (2) the two ways to reach XᵀXw = Xᵀy, and (3) what λI does to the eigenvalues of XᵀX.
Stretch: swap in np.linalg.lstsq(X, y) and confirm the same [2.2, 0.6]. Then make a feature collinear and watch solve raise while lstsq and ridge survive — the exact robustness gap the exam likes to probe.
Counterexample
Discussion prompt
Slides closed, out loud: explain (1) why the residual must be perpendicular to X's columns, (2) the two ways to reach XᵀXw = Xᵀy, and (3) what λI does to the eigenvalues of XᵀX.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The problem: no exact answer · Derivation 1: projection · Derivation 2: calculus · Solve it by hand · When is XᵀX invertible? · Ridge regression. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
y onto the column space of X — the residual comes out perpendicular, Xᵀr = 0XᵀXw = Xᵀy from that perpendicularity or from ∇‖Xw − y‖² = 0 — the two agree line for line2×2 by hand and confirm w = [2.2, 0.6], error 2.4, R² = 0.6XᵀX is invertible iff X has full column rank, and derive ridge (XᵀX + λI)w = Xᵀy when it isn'tsolve / lstsq (never inv) and match sklearn| move | the one thing to remember |
|---|---|
| normal equations | Xᵀ on the LEFT — XᵀX w = Xᵀy |
| derivation | perpendicular residual = zero gradient |
| bias | prepend a ones column or you fit through the origin |
| invertible | full column rank; collinear ⇒ singular ⇒ ridge |
| solving | solve / lstsq, never inv; ridge adds λI |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.