USAAIO Lesson 40, from Week 14 of Phase 2, fully worked on ONE tiny dataset, x=[1,2,3,4] and y=[3,4,7,8]. Closed-form OLS is solved as a 2x2 system by hand, entry by entry, the MSE gradient is derived term by term, and gradient descent is stepped by hand until it lands on the closed-form w=[1.0,1.8]. The canonical PyTorch training loop is then dissected step by step - tensors, nn.Linear, MSELoss, autograd, and SGD - with one step hand-checked and a full loss trace to convergence, and the zero_grad and learning-rate traps are shown through real divergence. It closes with weight_decay and ridge, polynomial features, R2 and RMSE, and a match against sklearn. Every snippet runs standalone in a fresh interpreter, and every number came from real torch and numpy execution. The lesson runs to 63 slides.
Subject: Machine Learning · 108 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 40 · Week 14 (Phase 2 begins)
The first real model, trained two ways. We solve one tiny dataset in closed form by hand, then rebuild the answer with gradient descent, then wrap it in the PyTorch training loop every later model reuses — no step skipped, every number from real execution.
Objectives
XᵀX, Xᵀy, invert the 2×2, and read off w∇L = (2/n)Xᵀ(Xw − y) term by term and step gradient descent by hand until it lands on the closed-form wnn.Linear, MSELoss, SGD) and hand-check a single step against autogradzero_grad() and a too-large learning rate — from real divergenceweight_decay, fit polynomial features, and score the fit with R² and RMSE, matching sklearnWarm-up
Discussion prompt
Before we open Lesson 40: Linear Regression & the PyTorch Training Loop: without looking back, what was the main idea of Mock Exam — Phase 1 Coding, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
a timed coding mock exam — implement PCA, gradient descent, MLE, and k-fold cross-validation from scratch in NumPy, each with a theory sub-question, verified against the libraries, plus a code-review pass. The Phase-1 coding milestone.
Section
Part 1 of 6 — recap & by-hand OLS
Concept
We keep one tiny dataset for the whole lesson so every number is checkable by hand. Four points — think of x as hours studied and y as a score:
| i | x | y |
|---|---|---|
| 1 | 1 | 3 |
| 2 | 2 | 4 |
| 3 | 3 | 7 |
| 4 | 4 | 8 |
We want the best line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁. Four points, two unknowns: overdetermined, so no line hits all four.
Comparison
Comparison matrix
From The running example: refill the x column from what you know. The rest of the table is as it appeared.
| i | x | y |
|---|---|---|
| 1 | 1 | 3 |
| 2 | 2 | 4 |
| 3 | 3 | 7 |
| 4 | 4 | 8 |
Intuition
Since no line is exact, we pick the one whose predictions miss by the least overall. 'Least' means the sum of squared misses — squaring makes every error positive and punishes big misses hardest.
That single number, averaged over the points, is the mean squared error — the loss we minimize both by algebra (closed form) and by iteration (gradient descent). Two roads, same destination.
Concept
Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:
\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \end{bmatrix}, \quad w = \begin{bmatrix} w_0 \\ w_1 \end{bmatrix}, \quad y = \begin{bmatrix} 3 \\ 4 \\ 7 \\ 8 \end{bmatrix} \]
\[ L(w) = \frac{1}{n}\lVert Xw - y \rVert^2 = \frac{1}{n}\sum_{i=1}^{n}\big(w_0 + w_1 x_i - y_i\big)^2 \]
Here n = 4; the 1/n makes it a MEAN squared error
Why: Dividing by n makes the loss independent of dataset size, so a learning rate that works on 4 points also works on 4 million. PyTorch's MSELoss uses this 1/n mean by default.
Counterexample
Discussion prompt
Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:
\[ \boxed{\,X^\top X\, w^\star = X^\top y\,} \]
We'll solve this 2×2 by hand this lesson, then spend most of the deck reaching the same w iteratively — because iteration is what generalizes to every model with no closed form.
Analogy
Discussion prompt
Explain The closed form (Lesson 7, in one line) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:
Ranking
Put in order
Put the moves of Build XᵀX, entry by entry into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Four 1s dotted with four 1s: 1+1+1+1 = 4.
Worked example
(XᵀX)ⱼₖ = column j of X dotted with column k. Column 0 is ones, column 1 is x, so there are only three distinct entries (it is symmetric):
Top-left = ones·ones = Σ1 = 4
Why: Four 1s dotted with four 1s: 1+1+1+1 = 4. This entry is just n, the number of points.
Off-diagonal = ones·x = Σxᵢ = 1+2+3+4 = 10
Why: The ones vector dotted with x simply sums x. It fills both off-diagonal slots by symmetry.
Bottom-right = x·x = Σxᵢ² = 1+4+9+16 = 30
Why: Sum of squared features.
\[ X^\top X = \begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} \]
Notation
Annotate
From Build XᵀX, entry by entry — read this one piece at a time. What is each part doing?
On: \( X^\top X = \begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} \)
Worked example
Top = ones·y = Σyᵢ = 3+4+7+8 = 22
Why: The ones row of Xᵀ dotted with y sums the targets.
Bottom = x·y = Σxᵢyᵢ
Why: The x row of Xᵀ dotted with y: multiply pairwise, then add.
\[ \sum_i x_i y_i = 1\cdot3 + 2\cdot4 + 3\cdot7 + 4\cdot8 = 3 + 8 + 21 + 32 = 64 \]
\[ X^\top y = \begin{bmatrix} 22 \\ 64 \end{bmatrix} \]
Blank canvas
Draw it
Draw what Build Xᵀy, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Estimation
Predict first
For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:
Commit before you compute: what does Invert the 2×2 — determinant first come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: det = 20 ≠ 0 → invertible, unique solution
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique.
Worked example
For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:
\[ \det\!\begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} = 4\cdot 30 - 10\cdot 10 = 120 - 100 = 20 \]
det = 20 ≠ 0 → invertible, unique solution
Why: A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique. (det = 0 would mean collinear columns — the singular case ridge fixes.)
\[ (X^\top X)^{-1} = \frac{1}{20}\begin{bmatrix} 30 & -10 \\ -10 & 4 \end{bmatrix} = \begin{bmatrix} 1.5 & -0.5 \\ -0.5 & 0.2 \end{bmatrix} \]
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
det = 20 ≠ 0 → invertible, unique solution
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:
Missing information
Discussion prompt
w★ = (XᵀX)⁻¹ Xᵀy. Dot each row of the inverse with [22, 64], one entry at a time:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
First row of the inverse dotted with Xᵀy. Equivalently (30·22 − 10·64)/20 = (660 − 640)/20 = 20/20 = 1.0.
Worked example
w★ = (XᵀX)⁻¹ Xᵀy. Dot each row of the inverse with [22, 64], one entry at a time:
w₀ = 1.5·22 + (−0.5)·64 = 33 − 32 = 1.0
Why: First row of the inverse dotted with Xᵀy. Equivalently (30·22 − 10·64)/20 = (660 − 640)/20 = 20/20 = 1.0.
w₁ = (−0.5)·22 + 0.2·64 = −11 + 12.8 = 1.8
Why: Second row dotted with Xᵀy. Equivalently (−10·22 + 4·64)/20 = (−220 + 256)/20 = 36/20 = 1.8.
\[ w^\star = \begin{bmatrix} 1.0 \\ 1.8 \end{bmatrix} \;\Longrightarrow\; \widehat{y} = 1.0 + 1.8\,x \]
Translation
\( w^\star = \begin{bmatrix} 1.0 \\ 1.8 \end{bmatrix} \;\Longrightarrow\; \widehat{y} = 1.0 + 1.8\,x \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Pattern
Predict first
The table runs: XtX | [[4.0, 10.0], [10.0, 30.0]] · Xty | [22.0, 64.0]
In Confirm w in NumPy, given the rows so far: what is the next one — the row where object is w = [w₀, w₁]?
Correct: w = [w₀, w₁] | [1.0, 1.8]
| object | printed value (verified) |
|---|---|
| XtX | [[4.0, 10.0], [10.0, 30.0]] |
| Xty | [22.0, 64.0] |
| w = [w₀, w₁] | [1.0, 1.8] |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. np.linalg.solve factors the 2×2 (LU) and never forms an explicit inverse — faster and more numerically stable than inv.
Worked example
The whole by-hand derivation is one solve call. This block is self-contained — it re-imports NumPy and re-defines the data:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
XtX, Xty = X.T @ X, X.T @ y
w = np.linalg.solve(XtX, Xty) # never inv
print("XtX =", XtX.tolist())
print("Xty =", Xty.tolist())
print("w =", w.round(4).tolist())solve returns exactly the by-hand [1.0, 1.8]
Why: np.linalg.solve factors the 2×2 (LU) and never forms an explicit inverse — faster and more numerically stable than inv. The printed w matches our hand math to the digit.
| object | printed value (verified) |
|---|---|
| XtX | [[4.0, 10.0], [10.0, 30.0]] |
| Xty | [22.0, 64.0] |
| w = [w₀, w₁] | [1.0, 1.8] |
Trade off
Comparison matrix
From Confirm w in NumPy: every row here is a choice with a cost. Fill the printed value (verified) column, then say which row you would actually pick and what you give up for it.
| object | printed value (verified) |
|---|---|
| XtX | [[4.0, 10.0], [10.0, 30.0]] |
| Xty | [22.0, 64.0] |
| w = [w₀, w₁] | [1.0, 1.8] |
Worked example
Plug each x into ŷ = 1.0 + 1.8x, then subtract from the true y to get every residual and its square:
| x | y | ŷ = 1 + 1.8x | r = y − ŷ | r² |
|---|---|---|---|---|
| 1 | 3 | 2.8 | +0.2 | 0.04 |
| 2 | 4 | 4.6 | −0.6 | 0.36 |
| 3 | 7 | 6.4 | +0.6 | 0.36 |
| 4 | 8 | 8.2 | −0.2 | 0.04 |
Σr² = 0.04 + 0.36 + 0.36 + 0.04 = 0.8
Why: This is the minimized ‖Xw − y‖² = 0.8 — the smallest total squared error any line achieves on this data. The MSE is 0.8/4 = 0.2. We reuse both for R² and RMSE.
Pattern
Step through it
Step through Predictions and residuals — full trace one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
The normal equations promise Xᵀr = 0: the residual is perpendicular to every column of X. Confirm it on r = [0.2, −0.6, 0.6, −0.2]:
Row 0 (ones·r): 0.2 − 0.6 + 0.6 − 0.2 = 0
Why: Residuals sum to exactly zero — a direct consequence of the bias column. The fit neither systematically over- nor under-predicts.
Row 1 (x·r): 1(0.2) + 2(−0.6) + 3(0.6) + 4(−0.2) = 0.2 − 1.2 + 1.8 − 0.8 = 0
Why: Perpendicular to the feature column too. Xᵀr = [0, 0], so w is truly the least-squares solution — this is your always-available correctness check.
| condition | value (verified) |
|---|---|
| Σ rᵢ | 0.0 |
| Σ xᵢ rᵢ | 0.0 |
| → Xᵀr | [0, 0] |
Section
Part 2 of 6 — derive & step by hand
Concept
The closed form needs to solve a d×d system XᵀX w = Xᵀy. That is fine for d = 2, but costs O(d³) and only exists for linear least squares.
Gradient descent instead nudges w downhill on the loss, step by step. It scales to huge d and n, and — crucially — works for every model in Phase 2 (logistic regression, MLPs, transformers) where no closed form exists.
So we practice it here, on a problem where we already know the answer [1.0, 1.8] — the perfect place to watch iteration converge to truth.
Explain it
Discussion prompt
Explain Why iterate when we have the exact answer? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
The closed form needs to solve a d×d system XᵀX w = Xᵀy. That is fine for d = 2, but costs O(d³) and only exists for linear least squares.
Concept
Gradient descent repeatedly steps against the gradient (the uphill direction), scaled by a learning rate η:
\[ w \;\leftarrow\; w - \eta\,\nabla L(w) \]
The gradient points uphill, so subtracting it walks downhill. η sets the step size: too small and it crawls, too large and it overshoots and diverges. To run it we need ∇L explicitly — derive it next.
Ranking
Put in order
Put the moves of Derive the MSE gradient, term by term into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The mean squared error on the design matrix — the exact loss PyTorch's MSELoss computes.
Worked example
Start from L(w) = (1/n)‖Xw − y‖²
Why: The mean squared error on the design matrix — the exact loss PyTorch's MSELoss computes.
\[ L(w) = \tfrac{1}{n}(Xw - y)^\top (Xw - y) \]
Differentiate the squared norm: ∇‖Xw − y‖² = 2Xᵀ(Xw − y)
Why: Chain rule on vᵀv with v = Xw − y: the outer derivative gives 2v, the inner derivative of v w.r.t. w is Xᵀ (from Lesson 7's ∇(wᵀAw) = 2Aw with A = XᵀX).
\[ \nabla \lVert Xw - y \rVert^2 = 2\,X^\top(Xw - y) \]
Carry the 1/n through
Why: The 1/n is a constant multiplier, so it rides along untouched into the gradient.
\[ \boxed{\,\nabla L(w) = \tfrac{2}{n}\,X^\top(Xw - y)\,} \]
Blank canvas
Draw it
Draw what Derive the MSE gradient, term by term just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From The gradient in code, checked two ways — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
def grad(w): return (2.0/n) * (X.T @ (X @ w - y))
print("grad at [0, 0] =", grad(np.array([0., 0.])).tolist())
print("grad at [1.0, 1.8] =", grad(np.array([1.0, 1.8])).round(6).tolist())
Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. grad([0,0]) = (2/4)·Xᵀ(−y) = 0.5·(−[22,64]) = [−11, −32]: strongly downhill.
Worked example
Evaluate ∇L = (2/n)Xᵀ(Xw − y) at the start w = [0, 0] and at the known optimum w = [1.0, 1.8]. Self-contained block:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
def grad(w): return (2.0/n) * (X.T @ (X @ w - y))
print("grad at [0, 0] =", grad(np.array([0., 0.])).tolist())
print("grad at [1.0, 1.8] =", grad(np.array([1.0, 1.8])).round(6).tolist())At the start the gradient is steep; at the optimum it is zero
Why: grad([0,0]) = (2/4)·Xᵀ(−y) = 0.5·(−[22,64]) = [−11, −32]: strongly downhill. grad([1.0,1.8]) = [0, 0]: flat, confirming [1.0, 1.8] is the minimum GD must find.
| w | ∇L (verified) |
|---|---|
| [0, 0] | [−11.0, −32.0] |
| [1.0, 1.8] | [0.0, 0.0] |
Worked example
Start at w = [0, 0] with learning rate η = 0.1. We just found ∇L([0,0]) = [−11, −32]. Apply the update:
w₀: 0 − 0.1·(−11) = 0 + 1.1 = 1.1
Why: Subtracting a negative gradient moves the intercept UP toward its target of 1.0 — the right direction.
w₁: 0 − 0.1·(−32) = 0 + 3.2 = 3.2
Why: The slope jumps to 3.2 — overshooting the target 1.8, because the gradient was so steep at the start. GD will correct this on the next step.
\[ w^{(1)} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} - 0.1\begin{bmatrix} -11 \\ -32 \end{bmatrix} = \begin{bmatrix} 1.1 \\ 3.2 \end{bmatrix} \]
Notation
Annotate
From Step 1 of gradient descent, by hand — read this one piece at a time. What is each part doing?
On: \( w^{(1)} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} - 0.1\begin{bmatrix} -11 \\ -32 \end{bmatrix} = \begin{bmatrix} 1.1 \\ 3.2 \end{bmatrix} \)
Faded example
Fill in the blanks
Watch four steps descend, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(4):
g = (2.0/n) * (X.T @ (X @ w - y))
w = w - eta * g
loss = ((X @ w - y)2).mean()**
print(f"step ___: w = ___ loss = ___")
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Each step over/undershoots the target 1.8, but the loss shrinks 15.6 → 7.1 → 3.3 → 1.6 — the bowl is convex, so downhill always makes progress even when a single coordinate zig-zags.
Worked example
Run the same by-hand update in a loop and print w and the loss each step. Self-contained:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(4):
g = (2.0/n) * (X.T @ (X @ w - y))
w = w - eta * g
loss = ((X @ w - y)**2).mean()
print(f"step {step+1}: w = {w.round(4).tolist()} loss = {loss:.4f}")The slope oscillates (3.2 → 1.05 → 2.49 → 1.52) but the loss falls every step
Why: Each step over/undershoots the target 1.8, but the loss shrinks 15.6 → 7.1 → 3.3 → 1.6 — the bowl is convex, so downhill always makes progress even when a single coordinate zig-zags.
| step | w = [w₀, w₁] | loss (verified) |
|---|---|---|
| 1 | [1.1, 3.2] | 15.6100 |
| 2 | [0.38, 1.05] | 7.1281 |
| 3 | [0.879, 2.485] | 3.3194 |
| 4 | [0.5607, 1.518] | 1.6088 |
Worked example
Keep stepping. After enough iterations GD reaches the closed-form answer. Same loop, more steps:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(2000):
w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
print("GD w =", w.round(4).tolist())
print("closed form =", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())GD lands on [1.0, 1.8] — identical to the closed form
Why: Because MSE is convex with a single global minimum, gradient descent from any start converges to the SAME w the normal equations give. Two derivations, one answer — proven numerically.
| method | w (verified) |
|---|---|
| gradient descent (2000 steps) | [1.0, 1.8] |
| closed-form solve | [1.0, 1.8] |
| match | ✓ |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Bigger steps must reach the minimum faster — crank η up to 0.5.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up.
Keep η small enough that each step actually decreases the loss — here 0.1 works.
Why: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up. Too-large η diverges — the opposite of what you wanted.
Trap
Bigger steps must reach the minimum faster — crank η up to 0.5.
import numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]; n = 4
w = np.array([0., 0.])
eta = 0.5 # too big
for step in range(4):
w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
print(f"step {step+1}: loss = {((X@w - y)**2).mean():.2f}")| step | loss (verified) |
|---|---|
| 1 | 1852.25 |
| 2 | 100060.09 |
| 3 | 5405933.68 |
| 4 | 292066268.61 |
Loss EXPLODES to infinity
Why: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up. Too-large η diverges — the opposite of what you wanted.
Keep η small enough that each step actually decreases the loss — here 0.1 works.
import numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]; n = 4
w = np.array([0., 0.])
eta = 0.1 # stable
for step in range(4):
w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
print(f"step {step+1}: loss = {((X@w - y)**2).mean():.2f}")| step | loss (verified) |
|---|---|
| 1 | 15.61 |
| 2 | 7.13 |
| 3 | 3.32 |
| 4 | 1.61 |
Loss falls monotonically toward 0.2
Why: A small enough step size guarantees descent on a convex loss. If your loss ever increases or turns to NaN, halve η first — it's the #1 cause of a 'broken' training run.
Error analysis
Annotate
Walk the callouts on Trap: a learning rate that is too large. Each one is a place this is easy to get subtly wrong.
Section
Part 3 of 6 — every line
Intuition
We hand-derived ∇L because the model was simple. Real models have millions of parameters — deriving gradients by hand is hopeless.
PyTorch's autograd computes the gradient of any loss automatically by tracking the operations you run. Your job shrinks to: define the model, define the loss, and run a fixed five-step loop. That loop is the same for every model you will ever train.
Worked example
nn.Linear expects a 2-D input: one row per sample, one column per feature. unsqueeze(1) turns a length-4 vector into a (4, 1) column:
import torch
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
print("xt.shape =", tuple(xt.shape))
print("yt.shape =", tuple(yt.shape))Both become shape (4, 1) — 4 rows, 1 column
Why: Feed a flat (4,) vector to nn.Linear and the shapes silently mis-broadcast, giving a wrong loss with no error. Always reshape features to (n, features) — here (4, 1).
| tensor | shape (verified) |
|---|---|
| xt | (4, 1) |
| yt | (4, 1) |
| xt values | [1.0, 2.0, 3.0, 4.0] |
Estimation
Predict first
nn.Linear(1, 1) holds one weight and one bias, initialized randomly. Seed first so the numbers are reproducible:
Commit before you compute: what does Step B — nn.Linear is ŷ = wx + b come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: It starts with random weight −0.0075, bias 0.5364 — a bad line
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With torch.manual_seed(0), nn.Linear(1,1) initializes to these exact values.
Worked example
nn.Linear(1, 1) holds one weight and one bias, initialized randomly. Seed first so the numbers are reproducible:
import torch
torch.manual_seed(0)
layer = torch.nn.Linear(1, 1)
print("weight =", round(layer.weight.item(), 4))
print("bias =", round(layer.bias.item(), 4))
xt = torch.tensor([[1.], [2.], [3.], [4.]])
print("forward =", layer(xt).squeeze().round(decimals=4).tolist())It starts with random weight −0.0075, bias 0.5364 — a bad line
Why: With torch.manual_seed(0), nn.Linear(1,1) initializes to these exact values. The forward pass ŷ = weight·x + bias is nowhere near the data yet — training's whole job is to move these two numbers to [1.8, 1.0].
| parameter | initial value (seed 0) |
|---|---|
| weight (slope) | −0.0075 |
| bias (intercept) | 0.5364 |
| forward on x=[1,2,3,4] | [0.529, 0.521, 0.514, 0.507] |
Sorting
Sort into buckets
These are the pieces of Lesson 40: Linear Regression & the PyTorch Training Loop, out of order. Put each one back under the part of the lesson it belongs to.
Worked example
nn.MSELoss() computes the mean of the squared differences — exactly the L(w) we minimized. A quick check with hand-picked predictions:
import torch
pred = torch.tensor([[2.], [3.], [6.], [7.]])
target = torch.tensor([[3.], [4.], [7.], [8.]])
loss = torch.nn.MSELoss()(pred, target)
print("MSE =", loss.item())Every error is −1, squared is 1, mean of four 1s is 1.0
Why: MSELoss = mean((pred − target)²). Here pred is 1 below target everywhere, so each squared error is 1 and the mean is exactly 1.0 — confirming MSELoss is the same 1/n·Σ(ŷ−y)² we minimize.
| pred | target | (pred − target)² |
|---|---|---|
| 2 | 3 | 1 |
| 3 | 4 | 1 |
| 6 | 7 | 1 |
| 7 | 8 | 1 |
Invariant
Step through it
Step through Step C — MSELoss scores the fit one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
Every PyTorch model trains with the same five steps per iteration. Memorize this order — it never changes:
opt.zero_grad() — clear last step's gradientsout = model(xt) — forward pass (predict)loss = lossfn(out, yt) — score the predictionsloss.backward() — autograd fills every .gradopt.step() — the optimizer updates the parametersSteps 4–5 are exactly our w ← w − η∇L: backward() computes ∇L, step() applies the subtraction. PyTorch just automates the gradient.
Fill the middle
Fill in the blanks
From One iteration, hand-checked against autograd — one line has had its right-hand side removed. Put it back.
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward()
print("weight.grad =", round(model.weight.grad.item(), 4))
print("bias.grad =", round(model.bias.grad.item(), 4))
opt.step()
print("weight now =", round(model.weight.item(), 4))
print("bias now =", round(model.bias.item(), 4))
Why: loss is what everything below it consumes, so the wrong expression here fails later and somewhere else. SGD applied w ← w − lr·grad exactly.
Worked example
Run one loop iteration and print the gradients autograd computed and the updated weights. We'll check they obey w ← w − η·grad:
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward()
print("weight.grad =", round(model.weight.grad.item(), 4))
print("bias.grad =", round(model.bias.grad.item(), 4))
opt.step()
print("weight now =", round(model.weight.item(), 4))
print("bias now =", round(model.bias.item(), 4))Check the update: −0.0075 − 0.1·(−29.4301) = 2.9355 ✓
Why: SGD applied w ← w − lr·grad exactly. weight: −0.0075 − 0.1·(−29.4301) = 2.9355 (printed). bias: 0.5364 − 0.1·(−9.9645) = 1.5329 (printed). Autograd's gradient plugged straight into our by-hand rule.
| quantity | value (verified) |
|---|---|
| weight.grad | −29.4301 |
| bias.grad | −9.9645 |
| weight after step | 2.9355 |
| bias after step | 1.5329 |
Worked example
Wrap the five steps in a for loop and run it. Watch it reach the closed-form [1.0, 1.8]:
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
lossfn = torch.nn.MSELoss()
for epoch in range(1000):
opt.zero_grad()
loss = lossfn(model(xt), yt)
loss.backward()
opt.step()
print("final loss =", round(loss.item(), 4))
print("slope =", round(model.weight.item(), 4))
print("intercept =", round(model.bias.item(), 4))loss → 0.2, slope → 1.8, intercept → 1.0
Why: The training loop lands on the exact closed-form solution: slope 1.8, intercept 1.0, minimum MSE 0.2 (= 0.8/4). The framework did what our hand-GD did, with autograd computing the gradient for us.
| source | slope | intercept | loss |
|---|---|---|---|
| closed-form OLS | 1.8 | 1.0 | 0.2 |
| PyTorch SGD (1000 ep) | 1.8 | 1.0 | 0.2 |
| match | ✓ | ✓ | ✓ |
Comparison
Comparison matrix
From The full loop — train to convergence: refill the intercept column from what you know. The rest of the table is as it appeared.
| source | slope | intercept | loss |
|---|---|---|---|
| closed-form OLS | 1.8 | 1.0 | 0.2 |
| PyTorch SGD (1000 ep) | 1.8 | 1.0 | 0.2 |
| match | ✓ | ✓ | ✓ |
Worked example
Print the loss and parameters at chosen epochs to see the descent curve — fast at first, then fine-tuning:
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
lossfn = torch.nn.MSELoss()
for epoch in range(1001):
opt.zero_grad()
loss = lossfn(model(xt), yt)
loss.backward()
opt.step()
if epoch in (0, 1, 10, 100, 1000):
print(f"epoch {epoch:4d}: loss={loss.item():.4f} slope={model.weight.item():.4f} intercept={model.bias.item():.4f}")loss 29.1 → 0.2 in about 100 epochs; the rest is polish
Why: The loss drops fastest early (steep gradient), then the slope/intercept settle onto 1.8/1.0. By epoch 100 the loss is already 0.2000 to four places — later epochs only refine the trailing digits.
| epoch | loss | slope | intercept |
|---|---|---|---|
| 0 | 29.1068 | 2.9355 | 1.5329 |
| 1 | 13.1801 | 0.9658 | 0.8586 |
| 10 | 0.2113 | 1.7885 | 1.1043 |
| 100 | 0.2000 | 1.7979 | 1.0063 |
| 1000 | 0.2000 | 1.8000 | 1.0000 |
Pattern
Step through it
Step through A loss trace: watch it converge one row at a time. What is driving the change, and what would the row after the last one be?
Anomaly
Predict first
A student writes this, and it looks reasonable:
The gradient is fresh each pass, so just backward() then step() — skip zero_grad().
It is wrong. Say what breaks — and say it before you turn the page.
Correct: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad.
Call opt.zero_grad() at the top of every iteration.
Why: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad. Without zeroing, step 2 uses grad₁+grad₂, so the optimizer takes wrong-sized steps and the loss bounces (43.30, then 49.75) instead of descending.
Trap
The gradient is fresh each pass, so just backward() then step() — skip zero_grad().
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(6):
# BUG: no opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward(); opt.step()
print(f"epoch {epoch}: loss = {loss.item():.2f}")| epoch | loss (verified) |
|---|---|
| 0 | 29.11 |
| 1 | 13.18 |
| 2 | 43.30 |
| 3 | 2.27 |
| 4 | 49.75 |
Loss lurches around instead of falling smoothly
Why: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad. Without zeroing, step 2 uses grad₁+grad₂, so the optimizer takes wrong-sized steps and the loss bounces (43.30, then 49.75) instead of descending.
Call opt.zero_grad() at the top of every iteration.
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(6):
opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward(); opt.step()
print(f"epoch {epoch}: loss = {loss.item():.2f}")| epoch | loss (verified) |
|---|---|
| 0 | 29.11 |
| 1 | 13.18 |
| 2 | 6.36 |
| 3 | 3.24 |
| 4 | 1.75 |
Loss falls monotonically — the fix is one line
Why: zero_grad() resets every .grad to zero so each step uses only THIS iteration's gradient. This is the single most common training-loop bug; put it first, every loop, always.
Error analysis
Annotate
Walk the callouts on Trap: forgetting zero_grad(). Each one is a place this is easy to get subtly wrong.
Section
Part 4 of 6 — weight_decay & polynomials
Concept
Ridge (Lesson 12) adds an L2 penalty on weight size to the loss. In PyTorch it is a one-word optimizer argument, weight_decay:
\[ \min_w \; \lVert Xw - y \rVert^2 + \lambda \lVert w \rVert^2 \;\equiv\; \texttt{SGD(..., weight\_decay=}\lambda\texttt{)} \]
Big weights now cost something, so the fit is pulled toward 0 — trading a little training accuracy for stability. It's ridge, wired into the optimizer.
Faded example
Fill in the blanks
weight_decay shrinks the fit, with the scaffolding fading: two lines are gone now — fill both.
import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
torch.manual_seed(0)
m = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
for _ in range(1000):
opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
print(f"wd=___: slope=___ intercept=___")
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. With weight_decay=0.5 the penalty pulls the parameters toward zero, so the fit backs off the exact OLS solution.
Worked example
Train the same loop with weight_decay = 0 and 0.5, and compare the learned slope and intercept:
import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
torch.manual_seed(0)
m = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
for _ in range(1000):
opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
print(f"wd={wd}: slope={m.weight.item():.4f} intercept={m.bias.item():.4f}")The intercept shrinks 1.00 → 0.76 under the penalty
Why: With weight_decay=0.5 the penalty pulls the parameters toward zero, so the fit backs off the exact OLS solution. The intercept drops from 1.00 to 0.76; the slope barely moves (1.80 → 1.82) — the penalty hits whichever weights it can shrink cheaply.
| weight_decay | slope | intercept (verified) |
|---|---|---|
| 0.0 (plain OLS) | 1.8000 | 1.0000 |
| 0.5 (ridge) | 1.8182 | 0.7636 |
Concept
Linear regression is only linear in the weights, not the features. Map x → [1, x, x², …, xᵈ] and the same machinery fits a curve — the design matrix just gets more columns.
\[ \hat y = w_0 + w_1 x + w_2 x^2 + \cdots + w_d x^d \]
The gradient, the loop, and weight_decay are unchanged — you only widen X. But more columns means more power to fit, and past a point, to overfit.
Fill the middle
Fill in the blanks
From Degree-2 features on our data — one line has had its right-hand side removed. Put it back.
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
Xp = np.c_[np.ones(4), x, x2] # [1, x, x^2]
w = np.linalg.solve(Xp.T@Xp, Xp.T@y)
print("w (deg 2) =", w.round(4).tolist())
print("SSE =", round(((Xp@w - y)2).sum(), 4))
Why: Xp is what everything below it consumes, so the wrong expression here fails later and somewhere else. OLS finds w = [1.0, 1.8, 0.0]: the quadratic term is exactly zero because our y = [3,4,7,8] has no curvature.
Worked example
Add an x² column and solve OLS on the widened matrix. Self-contained:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
Xp = np.c_[np.ones(4), x, x**2] # [1, x, x^2]
w = np.linalg.solve(Xp.T@Xp, Xp.T@y)
print("w (deg 2) =", w.round(4).tolist())
print("SSE =", round(((Xp@w - y)**2).sum(), 4))The x² coefficient comes out 0.0 — the data really is linear
Why: OLS finds w = [1.0, 1.8, 0.0]: the quadratic term is exactly zero because our y = [3,4,7,8] has no curvature. SSE stays 0.8, same as the line. Extra capacity that the data doesn't need sits idle here.
| degree | w | SSE (verified) |
|---|---|---|
| 1 (line) | [1.0, 1.8] | 0.8 |
| 2 (parabola) | [1.0, 1.8, 0.0] | 0.8 |
Pattern
Predict first
The table runs: 1 | 2 | 0.800000 · 2 | 3 | 0.800000
In Degree 3 overfits: SSE hits zero, given the rows so far: what is the next one — the row where degree is 3?
Correct: 3 | 4 | 0.000000
| degree | #params | training SSE (verified) |
|---|---|---|
| 1 | 2 | 0.800000 |
| 2 | 3 | 0.800000 |
| 3 | 4 | 0.000000 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. 4 params for 4 points interpolates exactly: zero training error.
Worked example
Push the degree up. With 4 points, a degree-3 polynomial has 4 parameters — enough to pass through every point exactly:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
for deg in (1, 2, 3):
Xp = np.vander(x, deg+1, increasing=True) # [1, x, ..., x^deg]
w = np.linalg.lstsq(Xp, y, rcond=None)[0]
sse = ((Xp@w - y)**2).sum()
print(f"degree {deg}: SSE = {sse:.6f}, #params = {deg+1}")Degree 3 drives SSE to 0 — but that is memorizing, not learning
Why: 4 params for 4 points interpolates exactly: zero training error. On NEW data such a wiggly curve does worse than the honest line — the definition of overfitting. Low training error is not the goal; low validation error is.
| degree | #params | training SSE (verified) |
|---|---|---|
| 1 | 2 | 0.800000 |
| 2 | 3 | 0.800000 |
| 3 | 4 | 0.000000 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Higher degree gives lower SSE, so pick the degree with the smallest training error.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points.
Choose the degree with the smallest validation error — error on data the model did not train on.
Why: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points. Optimizing training error alone always picks the most complex, most overfit model.
Trap
Higher degree gives lower SSE, so pick the degree with the smallest training error.
Choose degree 3 because its training SSE = 0
Why: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points. Optimizing training error alone always picks the most complex, most overfit model.
Choose the degree with the smallest validation error — error on data the model did not train on.
Hold out data; pick the degree that generalizes best
Why: Validation error falls then RISES as overfitting kicks in (bias–variance, Lesson 32). The minimum of the validation curve — here the honest degree-1 line — is the right complexity, not the training-error minimum.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x as hours studied and y as a score:; Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:; Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:η up to 0.5.; The gradient is fresh each pass, so just backward() then step() — skip zero_grad().Section
Part 5 of 6 — R², RMSE, sklearn
Concept
Two standard scores. R² is the fraction of the target's variance the model explains (1 = perfect, 0 = no better than the mean). RMSE is the typical error, in the target's own units:
\[ R^2 = 1 - \frac{\sum_i (y_i - \hat y_i)^2}{\sum_i (y_i - \bar y)^2}, \qquad \mathrm{RMSE} = \sqrt{\tfrac{1}{n}\textstyle\sum_i (y_i - \hat y_i)^2} \]
We already have the pieces: SS_res = Σr² = 0.8 and MSE = 0.2. Only SS_tot (the mean-only baseline error) is new.
Explain it
Discussion prompt
Explain R² and RMSE to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Two standard scores. R² is the fraction of the target's variance the model explains (1 = perfect, 0 = no better than the mean). RMSE is the typical error, in the target's own units:
Missing information
Discussion prompt
ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
The mean-only baseline leaves total error 17.0. Our line leaves only 0.8, so R² = 1 − 0.8/17 = 0.9529: the line explains 95% of the variance in y.
Worked example
ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
w = np.linalg.solve(X.T@X, X.T@y)
yhat = X @ w
sse = ((y - yhat)**2).sum()
r2 = 1 - sse / ((y - y.mean())**2).sum()
rmse = np.sqrt(sse / len(y))
print("R2 =", round(r2, 4))
print("RMSE =", round(rmse, 4))SS_tot = 6.25 + 2.25 + 2.25 + 6.25 = 17.0
Why: The mean-only baseline leaves total error 17.0. Our line leaves only 0.8, so R² = 1 − 0.8/17 = 0.9529: the line explains 95% of the variance in y.
| quantity | value (verified) |
|---|---|
| SS_res = Σ(y − ŷ)² | 0.8 |
| SS_tot = Σ(y − ȳ)² | 17.0 |
| R² = 1 − SS_res/SS_tot | 0.9529 |
| RMSE = √(SS_res/n) | 0.4472 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
SS_tot = 6.25 + 2.25 + 2.25 + 6.25 = 17.0
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:
Estimation
Predict first
The final sanity check: fit sklearn.LinearRegression and confirm it prints the same intercept, slope, and R²:
Commit before you compute: what does Cross-check against sklearn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: sklearn agrees to the digit: intercept 1.0, slope 1.8, R² 0.9529
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Three independent routes — by-hand OLS, the PyTorch loop, and sklearn — all land on the identical [1.0, 1.8] with R² 0.9529.
Worked example
The final sanity check: fit sklearn.LinearRegression and confirm it prints the same intercept, slope, and R²:
import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
m = LinearRegression().fit(x.reshape(-1, 1), y)
print("intercept =", round(m.intercept_, 4))
print("slope =", round(m.coef_[0], 4))
print("R2 =", round(m.score(x.reshape(-1, 1), y), 4))sklearn agrees to the digit: intercept 1.0, slope 1.8, R² 0.9529
Why: Three independent routes — by-hand OLS, the PyTorch loop, and sklearn — all land on the identical [1.0, 1.8] with R² 0.9529. When your from-scratch result matches the library, you built it right.
| source | intercept | slope | R² |
|---|---|---|---|
| by-hand OLS | 1.0 | 1.8 | 0.9529 |
| PyTorch SGD | 1.0 | 1.8 | 0.9529 |
| sklearn | 1.0 | 1.8 | 0.9529 |
Invariant
Step through it
Step through Cross-check against sklearn one row at a time. One of these columns never changes — find it, and say why it cannot.
Concept
A high R² can still hide curvature or unequal spread. After fitting, plot residuals r = y − ŷ against the fitted ŷ: it should look like structureless scatter around zero.
A visible pattern (a U-shape, a fan) means the linear model is misspecified — add features or change the model. R² won't warn you; the residual plot will.
Analogy
Discussion prompt
Explain Don't trust R² alone — check residuals by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A high R² can still hide curvature or unequal spread. After fitting, plot residuals r = y − ŷ against the fitted ŷ: it should look like structureless scatter around zero.
Ranking
Put in order
These are the steps of The regression + training-loop recipe, scrambled. Put them back in order before the next slide shows you.
X = [1, features]; tensors reshaped to (n, features) for nn.Linearnn.Linear (add polynomial columns for curves); MSELosszero_grad → forward → loss → backward → step, with a small learning rateweight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation errorR², RMSE, a residual plot, and match a from-scratch / sklearn baselineWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
X = [1, features]; tensors reshaped to (n, features) for nn.Linearnn.Linear (add polynomial columns for curves); MSELosszero_grad → forward → loss → backward → step, with a small learning rateweight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation errorR², RMSE, a residual plot, and match a from-scratch / sklearn baselineEdge cases
Discussion prompt
The regression + training-loop recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
X = [1, features]; tensors reshaped to (n, features) for nn.Linearnn.Linear (add polynomial columns for curves); MSELosszero_grad → forward → loss → backward → step, with a small learning rateweight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation errorR², RMSE, a residual plot, and match a from-scratch / sklearn baselineElimination
Eliminate the wrong options
What is the correct order of one PyTorch training iteration?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Clear old grads, run the forward pass to predict, score the loss, backprop to fill every .grad, then let the optimizer update the parameters. Skipping zero_grad accumulates gradients across steps (the earlier trap).
Check
Recall the five-line rhythm. Think about which step needs the output of which.
Check your understanding
What is the correct order of one PyTorch training iteration?
Answer: A
Why: Clear old grads, run the forward pass to predict, score the loss, backprop to fill every .grad, then let the optimizer update the parameters. Skipping zero_grad accumulates gradients across steps (the earlier trap).
Prediction
Predict first
Linear regression has an exact closed form. Why train it with gradient descent anyway?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: GD scales to huge data and generalizes to models with no closed form
Why: The closed form needs to solve a d×d system (O(d³), and only exists for linear least squares). GD scales to large d/n and — the whole point of learning it now — works for every later model (logistic regression, MLPs, transformers) where no closed form exists.
Check
We got [1.0, 1.8] two ways. Why bother with the slower, iterative one?
Check your understanding
Linear regression has an exact closed form. Why train it with gradient descent anyway?
Answer: A
Why: The closed form needs to solve a d×d system (O(d³), and only exists for linear least squares). GD scales to large d/n and — the whole point of learning it now — works for every later model (logistic regression, MLPs, transformers) where no closed form exists.
Check
You derived it term by term. Read the dimensions carefully.
Check your understanding
For L(w) = (1/n)‖Xw − y‖², what is ∇L(w)?
Answer: A
Why: Chain rule: the outer 2(Xw − y) times the inner Xᵀ, all scaled by 1/n, gives (2/n)Xᵀ(Xw − y). At w = [1.0, 1.8] this evaluates to [0, 0], confirming the minimum.
Elimination
Eliminate the wrong options
Setting weight_decay=λ in a PyTorch optimizer implements:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: weight_decay adds the L2 penalty λ‖w‖² to the loss — exactly ridge regression (Lesson 12). It pulls parameters toward zero; in our run it shrank the intercept from 1.00 to 0.76.
Check
It shrank our intercept from 1.00 to 0.76. What is it?
Check your understanding
Setting weight_decay=λ in a PyTorch optimizer implements:
Answer: A
Why: weight_decay adds the L2 penalty λ‖w‖² to the loss — exactly ridge regression (Lesson 12). It pulls parameters toward zero; in our run it shrank the intercept from 1.00 to 0.76.
Section
Part 6 of 6 — the project
Concept
Fit ŷ = w₀ + w₁·x on our data x = [1,2,3,4], y = [3,4,7,8] three ways and prove they agree, then add weight_decay. You derived every piece — now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | closed-form OLS | np.linalg.solve |
| 2 | PyTorch training loop | nn.Linear, MSELoss, SGD |
| 3 | verify match; add weight_decay | compare slope/intercept |
Build rules: type every line yourself, reshape inputs to (n, 1) for nn.Linear, seed with torch.manual_seed(0), and never forget zero_grad() at the top of the loop.
Counterexample
Discussion prompt
Fit ŷ = w₀ + w₁·x on our data x = [1,2,3,4], y = [3,4,7,8] three ways and prove they agree, then add weight_decay. You derived every piece — now assemble it yourself.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: type every line yourself, reshape inputs to (n, 1) for nn.Linear, seed with torch.manual_seed(0), and never forget zero_grad() at the top of the loop.
Worked example
Your turn: build X with a ones column and solve the normal equations. Say the intercept and slope you expect (from the by-hand [1.0, 1.8]) before you print.
Hint: X = np.c_[np.ones(4), x], then np.linalg.solve(X.T@X, X.T@y). Do not use inv.
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
print("closed-form w =", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())| param | value (verified) |
|---|---|
| intercept w₀ | 1.0 |
| slope w₁ | 1.8 |
Worked example
Your turn: fit with nn.Linear + SGD. Predict whether the loss reaches near 0.2 (the minimum MSE) before you run it.
Hint: reshape to (4, 1) with unsqueeze(1); loop zero_grad → loss → backward → step for 1000 epochs at lr=0.1.
import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for _ in range(1000):
opt.zero_grad(); torch.nn.MSELoss()(model(xt), yt).backward(); opt.step()
print("slope =", round(model.weight.item(), 4), " intercept =", round(model.bias.item(), 4))| param | value (verified) |
|---|---|
| slope | 1.8 |
| intercept | 1.0 |
Worked example
Your turn: confirm SGD matched the closed form, then add weight_decay=0.5 and predict which way the intercept moves before printing.
Hint: SGD(..., weight_decay=0.5). The penalty pulls weights toward 0, so the intercept should shrink below 1.0.
import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
torch.manual_seed(0)
m = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
for _ in range(1000):
opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
print(f"wd={wd}: slope={m.weight.item():.4f} intercept={m.bias.item():.4f}")| setting | slope | intercept (verified) |
|---|---|---|
| wd = 0.0 (≈ OLS) | 1.8000 | 1.0000 |
| wd = 0.5 (ridge) | 1.8182 | 0.7636 |
Trade off
Comparison matrix
From Milestone 3 — verify & regularize: every row here is a choice with a cost. Fill the slope column, then say which row you would actually pick and what you give up for it.
| setting | slope | intercept (verified) |
|---|---|---|
| wd = 0.0 (≈ OLS) | 1.8000 | 1.0000 |
| wd = 0.5 (ridge) | 1.8182 | 0.7636 |
Concept
import numpy as np, torch
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
print("OLS closed form:", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())
torch.manual_seed(0)
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for _ in range(1000):
opt.zero_grad(); torch.nn.MSELoss()(model(xt), yt).backward(); opt.step()
print("PyTorch SGD: ", round(model.bias.item(), 4), round(model.weight.item(), 4))| printed line | value (verified) |
|---|---|
| OLS closed form | [1.0, 1.8] (intercept, slope) |
| PyTorch SGD (intercept, slope) | 1.0 1.8 |
| match | True |
If your closed form reads [1.0, 1.8] and the training loop lands on the same two numbers — you built the loop every model in Phase 2 reuses.
Comparison
Comparison matrix
From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.
| printed line | value (verified) |
|---|---|
| OLS closed form | [1.0, 1.8] (intercept, slope) |
| PyTorch SGD (intercept, slope) | 1.0 1.8 |
| match | True |
Concept
Slides closed, out loud: explain (1) why closed-form OLS and gradient descent reach the same w, (2) the five steps of the training loop and what backward() and step() each do, and (3) what weight_decay does to the weights.
Stretch: add a train/validation split and track both losses; fit [1, x, x², x³] and watch degree 3 interpolate the 4 points (SSE → 0) while a held-out point gets worse. Next up: logistic regression — the same loop, a new loss.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — One dataset, the closed form · The same w, by gradient descent · The PyTorch training loop · Regularize & bend the line · Score the fit · Your turn: train it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
XᵀX = [[4,10],[10,30]], Xᵀy = [22,64], det = 20, w = [1.0, 1.8]∇L = (2/n)Xᵀ(Xw − y) and step gradient descent by hand until it lands on that same [1.0, 1.8]zero_grad → forward → loss → backward → step — and hand-check one step against autogradzero_grad() and a too-large learning rate from real divergenceweight_decay, fit polynomial features (and spot overfitting), and score with R² = 0.9529, RMSE = 0.4472, matching sklearn| idea | the one thing to remember |
|---|---|
| closed form vs GD | same w [1.0, 1.8]; GD scales & generalizes |
| MSE gradient | ∇L = (2/n)Xᵀ(Xw − y) |
| training loop | zero_grad FIRST, every step |
| learning rate | too big diverges; if loss rises, halve η |
| weight_decay | = L2 / ridge; shrinks weights toward 0 |
| evaluation | R²/RMSE + residual plot; match a baseline |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.