Lesson 40: Linear Regression & the PyTorch Training Loop

USAAIO Lesson 40, from Week 14 of Phase 2, fully worked on ONE tiny dataset, x=[1,2,3,4] and y=[3,4,7,8]. Closed-form OLS is solved as a 2x2 system by hand, entry by entry, the MSE gradient is derived term by term, and gradient descent is stepped by hand until it lands on the closed-form w=[1.0,1.8]. The canonical PyTorch training loop is then dissected step by step - tensors, nn.Linear, MSELoss, autograd, and SGD - with one step hand-checked and a full loss trace to convergence, and the zero_grad and learning-rate traps are shown through real divergence. It closes with weight_decay and ridge, polynomial features, R2 and RMSE, and a match against sklearn. Every snippet runs standalone in a fresh interpreter, and every number came from real torch and numpy execution. The lesson runs to 63 slides.

Subject: Machine Learning · 108 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Linear Regression & the PyTorch Training Loop

Title

USAAIO · Lesson 40 · Week 14 (Phase 2 begins)

The first real model, trained two ways. We solve one tiny dataset in closed form by hand, then rebuild the answer with gradient descent, then wrap it in the PyTorch training loop every later model reuses — no step skipped, every number from real execution.

2. By the end of this lesson you can

Objectives

  1. Solve closed-form OLS by hand — build XᵀX, Xᵀy, invert the 2×2, and read off w
  2. Derive the MSE gradient ∇L = (2/n)Xᵀ(Xw − y) term by term and step gradient descent by hand until it lands on the closed-form w
  3. Write the canonical PyTorch training loop (nn.Linear, MSELoss, SGD) and hand-check a single step against autograd
  4. Diagnose the two classic bugs — a missing zero_grad() and a too-large learning rate — from real divergence
  5. Add L2 regularization with weight_decay, fit polynomial features, and score the fit with R² and RMSE, matching sklearn

3. What survived from Mock Exam — Phase 1 Coding?

Warm-up

Discussion prompt

Before we open Lesson 40: Linear Regression & the PyTorch Training Loop: without looking back, what was the main idea of Mock Exam — Phase 1 Coding, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

a timed coding mock exam — implement PCA, gradient descent, MLE, and k-fold cross-validation from scratch in NumPy, each with a theory sub-question, verified against the libraries, plus a code-review pass. The Phase-1 coding milestone.

4. One dataset, the closed form

Section

Part 1 of 6 — recap & by-hand OLS

5. The running example

Concept

We keep one tiny dataset for the whole lesson so every number is checkable by hand. Four points — think of x as hours studied and y as a score:

ixy
113
224
337
448

We want the best line ŷ = w₀ + w₁·x — an intercept w₀ and a slope w₁. Four points, two unknowns: overdetermined, so no line hits all four.

6. Fill in: x for The running example

Comparison

Comparison matrix

From The running example: refill the x column from what you know. The rest of the table is as it appeared.

ixy
113
224
337
448

7. Best line = smallest total squared miss

Intuition

Since no line is exact, we pick the one whose predictions miss by the least overall. 'Least' means the sum of squared misses — squaring makes every error positive and punishes big misses hardest.

That single number, averaged over the points, is the mean squared error — the loss we minimize both by algebra (closed form) and by iteration (gradient descent). Two roads, same destination.

8. The design matrix and MSE

Concept

Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:

\[ X = \begin{bmatrix} 1 & 1 \\ 1 & 2 \\ 1 & 3 \\ 1 & 4 \end{bmatrix}, \quad w = \begin{bmatrix} w_0 \\ w_1 \end{bmatrix}, \quad y = \begin{bmatrix} 3 \\ 4 \\ 7 \\ 8 \end{bmatrix} \]

\[ L(w) = \frac{1}{n}\lVert Xw - y \rVert^2 = \frac{1}{n}\sum_{i=1}^{n}\big(w_0 + w_1 x_i - y_i\big)^2 \]

Here n = 4; the 1/n makes it a MEAN squared error

Why: Dividing by n makes the loss independent of dataset size, so a learning rate that works on 4 points also works on 4 million. PyTorch's MSELoss uses this 1/n mean by default.

9. Break it if you can: The design matrix and MSE

Counterexample

Discussion prompt

Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

10. The closed form (Lesson 7, in one line)

Concept

Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:

\[ \boxed{\,X^\top X\, w^\star = X^\top y\,} \]

We'll solve this 2×2 by hand this lesson, then spend most of the deck reaching the same w iteratively — because iteration is what generalizes to every model with no closed form.

11. By analogy: The closed form (Lesson 7, in one line)

Analogy

Discussion prompt

Explain The closed form (Lesson 7, in one line) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:

12. What has to happen first: Build XᵀX, entry by entry

Ranking

Put in order

Put the moves of Build XᵀX, entry by entry into the order they have to happen.

  1. Top-left = ones·ones = Σ1 = 4
  2. Off-diagonal = ones·x = Σxᵢ = 1+2+3+4 = 10
  3. Bottom-right = x·x = Σxᵢ² = 1+4+9+16 = 30

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Four 1s dotted with four 1s: 1+1+1+1 = 4.

13. Build XᵀX, entry by entry

Worked example

(XᵀX)ⱼₖ = column j of X dotted with column k. Column 0 is ones, column 1 is x, so there are only three distinct entries (it is symmetric):

Top-left = ones·ones = Σ1 = 4

Why: Four 1s dotted with four 1s: 1+1+1+1 = 4. This entry is just n, the number of points.

Off-diagonal = ones·x = Σxᵢ = 1+2+3+4 = 10

Why: The ones vector dotted with x simply sums x. It fills both off-diagonal slots by symmetry.

Bottom-right = x·x = Σxᵢ² = 1+4+9+16 = 30

Why: Sum of squared features.

\[ X^\top X = \begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} \]

14. Decode the notation: Build XᵀX, entry by entry

Notation

Annotate

From Build XᵀX, entry by entry — read this one piece at a time. What is each part doing?

On: \( X^\top X = \begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} \)

  • Four 1s dotted with four 1s: 1+1+1+1 = 4. This entry is just n, the number of points.
  • The ones vector dotted with x simply sums x. It fills both off-diagonal slots by symmetry.

15. Build Xᵀy, entry by entry

Worked example

Top = ones·y = Σyᵢ = 3+4+7+8 = 22

Why: The ones row of Xᵀ dotted with y sums the targets.

Bottom = x·y = Σxᵢyᵢ

Why: The x row of Xᵀ dotted with y: multiply pairwise, then add.

\[ \sum_i x_i y_i = 1\cdot3 + 2\cdot4 + 3\cdot7 + 4\cdot8 = 3 + 8 + 21 + 32 = 64 \]

\[ X^\top y = \begin{bmatrix} 22 \\ 64 \end{bmatrix} \]

16. Draw the shape of it: Build Xᵀy, entry by entry

Blank canvas

Draw it

Draw what Build Xᵀy, entry by entry just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

17. Guess the shape of the answer: Invert the 2×2 — determinant first

Estimation

Predict first

For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:

Commit before you compute: what does Invert the 2×2 — determinant first come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: det = 20 ≠ 0 → invertible, unique solution

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique.

18. Invert the 2×2 — determinant first

Worked example

For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:

\[ \det\!\begin{bmatrix} 4 & 10 \\ 10 & 30 \end{bmatrix} = 4\cdot 30 - 10\cdot 10 = 120 - 100 = 20 \]

det = 20 ≠ 0 → invertible, unique solution

Why: A nonzero determinant guarantees XᵀX is invertible, so w★ exists and is unique. (det = 0 would mean collinear columns — the singular case ridge fixes.)

\[ (X^\top X)^{-1} = \frac{1}{20}\begin{bmatrix} 30 & -10 \\ -10 & 4 \end{bmatrix} = \begin{bmatrix} 1.5 & -0.5 \\ -0.5 & 0.2 \end{bmatrix} \]

19. Work backwards from the answer: Invert the 2×2 — determinant first

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

det = 20 ≠ 0 → invertible, unique solution

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

For [[a, b], [c, d]] the inverse is 1/det · [[d, −b], [−c, a]]. First the determinant det = ad − bc:

20. What has to be given first: Multiply through to get w★

Missing information

Discussion prompt

w★ = (XᵀX)⁻¹ Xᵀy. Dot each row of the inverse with [22, 64], one entry at a time:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

First row of the inverse dotted with Xᵀy. Equivalently (30·22 − 10·64)/20 = (660 − 640)/20 = 20/20 = 1.0.

21. Multiply through to get w★

Worked example

w★ = (XᵀX)⁻¹ Xᵀy. Dot each row of the inverse with [22, 64], one entry at a time:

w₀ = 1.5·22 + (−0.5)·64 = 33 − 32 = 1.0

Why: First row of the inverse dotted with Xᵀy. Equivalently (30·22 − 10·64)/20 = (660 − 640)/20 = 20/20 = 1.0.

w₁ = (−0.5)·22 + 0.2·64 = −11 + 12.8 = 1.8

Why: Second row dotted with Xᵀy. Equivalently (−10·22 + 4·64)/20 = (−220 + 256)/20 = 36/20 = 1.8.

\[ w^\star = \begin{bmatrix} 1.0 \\ 1.8 \end{bmatrix} \;\Longrightarrow\; \widehat{y} = 1.0 + 1.8\,x \]

22. Say it in words: Multiply through to get w★

Translation

\( w^\star = \begin{bmatrix} 1.0 \\ 1.8 \end{bmatrix} \;\Longrightarrow\; \widehat{y} = 1.0 + 1.8\,x \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

23. Predict the next row: Confirm w in NumPy

Pattern

Predict first

The table runs: XtX | [[4.0, 10.0], [10.0, 30.0]] · Xty | [22.0, 64.0]

In Confirm w in NumPy, given the rows so far: what is the next one — the row where object is w = [w₀, w₁]?

Correct: w = [w₀, w₁] | [1.0, 1.8]

objectprinted value (verified)
XtX[[4.0, 10.0], [10.0, 30.0]]
Xty[22.0, 64.0]
w = [w₀, w₁][1.0, 1.8]

Why: The relationship between the columns, not the individual numbers, is what generates the next row. np.linalg.solve factors the 2×2 (LU) and never forms an explicit inverse — faster and more numerically stable than inv.

24. Confirm w in NumPy

Worked example

The whole by-hand derivation is one solve call. This block is self-contained — it re-imports NumPy and re-defines the data:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
XtX, Xty = X.T @ X, X.T @ y
w = np.linalg.solve(XtX, Xty)   # never inv
print("XtX =", XtX.tolist())
print("Xty =", Xty.tolist())
print("w   =", w.round(4).tolist())

solve returns exactly the by-hand [1.0, 1.8]

Why: np.linalg.solve factors the 2×2 (LU) and never forms an explicit inverse — faster and more numerically stable than inv. The printed w matches our hand math to the digit.

objectprinted value (verified)
XtX[[4.0, 10.0], [10.0, 30.0]]
Xty[22.0, 64.0]
w = [w₀, w₁][1.0, 1.8]

25. What each one costs: Confirm w in NumPy

Trade off

Comparison matrix

From Confirm w in NumPy: every row here is a choice with a cost. Fill the printed value (verified) column, then say which row you would actually pick and what you give up for it.

objectprinted value (verified)
XtX[[4.0, 10.0], [10.0, 30.0]]
Xty[22.0, 64.0]
w = [w₀, w₁][1.0, 1.8]

26. Predictions and residuals — full trace

Worked example

Plug each x into ŷ = 1.0 + 1.8x, then subtract from the true y to get every residual and its square:

xyŷ = 1 + 1.8xr = y − ŷr²
132.8+0.20.04
244.6−0.60.36
376.4+0.60.36
488.2−0.20.04

Σr² = 0.04 + 0.36 + 0.36 + 0.04 = 0.8

Why: This is the minimized ‖Xw − y‖² = 0.8 — the smallest total squared error any line achieves on this data. The MSE is 0.8/4 = 0.2. We reuse both for R² and RMSE.

27. Watch it run: Predictions and residuals — full trace

Pattern

Step through it

Step through Predictions and residuals — full trace one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: x is 1
  2. Step 2: x is 2
  3. Step 3: x is 3
  4. Step 4: x is 4

28. The residual is perpendicular (a free check)

Worked example

The normal equations promise Xᵀr = 0: the residual is perpendicular to every column of X. Confirm it on r = [0.2, −0.6, 0.6, −0.2]:

Row 0 (ones·r): 0.2 − 0.6 + 0.6 − 0.2 = 0

Why: Residuals sum to exactly zero — a direct consequence of the bias column. The fit neither systematically over- nor under-predicts.

Row 1 (x·r): 1(0.2) + 2(−0.6) + 3(0.6) + 4(−0.2) = 0.2 − 1.2 + 1.8 − 0.8 = 0

Why: Perpendicular to the feature column too. Xᵀr = [0, 0], so w is truly the least-squares solution — this is your always-available correctness check.

conditionvalue (verified)
Σ rᵢ0.0
Σ xᵢ rᵢ0.0
→ Xᵀr[0, 0]

29. The same w, by gradient descent

Section

Part 2 of 6 — derive & step by hand

30. Why iterate when we have the exact answer?

Concept

The closed form needs to solve a d×d system XᵀX w = Xᵀy. That is fine for d = 2, but costs O(d³) and only exists for linear least squares.

Gradient descent instead nudges w downhill on the loss, step by step. It scales to huge d and n, and — crucially — works for every model in Phase 2 (logistic regression, MLPs, transformers) where no closed form exists.

So we practice it here, on a problem where we already know the answer [1.0, 1.8] — the perfect place to watch iteration converge to truth.

31. Teach it back: Why iterate when we have the exact answer?

Explain it

Discussion prompt

Explain Why iterate when we have the exact answer? to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

The closed form needs to solve a d×d system XᵀX w = Xᵀy. That is fine for d = 2, but costs O(d³) and only exists for linear least squares.

32. The update rule

Concept

Gradient descent repeatedly steps against the gradient (the uphill direction), scaled by a learning rate η:

\[ w \;\leftarrow\; w - \eta\,\nabla L(w) \]

The gradient points uphill, so subtracting it walks downhill. η sets the step size: too small and it crawls, too large and it overshoots and diverges. To run it we need ∇L explicitly — derive it next.

33. What has to happen first: Derive the MSE gradient, term by term

Ranking

Put in order

Put the moves of Derive the MSE gradient, term by term into the order they have to happen.

  1. Start from L(w) = (1/n)‖Xw − y‖²
  2. Differentiate the squared norm: ∇‖Xw − y‖² = 2Xᵀ(Xw − y)
  3. Carry the 1/n through

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The mean squared error on the design matrix — the exact loss PyTorch's MSELoss computes.

34. Derive the MSE gradient, term by term

Worked example

Start from L(w) = (1/n)‖Xw − y‖²

Why: The mean squared error on the design matrix — the exact loss PyTorch's MSELoss computes.

\[ L(w) = \tfrac{1}{n}(Xw - y)^\top (Xw - y) \]

Differentiate the squared norm: ∇‖Xw − y‖² = 2Xᵀ(Xw − y)

Why: Chain rule on vᵀv with v = Xw − y: the outer derivative gives 2v, the inner derivative of v w.r.t. w is Xᵀ (from Lesson 7's ∇(wᵀAw) = 2Aw with A = XᵀX).

\[ \nabla \lVert Xw - y \rVert^2 = 2\,X^\top(Xw - y) \]

Carry the 1/n through

Why: The 1/n is a constant multiplier, so it rides along untouched into the gradient.

\[ \boxed{\,\nabla L(w) = \tfrac{2}{n}\,X^\top(Xw - y)\,} \]

35. Draw the shape of it: Derive the MSE gradient, term by term

Blank canvas

Draw it

Draw what Derive the MSE gradient, term by term just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

36. Restore the missing line: The gradient in code, checked two ways

Fill the middle

Fill in the blanks

From The gradient in code, checked two ways — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
def grad(w): return (2.0/n) * (X.T @ (X @ w - y))
print("grad at [0, 0] =", grad(np.array([0., 0.])).tolist())
print("grad at [1.0, 1.8] =", grad(np.array([1.0, 1.8])).round(6).tolist())

Why: y is what everything below it consumes, so the wrong expression here fails later and somewhere else. grad([0,0]) = (2/4)·Xᵀ(−y) = 0.5·(−[22,64]) = [−11, −32]: strongly downhill.

37. The gradient in code, checked two ways

Worked example

Evaluate ∇L = (2/n)Xᵀ(Xw − y) at the start w = [0, 0] and at the known optimum w = [1.0, 1.8]. Self-contained block:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
def grad(w): return (2.0/n) * (X.T @ (X @ w - y))
print("grad at [0, 0]     =", grad(np.array([0., 0.])).tolist())
print("grad at [1.0, 1.8] =", grad(np.array([1.0, 1.8])).round(6).tolist())

At the start the gradient is steep; at the optimum it is zero

Why: grad([0,0]) = (2/4)·Xᵀ(−y) = 0.5·(−[22,64]) = [−11, −32]: strongly downhill. grad([1.0,1.8]) = [0, 0]: flat, confirming [1.0, 1.8] is the minimum GD must find.

w∇L (verified)
[0, 0][−11.0, −32.0]
[1.0, 1.8][0.0, 0.0]

38. Step 1 of gradient descent, by hand

Worked example

Start at w = [0, 0] with learning rate η = 0.1. We just found ∇L([0,0]) = [−11, −32]. Apply the update:

w₀: 0 − 0.1·(−11) = 0 + 1.1 = 1.1

Why: Subtracting a negative gradient moves the intercept UP toward its target of 1.0 — the right direction.

w₁: 0 − 0.1·(−32) = 0 + 3.2 = 3.2

Why: The slope jumps to 3.2 — overshooting the target 1.8, because the gradient was so steep at the start. GD will correct this on the next step.

\[ w^{(1)} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} - 0.1\begin{bmatrix} -11 \\ -32 \end{bmatrix} = \begin{bmatrix} 1.1 \\ 3.2 \end{bmatrix} \]

39. Decode the notation: Step 1 of gradient descent, by hand

Notation

Annotate

From Step 1 of gradient descent, by hand — read this one piece at a time. What is each part doing?

On: \( w^{(1)} = \begin{bmatrix} 0 \\ 0 \end{bmatrix} - 0.1\begin{bmatrix} -11 \\ -32 \end{bmatrix} = \begin{bmatrix} 1.1 \\ 3.2 \end{bmatrix} \)

  • Subtracting a negative gradient moves the intercept UP toward its target of 1.0 — the right direction.
  • The slope jumps to 3.2 — overshooting the target 1.8, because the gradient was so steep at the start. GD will correct this on the next step.

40. Finish it with less help: Watch four steps descend

Faded example

Fill in the blanks

Watch four steps descend, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(4):
g = (2.0/n) * (X.T @ (X @ w - y))
w = w - eta * g
loss = ((X @ w - y)2).mean()**
print(f"step ___: w = ___ loss = ___")

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Each step over/undershoots the target 1.8, but the loss shrinks 15.6 → 7.1 → 3.3 → 1.6 — the bowl is convex, so downhill always makes progress even when a single coordinate zig-zags.

41. Watch four steps descend

Worked example

Run the same by-hand update in a loop and print w and the loss each step. Self-contained:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(4):
    g = (2.0/n) * (X.T @ (X @ w - y))
    w = w - eta * g
    loss = ((X @ w - y)**2).mean()
    print(f"step {step+1}: w = {w.round(4).tolist()}  loss = {loss:.4f}")

The slope oscillates (3.2 → 1.05 → 2.49 → 1.52) but the loss falls every step

Why: Each step over/undershoots the target 1.8, but the loss shrinks 15.6 → 7.1 → 3.3 → 1.6 — the bowl is convex, so downhill always makes progress even when a single coordinate zig-zags.

stepw = [w₀, w₁]loss (verified)
1[1.1, 3.2]15.6100
2[0.38, 1.05]7.1281
3[0.879, 2.485]3.3194
4[0.5607, 1.518]1.6088

42. Run it to convergence

Worked example

Keep stepping. After enough iterations GD reaches the closed-form answer. Same loop, more steps:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
n = 4
w = np.array([0., 0.]); eta = 0.1
for step in range(2000):
    w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
print("GD w         =", w.round(4).tolist())
print("closed form  =", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())

GD lands on [1.0, 1.8] — identical to the closed form

Why: Because MSE is convex with a single global minimum, gradient descent from any start converges to the SAME w the normal equations give. Two derivations, one answer — proven numerically.

methodw (verified)
gradient descent (2000 steps)[1.0, 1.8]
closed-form solve[1.0, 1.8]
match✓

43. Something is wrong here: a learning rate that is too large

Anomaly

Predict first

A student writes this, and it looks reasonable:

Bigger steps must reach the minimum faster — crank η up to 0.5.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up.

Keep η small enough that each step actually decreases the loss — here 0.1 works.

Why: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up. Too-large η diverges — the opposite of what you wanted.

44. Trap: a learning rate that is too large

Trap

The trap

Bigger steps must reach the minimum faster — crank η up to 0.5.

import numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]; n = 4
w = np.array([0., 0.])
eta = 0.5          # too big
for step in range(4):
    w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
    print(f"step {step+1}: loss = {((X@w - y)**2).mean():.2f}")
steploss (verified)
11852.25
2100060.09
35405933.68
4292066268.61

Loss EXPLODES to infinity

Why: Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up. Too-large η diverges — the opposite of what you wanted.

The fix

Keep η small enough that each step actually decreases the loss — here 0.1 works.

import numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]; n = 4
w = np.array([0., 0.])
eta = 0.1          # stable
for step in range(4):
    w = w - eta * (2.0/n) * (X.T @ (X @ w - y))
    print(f"step {step+1}: loss = {((X@w - y)**2).mean():.2f}")
steploss (verified)
115.61
27.13
33.32
41.61

Loss falls monotonically toward 0.2

Why: A small enough step size guarantees descent on a convex loss. If your loss ever increases or turns to NaN, halve η first — it's the #1 cause of a 'broken' training run.

45. Inspect it line by line: Trap: a learning rate that is too large

Error analysis

Annotate

Walk the callouts on Trap: a learning rate that is too large. Each one is a place this is easy to get subtly wrong.

  • Each step overshoots the minimum and lands farther out than it started, so the gradient grows, the next step overshoots more, and the loss blows up. Too-large η diverges — the opposite of what you wanted.
  • A small enough step size guarantees descent on a convex loss. If your loss ever increases or turns to NaN, halve η first — it's the #1 cause of a 'broken' training run.

46. The PyTorch training loop

Section

Part 3 of 6 — every line

47. Let the framework compute the gradient

Intuition

We hand-derived ∇L because the model was simple. Real models have millions of parameters — deriving gradients by hand is hopeless.

PyTorch's autograd computes the gradient of any loss automatically by tracking the operations you run. Your job shrinks to: define the model, define the loss, and run a fixed five-step loop. That loop is the same for every model you will ever train.

48. Step A — data as tensors of shape (n, 1)

Worked example

nn.Linear expects a 2-D input: one row per sample, one column per feature. unsqueeze(1) turns a length-4 vector into a (4, 1) column:

import torch
import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
print("xt.shape =", tuple(xt.shape))
print("yt.shape =", tuple(yt.shape))

Both become shape (4, 1) — 4 rows, 1 column

Why: Feed a flat (4,) vector to nn.Linear and the shapes silently mis-broadcast, giving a wrong loss with no error. Always reshape features to (n, features) — here (4, 1).

tensorshape (verified)
xt(4, 1)
yt(4, 1)
xt values[1.0, 2.0, 3.0, 4.0]

49. Guess the shape of the answer: Step B — nn.Linear is ŷ = wx + b

Estimation

Predict first

nn.Linear(1, 1) holds one weight and one bias, initialized randomly. Seed first so the numbers are reproducible:

Commit before you compute: what does Step B — nn.Linear is ŷ = wx + b come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: It starts with random weight −0.0075, bias 0.5364 — a bad line

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. With torch.manual_seed(0), nn.Linear(1,1) initializes to these exact values.

50. Step B — nn.Linear is ŷ = wx + b

Worked example

nn.Linear(1, 1) holds one weight and one bias, initialized randomly. Seed first so the numbers are reproducible:

import torch
torch.manual_seed(0)
layer = torch.nn.Linear(1, 1)
print("weight =", round(layer.weight.item(), 4))
print("bias   =", round(layer.bias.item(), 4))
xt = torch.tensor([[1.], [2.], [3.], [4.]])
print("forward =", layer(xt).squeeze().round(decimals=4).tolist())

It starts with random weight −0.0075, bias 0.5364 — a bad line

Why: With torch.manual_seed(0), nn.Linear(1,1) initializes to these exact values. The forward pass ŷ = weight·x + bias is nowhere near the data yet — training's whole job is to move these two numbers to [1.8, 1.0].

parameterinitial value (seed 0)
weight (slope)−0.0075
bias (intercept)0.5364
forward on x=[1,2,3,4][0.529, 0.521, 0.514, 0.507]

51. Where does each piece belong: Lesson 40: Linear Regression & the PyTorch…

Sorting

Sort into buckets

These are the pieces of Lesson 40: Linear Regression & the PyTorch Training Loop, out of order. Put each one back under the part of the lesson it belongs to.

One dataset, the closed form
The running example; Best line = smallest total squared miss; The design matrix and MSE
The same w, by gradient descent
Why iterate when we have the exact answer?; The update rule; Derive the MSE gradient, term by term
The PyTorch training loop
Let the framework compute the gradient; Step A — data as tensors of shape (n, 1); Step B — nn.Linear is ŷ = wx + b
s1
One dataset, the closed form is where Lesson 40: Linear Regression & the PyTorch Training Loop puts The running example, Best line = smallest total squared miss, The design matrix and MSE. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
The same w, by gradient descent is where Lesson 40: Linear Regression & the PyTorch Training Loop puts Why iterate when we have the exact answer?, The update rule, Derive the MSE gradient, term by term. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
The PyTorch training loop is where Lesson 40: Linear Regression & the PyTorch Training Loop puts Let the framework compute the gradient, Step A — data as tensors of shape (n, 1), Step B — nn.Linear is ŷ = wx + b. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

52. Step C — MSELoss scores the fit

Worked example

nn.MSELoss() computes the mean of the squared differences — exactly the L(w) we minimized. A quick check with hand-picked predictions:

import torch
pred   = torch.tensor([[2.], [3.], [6.], [7.]])
target = torch.tensor([[3.], [4.], [7.], [8.]])
loss = torch.nn.MSELoss()(pred, target)
print("MSE =", loss.item())

Every error is −1, squared is 1, mean of four 1s is 1.0

Why: MSELoss = mean((pred − target)²). Here pred is 1 below target everywhere, so each squared error is 1 and the mean is exactly 1.0 — confirming MSELoss is the same 1/n·Σ(ŷ−y)² we minimize.

predtarget(pred − target)²
231
341
671
781

53. What stays fixed: Step C — MSELoss scores the fit

Invariant

Step through it

Step through Step C — MSELoss scores the fit one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: pred is 2
  2. Step 2: pred is 3
  3. Step 3: pred is 6
  4. Step 4: pred is 7

54. The five-line rhythm

Concept

Every PyTorch model trains with the same five steps per iteration. Memorize this order — it never changes:

  1. opt.zero_grad() — clear last step's gradients
  2. out = model(xt) — forward pass (predict)
  3. loss = lossfn(out, yt) — score the predictions
  4. loss.backward() — autograd fills every .grad
  5. opt.step() — the optimizer updates the parameters

Steps 4–5 are exactly our w ← w − η∇L: backward() computes ∇L, step() applies the subtraction. PyTorch just automates the gradient.

55. Restore the missing line: One iteration, hand-checked against autograd

Fill the middle

Fill in the blanks

From One iteration, hand-checked against autograd — one line has had its right-hand side removed. Put it back.

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward()
print("weight.grad =", round(model.weight.grad.item(), 4))
print("bias.grad =", round(model.bias.grad.item(), 4))
opt.step()
print("weight now =", round(model.weight.item(), 4))
print("bias now =", round(model.bias.item(), 4))

Why: loss is what everything below it consumes, so the wrong expression here fails later and somewhere else. SGD applied w ← w − lr·grad exactly.

56. One iteration, hand-checked against autograd

Worked example

Run one loop iteration and print the gradients autograd computed and the updated weights. We'll check they obey w ← w − η·grad:

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
opt.zero_grad()
loss = torch.nn.MSELoss()(model(xt), yt)
loss.backward()
print("weight.grad =", round(model.weight.grad.item(), 4))
print("bias.grad   =", round(model.bias.grad.item(), 4))
opt.step()
print("weight now  =", round(model.weight.item(), 4))
print("bias now    =", round(model.bias.item(), 4))

Check the update: −0.0075 − 0.1·(−29.4301) = 2.9355 ✓

Why: SGD applied w ← w − lr·grad exactly. weight: −0.0075 − 0.1·(−29.4301) = 2.9355 (printed). bias: 0.5364 − 0.1·(−9.9645) = 1.5329 (printed). Autograd's gradient plugged straight into our by-hand rule.

quantityvalue (verified)
weight.grad−29.4301
bias.grad−9.9645
weight after step2.9355
bias after step1.5329

57. The full loop — train to convergence

Worked example

Wrap the five steps in a for loop and run it. Watch it reach the closed-form [1.0, 1.8]:

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
lossfn = torch.nn.MSELoss()
for epoch in range(1000):
    opt.zero_grad()
    loss = lossfn(model(xt), yt)
    loss.backward()
    opt.step()
print("final loss =", round(loss.item(), 4))
print("slope      =", round(model.weight.item(), 4))
print("intercept  =", round(model.bias.item(), 4))

loss → 0.2, slope → 1.8, intercept → 1.0

Why: The training loop lands on the exact closed-form solution: slope 1.8, intercept 1.0, minimum MSE 0.2 (= 0.8/4). The framework did what our hand-GD did, with autograd computing the gradient for us.

sourceslopeinterceptloss
closed-form OLS1.81.00.2
PyTorch SGD (1000 ep)1.81.00.2
match✓✓✓

58. Fill in: intercept for The full loop — train to convergence

Comparison

Comparison matrix

From The full loop — train to convergence: refill the intercept column from what you know. The rest of the table is as it appeared.

sourceslopeinterceptloss
closed-form OLS1.81.00.2
PyTorch SGD (1000 ep)1.81.00.2
match✓✓✓

59. A loss trace: watch it converge

Worked example

Print the loss and parameters at chosen epochs to see the descent curve — fast at first, then fine-tuning:

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
lossfn = torch.nn.MSELoss()
for epoch in range(1001):
    opt.zero_grad()
    loss = lossfn(model(xt), yt)
    loss.backward()
    opt.step()
    if epoch in (0, 1, 10, 100, 1000):
        print(f"epoch {epoch:4d}: loss={loss.item():.4f} slope={model.weight.item():.4f} intercept={model.bias.item():.4f}")

loss 29.1 → 0.2 in about 100 epochs; the rest is polish

Why: The loss drops fastest early (steep gradient), then the slope/intercept settle onto 1.8/1.0. By epoch 100 the loss is already 0.2000 to four places — later epochs only refine the trailing digits.

epochlossslopeintercept
029.10682.93551.5329
113.18010.96580.8586
100.21131.78851.1043
1000.20001.79791.0063
10000.20001.80001.0000

60. Watch it run: A loss trace: watch it converge

Pattern

Step through it

Step through A loss trace: watch it converge one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 1
  3. Step 3: epoch is 10
  4. Step 4: epoch is 100
  5. Step 5: epoch is 1000

61. Something is wrong here: forgetting zero_grad()

Anomaly

Predict first

A student writes this, and it looks reasonable:

The gradient is fresh each pass, so just backward() then step() — skip zero_grad().

It is wrong. Say what breaks — and say it before you turn the page.

Correct: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad.

Call opt.zero_grad() at the top of every iteration.

Why: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad. Without zeroing, step 2 uses grad₁+grad₂, so the optimizer takes wrong-sized steps and the loss bounces (43.30, then 49.75) instead of descending.

62. Trap: forgetting zero_grad()

Trap

The trap

The gradient is fresh each pass, so just backward() then step() — skip zero_grad().

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(6):
    # BUG: no opt.zero_grad()
    loss = torch.nn.MSELoss()(model(xt), yt)
    loss.backward(); opt.step()
    print(f"epoch {epoch}: loss = {loss.item():.2f}")
epochloss (verified)
029.11
113.18
243.30
32.27
449.75

Loss lurches around instead of falling smoothly

Why: PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad. Without zeroing, step 2 uses grad₁+grad₂, so the optimizer takes wrong-sized steps and the loss bounces (43.30, then 49.75) instead of descending.

The fix

Call opt.zero_grad() at the top of every iteration.

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for epoch in range(6):
    opt.zero_grad()
    loss = torch.nn.MSELoss()(model(xt), yt)
    loss.backward(); opt.step()
    print(f"epoch {epoch}: loss = {loss.item():.2f}")
epochloss (verified)
029.11
113.18
26.36
33.24
41.75

Loss falls monotonically — the fix is one line

Why: zero_grad() resets every .grad to zero so each step uses only THIS iteration's gradient. This is the single most common training-loop bug; put it first, every loop, always.

63. Inspect it line by line: Trap: forgetting zero_grad()

Error analysis

Annotate

Walk the callouts on Trap: forgetting zero_grad(). Each one is a place this is easy to get subtly wrong.

  • PyTorch ACCUMULATES gradients: each backward() ADDS to the existing .grad. Without zeroing, step 2 uses grad₁+grad₂, so the optimizer takes wrong-sized steps and the loss bounces (43.30, then 49.75) instead of descending.
  • zero_grad() resets every .grad to zero so each step uses only THIS iteration's gradient. This is the single most common training-loop bug; put it first, every loop, always.

64. Regularize & bend the line

Section

Part 4 of 6 — weight_decay & polynomials

65. L2 regularization = weight_decay

Concept

Ridge (Lesson 12) adds an L2 penalty on weight size to the loss. In PyTorch it is a one-word optimizer argument, weight_decay:

\[ \min_w \; \lVert Xw - y \rVert^2 + \lambda \lVert w \rVert^2 \;\equiv\; \texttt{SGD(..., weight\_decay=}\lambda\texttt{)} \]

Big weights now cost something, so the fit is pulled toward 0 — trading a little training accuracy for stability. It's ridge, wired into the optimizer.

66. Finish it with less help: weight_decay shrinks the fit

Faded example

Fill in the blanks

weight_decay shrinks the fit, with the scaffolding fading: two lines are gone now — fill both.

import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
torch.manual_seed(0)
m = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
for _ in range(1000):
opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
print(f"wd=___: slope=___ intercept=___")

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. With weight_decay=0.5 the penalty pulls the parameters toward zero, so the fit backs off the exact OLS solution.

67. weight_decay shrinks the fit

Worked example

Train the same loop with weight_decay = 0 and 0.5, and compare the learned slope and intercept:

import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
    torch.manual_seed(0)
    m = torch.nn.Linear(1, 1)
    opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
    for _ in range(1000):
        opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
    print(f"wd={wd}: slope={m.weight.item():.4f} intercept={m.bias.item():.4f}")

The intercept shrinks 1.00 → 0.76 under the penalty

Why: With weight_decay=0.5 the penalty pulls the parameters toward zero, so the fit backs off the exact OLS solution. The intercept drops from 1.00 to 0.76; the slope barely moves (1.80 → 1.82) — the penalty hits whichever weights it can shrink cheaply.

weight_decayslopeintercept (verified)
0.0 (plain OLS)1.80001.0000
0.5 (ridge)1.81820.7636

68. Polynomial features bend the line

Concept

Linear regression is only linear in the weights, not the features. Map x → [1, x, x², …, xᵈ] and the same machinery fits a curve — the design matrix just gets more columns.

\[ \hat y = w_0 + w_1 x + w_2 x^2 + \cdots + w_d x^d \]

The gradient, the loop, and weight_decay are unchanged — you only widen X. But more columns means more power to fit, and past a point, to overfit.

69. Restore the missing line: Degree-2 features on our data

Fill the middle

Fill in the blanks

From Degree-2 features on our data — one line has had its right-hand side removed. Put it back.

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
Xp = np.c_[np.ones(4), x, x2] # [1, x, x^2]
w = np.linalg.solve(Xp.T@Xp, Xp.T@y)
print("w (deg 2) =", w.round(4).tolist())
print("SSE =", round(((Xp@w - y)
2).sum(), 4))

Why: Xp is what everything below it consumes, so the wrong expression here fails later and somewhere else. OLS finds w = [1.0, 1.8, 0.0]: the quadratic term is exactly zero because our y = [3,4,7,8] has no curvature.

70. Degree-2 features on our data

Worked example

Add an x² column and solve OLS on the widened matrix. Self-contained:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
Xp = np.c_[np.ones(4), x, x**2]        # [1, x, x^2]
w = np.linalg.solve(Xp.T@Xp, Xp.T@y)
print("w (deg 2) =", w.round(4).tolist())
print("SSE       =", round(((Xp@w - y)**2).sum(), 4))

The x² coefficient comes out 0.0 — the data really is linear

Why: OLS finds w = [1.0, 1.8, 0.0]: the quadratic term is exactly zero because our y = [3,4,7,8] has no curvature. SSE stays 0.8, same as the line. Extra capacity that the data doesn't need sits idle here.

degreewSSE (verified)
1 (line)[1.0, 1.8]0.8
2 (parabola)[1.0, 1.8, 0.0]0.8

71. Predict the next row: Degree 3 overfits: SSE hits zero

Pattern

Predict first

The table runs: 1 | 2 | 0.800000 · 2 | 3 | 0.800000

In Degree 3 overfits: SSE hits zero, given the rows so far: what is the next one — the row where degree is 3?

Correct: 3 | 4 | 0.000000

degree#paramstraining SSE (verified)
120.800000
230.800000
340.000000

Why: The relationship between the columns, not the individual numbers, is what generates the next row. 4 params for 4 points interpolates exactly: zero training error.

72. Degree 3 overfits: SSE hits zero

Worked example

Push the degree up. With 4 points, a degree-3 polynomial has 4 parameters — enough to pass through every point exactly:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
for deg in (1, 2, 3):
    Xp = np.vander(x, deg+1, increasing=True)   # [1, x, ..., x^deg]
    w = np.linalg.lstsq(Xp, y, rcond=None)[0]
    sse = ((Xp@w - y)**2).sum()
    print(f"degree {deg}: SSE = {sse:.6f}, #params = {deg+1}")

Degree 3 drives SSE to 0 — but that is memorizing, not learning

Why: 4 params for 4 points interpolates exactly: zero training error. On NEW data such a wiggly curve does worse than the honest line — the definition of overfitting. Low training error is not the goal; low validation error is.

degree#paramstraining SSE (verified)
120.800000
230.800000
340.000000

73. Something is wrong here: picking degree by training error

Anomaly

Predict first

A student writes this, and it looks reasonable:

Higher degree gives lower SSE, so pick the degree with the smallest training error.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points.

Choose the degree with the smallest validation error — error on data the model did not train on.

Why: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points. Optimizing training error alone always picks the most complex, most overfit model.

74. Trap: picking degree by training error

Trap

The trap

Higher degree gives lower SSE, so pick the degree with the smallest training error.

Choose degree 3 because its training SSE = 0

Why: Training error always drops (or ties) as you add parameters — degree 3 hits 0 by interpolating the 4 points. Optimizing training error alone always picks the most complex, most overfit model.

The fix

Choose the degree with the smallest validation error — error on data the model did not train on.

Hold out data; pick the degree that generalizes best

Why: Validation error falls then RISES as overfitting kicks in (bias–variance, Lesson 32). The minimum of the validation curve — here the honest degree-1 line — is the right complexity, not the training-error minimum.

75. Which of these survive contact with Lesson 40: Linear Regression & the PyTorch…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
We keep one tiny dataset for the whole lesson so every number is checkable by hand. Four points — think of x as hours studied and y as a score:; Stack a ones column (for the intercept) next to x to get the design matrix X, so Xw produces all four predictions at once:; Setting ∇L = 0 and solving gives the normal equations — a square, solvable system whose unique answer is the least-squares w:
Breaks
Bigger steps must reach the minimum faster — crank η up to 0.5.; The gradient is fresh each pass, so just backward() then step() — skip zero_grad().
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 40: Linear Regression & the PyTorch Training Loop puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

76. Score the fit

Section

Part 5 of 6 — R², RMSE, sklearn

77. R² and RMSE

Concept

Two standard scores. R² is the fraction of the target's variance the model explains (1 = perfect, 0 = no better than the mean). RMSE is the typical error, in the target's own units:

\[ R^2 = 1 - \frac{\sum_i (y_i - \hat y_i)^2}{\sum_i (y_i - \bar y)^2}, \qquad \mathrm{RMSE} = \sqrt{\tfrac{1}{n}\textstyle\sum_i (y_i - \hat y_i)^2} \]

We already have the pieces: SS_res = Σr² = 0.8 and MSE = 0.2. Only SS_tot (the mean-only baseline error) is new.

78. Teach it back: R² and RMSE

Explain it

Discussion prompt

Explain R² and RMSE to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Two standard scores. R² is the fraction of the target's variance the model explains (1 = perfect, 0 = no better than the mean). RMSE is the typical error, in the target's own units:

79. What has to be given first: R² and RMSE by hand, then in code

Missing information

Discussion prompt

ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

The mean-only baseline leaves total error 17.0. Our line leaves only 0.8, so R² = 1 − 0.8/17 = 0.9529: the line explains 95% of the variance in y.

80. R² and RMSE by hand, then in code

Worked example

ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
w = np.linalg.solve(X.T@X, X.T@y)
yhat = X @ w
sse = ((y - yhat)**2).sum()
r2 = 1 - sse / ((y - y.mean())**2).sum()
rmse = np.sqrt(sse / len(y))
print("R2   =", round(r2, 4))
print("RMSE =", round(rmse, 4))

SS_tot = 6.25 + 2.25 + 2.25 + 6.25 = 17.0

Why: The mean-only baseline leaves total error 17.0. Our line leaves only 0.8, so R² = 1 − 0.8/17 = 0.9529: the line explains 95% of the variance in y.

quantityvalue (verified)
SS_res = Σ(y − ŷ)²0.8
SS_tot = Σ(y − ȳ)²17.0
R² = 1 − SS_res/SS_tot0.9529
RMSE = √(SS_res/n)0.4472

81. Work backwards from the answer: R² and RMSE by hand, then in code

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

SS_tot = 6.25 + 2.25 + 2.25 + 6.25 = 17.0

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

ȳ = 22/4 = 5.5. Deviations of y = [3,4,7,8] from 5.5 are [−2.5, −1.5, 1.5, 2.5]; squared and summed give SS_tot:

82. Guess the shape of the answer: Cross-check against sklearn

Estimation

Predict first

The final sanity check: fit sklearn.LinearRegression and confirm it prints the same intercept, slope, and R²:

Commit before you compute: what does Cross-check against sklearn come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: sklearn agrees to the digit: intercept 1.0, slope 1.8, R² 0.9529

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Three independent routes — by-hand OLS, the PyTorch loop, and sklearn — all land on the identical [1.0, 1.8] with R² 0.9529.

83. Cross-check against sklearn

Worked example

The final sanity check: fit sklearn.LinearRegression and confirm it prints the same intercept, slope, and R²:

import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
m = LinearRegression().fit(x.reshape(-1, 1), y)
print("intercept =", round(m.intercept_, 4))
print("slope     =", round(m.coef_[0], 4))
print("R2        =", round(m.score(x.reshape(-1, 1), y), 4))

sklearn agrees to the digit: intercept 1.0, slope 1.8, R² 0.9529

Why: Three independent routes — by-hand OLS, the PyTorch loop, and sklearn — all land on the identical [1.0, 1.8] with R² 0.9529. When your from-scratch result matches the library, you built it right.

sourceinterceptslopeR²
by-hand OLS1.01.80.9529
PyTorch SGD1.01.80.9529
sklearn1.01.80.9529

84. What stays fixed: Cross-check against sklearn

Invariant

Step through it

Step through Cross-check against sklearn one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: source is by-hand OLS
  2. Step 2: source is PyTorch SGD
  3. Step 3: source is sklearn

85. Don't trust R² alone — check residuals

Concept

A high R² can still hide curvature or unequal spread. After fitting, plot residuals r = y − ŷ against the fitted ŷ: it should look like structureless scatter around zero.

A visible pattern (a U-shape, a fan) means the linear model is misspecified — add features or change the model. R² won't warn you; the residual plot will.

86. By analogy: Don't trust R² alone — check residuals

Analogy

Discussion prompt

Explain Don't trust R² alone — check residuals by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A high R² can still hide curvature or unequal spread. After fitting, plot residuals r = y − ŷ against the fitted ŷ: it should look like structureless scatter around zero.

87. Rebuild the recipe: The regression + training-loop recipe

Ranking

Put in order

These are the steps of The regression + training-loop recipe, scrambled. Put them back in order before the next slide shows you.

  1. Shape the data: design matrix X = [1, features]; tensors reshaped to (n, features) for nn.Linear
  2. Model + loss: nn.Linear (add polynomial columns for curves); MSELoss
  3. Loop, in order: zero_grad → forward → loss → backward → step, with a small learning rate
  4. Regularize: weight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation error
  5. Score & check: R², RMSE, a residual plot, and match a from-scratch / sklearn baseline

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

88. The regression + training-loop recipe

Pattern

  1. Shape the data: design matrix X = [1, features]; tensors reshaped to (n, features) for nn.Linear
  2. Model + loss: nn.Linear (add polynomial columns for curves); MSELoss
  3. Loop, in order: zero_grad → forward → loss → backward → step, with a small learning rate
  4. Regularize: weight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation error
  5. Score & check: R², RMSE, a residual plot, and match a from-scratch / sklearn baseline

89. Where does it stop working: The regression + training-loop recipe

Edge cases

Discussion prompt

The regression + training-loop recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Shape the data: design matrix X = [1, features]; tensors reshaped to (n, features) for nn.Linear
  2. Model + loss: nn.Linear (add polynomial columns for curves); MSELoss
  3. Loop, in order: zero_grad → forward → loss → backward → step, with a small learning rate
  4. Regularize: weight_decay=λ (L2 / ridge) to shrink weights; pick complexity by validation error
  5. Score & check: R², RMSE, a residual plot, and match a from-scratch / sklearn baseline

90. Rule out three: Check yourself — the loop order

Elimination

Eliminate the wrong options

What is the correct order of one PyTorch training iteration?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. zero_grad → forward → loss → backward → step
  • B. forward → backward → step → zero_grad → loss
  • C. backward → forward → loss → step
  • D. forward → loss → step → backward

Survives elimination: A

Why: Clear old grads, run the forward pass to predict, score the loss, backprop to fill every .grad, then let the optimizer update the parameters. Skipping zero_grad accumulates gradients across steps (the earlier trap).

91. Check yourself — the loop order

Check

Recall the five-line rhythm. Think about which step needs the output of which.

Check your understanding

What is the correct order of one PyTorch training iteration?

  • A. zero_grad → forward → loss → backward → step (correct)
  • B. forward → backward → step → zero_grad → loss
  • C. backward → forward → loss → step
  • D. forward → loss → step → backward

Answer: A

Why: Clear old grads, run the forward pass to predict, score the loss, backprop to fill every .grad, then let the optimizer update the parameters. Skipping zero_grad accumulates gradients across steps (the earlier trap).

Why B tempts people
backward() differentiates the loss, so it must come AFTER loss is computed — not before it. And zero_grad belongs at the START of the iteration, not after step.
Why C tempts people
You cannot call backward() before a forward pass has built the computation graph — there is nothing to differentiate yet.
Why D tempts people
step() applies the gradients, so it must come AFTER backward() computes them. Stepping before backward updates with stale (or zero) gradients.

92. Answer it before you see the options: Check yourself — closed form vs GD

Prediction

Predict first

Linear regression has an exact closed form. Why train it with gradient descent anyway?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: GD scales to huge data and generalizes to models with no closed form

Why: The closed form needs to solve a d×d system (O(d³), and only exists for linear least squares). GD scales to large d/n and — the whole point of learning it now — works for every later model (logistic regression, MLPs, transformers) where no closed form exists.

93. Check yourself — closed form vs GD

Check

We got [1.0, 1.8] two ways. Why bother with the slower, iterative one?

Check your understanding

Linear regression has an exact closed form. Why train it with gradient descent anyway?

  • A. GD scales to huge data and generalizes to models with no closed form (correct)
  • B. GD reaches a lower loss than the closed form
  • C. The closed form is wrong when the data has noise
  • D. GD does not need a loss function

Answer: A

Why: The closed form needs to solve a d×d system (O(d³), and only exists for linear least squares). GD scales to large d/n and — the whole point of learning it now — works for every later model (logistic regression, MLPs, transformers) where no closed form exists.

Why B tempts people
On convex MSE both reach the SAME global minimum — we verified GD lands on the identical [1.0, 1.8], loss 0.2. GD is not more accurate, just more general.
Why C tempts people
The closed form IS the exact least-squares fit for noisy data — that is precisely what OLS computes. Noise is not a reason to abandon it.
Why D tempts people
GD is defined by differentiating a loss — step 3 of the loop computes the loss and step 4 backprops it. There is no avoiding the loss function.

94. Check yourself — the MSE gradient

Check

You derived it term by term. Read the dimensions carefully.

Check your understanding

For L(w) = (1/n)‖Xw − y‖², what is ∇L(w)?

  • A. (2/n) Xᵀ(Xw − y) (correct)
  • B. (2/n) X(Xw − y)
  • C. (2/n) Xᵀ(y − Xw)
  • D. (1/n) Xᵀ(Xw − y)

Answer: A

Why: Chain rule: the outer 2(Xw − y) times the inner Xᵀ, all scaled by 1/n, gives (2/n)Xᵀ(Xw − y). At w = [1.0, 1.8] this evaluates to [0, 0], confirming the minimum.

Why B tempts people
X (not Xᵀ) has shape (n×d); X·(n-vector) is undefined — the shapes don't conform. You must transpose: Xᵀ maps the n-vector residual back to a d-vector gradient.
Why C tempts people
This is the NEGATIVE gradient (y − Xw instead of Xw − y). Descending with it — w ← w − η·(this) — would move UPHILL and increase the loss.
Why D tempts people
The factor is 2/n, not 1/n: differentiating the square (·)² brings down a factor of 2 that does not cancel. Dropping it just rescales the effective learning rate, but the stated derivative is wrong.

95. Rule out three: Check yourself — weight_decay

Elimination

Eliminate the wrong options

Setting weight_decay=λ in a PyTorch optimizer implements:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. L2 regularization (ridge) — a λ‖w‖² penalty that shrinks weights toward 0
  • B. L1 regularization (lasso) — drives weights to exact zeros
  • C. Dropout — randomly zeroing activations during training
  • D. Early stopping — halting when validation loss rises

Survives elimination: A

Why: weight_decay adds the L2 penalty λ‖w‖² to the loss — exactly ridge regression (Lesson 12). It pulls parameters toward zero; in our run it shrank the intercept from 1.00 to 0.76.

96. Check yourself — weight_decay

Check

It shrank our intercept from 1.00 to 0.76. What is it?

Check your understanding

Setting weight_decay=λ in a PyTorch optimizer implements:

  • A. L2 regularization (ridge) — a λ‖w‖² penalty that shrinks weights toward 0 (correct)
  • B. L1 regularization (lasso) — drives weights to exact zeros
  • C. Dropout — randomly zeroing activations during training
  • D. Early stopping — halting when validation loss rises

Answer: A

Why: weight_decay adds the L2 penalty λ‖w‖² to the loss — exactly ridge regression (Lesson 12). It pulls parameters toward zero; in our run it shrank the intercept from 1.00 to 0.76.

Why B tempts people
L1/lasso (which yields exact zeros for sparsity) is NOT the built-in weight_decay. To get L1 you add λΣ|w| to the loss manually — there is no optimizer flag for it.
Why C tempts people
Dropout is a separate layer (nn.Dropout) that randomly zeroes activations; it is not an optimizer argument and does not penalize weight size.
Why D tempts people
Early stopping is a training-control policy based on a validation signal, applied outside the optimizer — not a per-step penalty like weight_decay.

97. Your turn: train it

Section

Part 6 of 6 — the project

98. Project: OLS and the training loop

Concept

Fit ŷ = w₀ + w₁·x on our data x = [1,2,3,4], y = [3,4,7,8] three ways and prove they agree, then add weight_decay. You derived every piece — now assemble it yourself.

#requirementtool
1closed-form OLSnp.linalg.solve
2PyTorch training loopnn.Linear, MSELoss, SGD
3verify match; add weight_decaycompare slope/intercept

Build rules: type every line yourself, reshape inputs to (n, 1) for nn.Linear, seed with torch.manual_seed(0), and never forget zero_grad() at the top of the loop.

99. Break it if you can: Project: OLS and the training loop

Counterexample

Discussion prompt

Fit ŷ = w₀ + w₁·x on our data x = [1,2,3,4], y = [3,4,7,8] three ways and prove they agree, then add weight_decay. You derived every piece — now assemble it yourself.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: type every line yourself, reshape inputs to (n, 1) for nn.Linear, seed with torch.manual_seed(0), and never forget zero_grad() at the top of the loop.

100. Milestone 1 — closed-form OLS

Worked example

Your turn: build X with a ones column and solve the normal equations. Say the intercept and slope you expect (from the by-hand [1.0, 1.8]) before you print.

Hint: X = np.c_[np.ones(4), x], then np.linalg.solve(X.T@X, X.T@y). Do not use inv.

import numpy as np
x = np.array([1., 2., 3., 4.])
y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
print("closed-form w =", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())
paramvalue (verified)
intercept w₀1.0
slope w₁1.8

101. Milestone 2 — the training loop

Worked example

Your turn: fit with nn.Linear + SGD. Predict whether the loss reaches near 0.2 (the minimum MSE) before you run it.

Hint: reshape to (4, 1) with unsqueeze(1); loop zero_grad → loss → backward → step for 1000 epochs at lr=0.1.

import torch, numpy as np
torch.manual_seed(0)
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for _ in range(1000):
    opt.zero_grad(); torch.nn.MSELoss()(model(xt), yt).backward(); opt.step()
print("slope =", round(model.weight.item(), 4), " intercept =", round(model.bias.item(), 4))
paramvalue (verified)
slope1.8
intercept1.0

102. Milestone 3 — verify & regularize

Worked example

Your turn: confirm SGD matched the closed form, then add weight_decay=0.5 and predict which way the intercept moves before printing.

Hint: SGD(..., weight_decay=0.5). The penalty pulls weights toward 0, so the intercept should shrink below 1.0.

import torch, numpy as np
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
lossfn = torch.nn.MSELoss()
for wd in (0.0, 0.5):
    torch.manual_seed(0)
    m = torch.nn.Linear(1, 1)
    opt = torch.optim.SGD(m.parameters(), lr=0.1, weight_decay=wd)
    for _ in range(1000):
        opt.zero_grad(); lossfn(m(xt), yt).backward(); opt.step()
    print(f"wd={wd}: slope={m.weight.item():.4f} intercept={m.bias.item():.4f}")
settingslopeintercept (verified)
wd = 0.0 (≈ OLS)1.80001.0000
wd = 0.5 (ridge)1.81820.7636

103. What each one costs: Milestone 3 — verify & regularize

Trade off

Comparison matrix

From Milestone 3 — verify & regularize: every row here is a choice with a cost. Fill the slope column, then say which row you would actually pick and what you give up for it.

settingslopeintercept (verified)
wd = 0.0 (≈ OLS)1.80001.0000
wd = 0.5 (ridge)1.81820.7636

104. The full program

Concept

import numpy as np, torch
x = np.array([1., 2., 3., 4.]); y = np.array([3., 4., 7., 8.])
X = np.c_[np.ones(4), x]
print("OLS closed form:", np.linalg.solve(X.T@X, X.T@y).round(4).tolist())

torch.manual_seed(0)
xt = torch.tensor(x, dtype=torch.float32).unsqueeze(1)
yt = torch.tensor(y, dtype=torch.float32).unsqueeze(1)
model = torch.nn.Linear(1, 1)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
for _ in range(1000):
    opt.zero_grad(); torch.nn.MSELoss()(model(xt), yt).backward(); opt.step()
print("PyTorch SGD:   ", round(model.bias.item(), 4), round(model.weight.item(), 4))
printed linevalue (verified)
OLS closed form[1.0, 1.8] (intercept, slope)
PyTorch SGD (intercept, slope)1.0 1.8
matchTrue

If your closed form reads [1.0, 1.8] and the training loop lands on the same two numbers — you built the loop every model in Phase 2 reuses.

105. Fill in: value (verified) for The full program

Comparison

Comparison matrix

From The full program: refill the value (verified) column from what you know. The rest of the table is as it appeared.

printed linevalue (verified)
OLS closed form[1.0, 1.8] (intercept, slope)
PyTorch SGD (intercept, slope)1.0 1.8
matchTrue

106. Show it off

Concept

Slides closed, out loud: explain (1) why closed-form OLS and gradient descent reach the same w, (2) the five steps of the training loop and what backward() and step() each do, and (3) what weight_decay does to the weights.

Stretch: add a train/validation split and track both losses; fit [1, x, x², x³] and watch degree 3 interpolate the 4 points (SSE → 0) while a held-out point gets worse. Next up: logistic regression — the same loop, a new loss.

107. Connect it up: Lesson 40: Linear Regression & the PyTorch Training Loop

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — One dataset, the closed form · The same w, by gradient descent · The PyTorch training loop · Regularize & bend the line · Score the fit · Your turn: train it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

108. What you can do now

Recap

ideathe one thing to remember
closed form vs GDsame w [1.0, 1.8]; GD scales & generalizes
MSE gradient∇L = (2/n)Xᵀ(Xw − y)
training loopzero_grad FIRST, every step
learning ratetoo big diverges; if loss rises, halve η
weight_decay= L2 / ridge; shrinks weights toward 0
evaluationR²/RMSE + residual plot; match a baseline

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 40 (Week 14 — Linear/Logistic Regression + Training Loop) — Barron · USAAIO Round 2 Preparation, 2026
  2. PyTorch nn.Linear / MSELoss / optim.SGD
  3. PyTorch autograd & the training loop
  4. Every trace table, gradient, loss and weight produced by real execution — torch 2.7.1+cpu + numpy 2.2.6 + scikit-learn, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108