Lesson 41: Logistic Regression

USAAIO Lesson 41, from Week 14 of Phase 2. You derive the logistic-regression loss from scratch and prove the gradient dL/dw = X^T(p−y), then implement binary logistic regression two ways: by manual backpropagation in NumPy, and in PyTorch with nn.Linear feeding a sigmoid and BCELoss. It covers BCEWithLogitsLoss and why it is numerically preferable, softmax regression for the multi-class case with CrossEntropyLoss, the intuition behind the decision boundary and how L2 regularization sharpens or softens it, and a comparison with the linear SVM. The lesson runs to 31 slides.

Subject: Machine Learning · 64 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Logistic Regression: Derive, Implement, Extend

Title

USAAIO · Lesson 41 · Week 14 (Phase 2)

From BCE loss derivation to multi-class softmax — the same training loop as L40, a fundamentally different loss. Binary and K-class classifiers in NumPy and PyTorch.

2. By the end of this lesson you can

Objectives

  1. Derive the binary cross-entropy loss from the log-likelihood of a Bernoulli model
  2. Prove the gradient dL/dw = X^T(p − y)/n from first principles (all steps)
  3. Implement logistic regression in NumPy backprop and PyTorch (nn.Linear → sigmoid → BCELoss)
  4. Explain why BCEWithLogitsLoss is preferred and reproduce the instability with extreme logits
  5. Extend to softmax regression (K-class) using CrossEntropyLoss and interpret the decision boundary

3. What survived from Linear Regression & the PyTorch Training Loop?

Warm-up

Discussion prompt

Before we open Lesson 41: Logistic Regression: without looking back, what was the main idea of Linear Regression & the PyTorch Training Loop, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

closed-form OLS vs gradient descent, batch vs stochastic GD, the canonical PyTorch training loop (nn.Linear, MSELoss, SGD), polynomial features and overfitting, L2 regularization via weight_decay, and R²/RMSE evaluation. Build LinearRegressionPyTorch and verify it converges to the closed-form solution.

4. Binary cross-entropy from scratch

Section

Part 1 of 3

5. The probabilistic model

Concept

Logistic regression models the probability of class 1 as a sigmoid of a linear score. The key constraint: this must stay in (0, 1).

\[ \hat{p}(y=1 \mid x) = \sigma(w^\top x) = \frac{1}{1+e^{-w^\top x}} \]

The loss function is not MSE — squared error on probabilities creates a non-convex surface. The right choice is negative log-likelihood of the Bernoulli distribution.

6. Break it if you can: The probabilistic model

Counterexample

Discussion prompt

Logistic regression models the probability of class 1 as a sigmoid of a linear score. The key constraint: this must stay in (0, 1).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The loss function is not MSE — squared error on probabilities creates a non-convex surface. The right choice is negative log-likelihood of the Bernoulli distribution.

7. What has to happen first: Deriving BCE from Bernoulli log-likelihood

Ranking

Put in order

Put the moves of Deriving BCE from Bernoulli log-likelihood into the order they have to happen.

  1. Write the Bernoulli likelihood for one example (y ∈ {0,1})
  2. Take the negative log to convert product to sum
  3. Note that L is convex in w (unlike MSE on probabilities)

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p(y|x) = p^y · (1-p)^(1-y) unifies both cases in one expression.

8. Deriving BCE from Bernoulli log-likelihood

Worked example

Write the Bernoulli likelihood for one example (y ∈ {0,1})

Why: p(y|x) = p^y · (1-p)^(1-y) unifies both cases in one expression.

\[ \mathcal{L}(w) = \prod_{i=1}^{n} \hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \]

Take the negative log to convert product to sum

Why: log turns products into sums; negation converts maximizing likelihood to minimizing loss.

\[ L(w) = -\frac{1}{n}\sum_{i=1}^{n}\bigl[y_i \log \hat{p}_i + (1-y_i)\log(1-\hat{p}_i)\bigr] \]

Note that L is convex in w (unlike MSE on probabilities)

Why: The sigmoid composed with log is concave in the linear score; negation makes the overall objective convex — so gradient descent finds the global minimum.

9. Decode the notation: Deriving BCE from Bernoulli log-likelihood

Notation

Annotate

From Deriving BCE from Bernoulli log-likelihood — read this one piece at a time. What is each part doing?

On: \( \mathcal{L}(w) = \prod_{i=1}^{n} \hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \)

  • p(y|x) = p^y · (1-p)^(1-y) unifies both cases in one expression.
  • log turns products into sums; negation converts maximizing likelihood to minimizing loss.
  • The sigmoid composed with log is concave in the linear score; negation makes the overall objective convex — so gradient descent finds the global minimum.

10. What has to happen first: Proving dL/dw = X^T(p − y)/n

Ranking

Put in order

Put the moves of Proving dL/dw = X^T(p − y)/n into the order they have to happen.

  1. Differentiate L w.r.t. the linear score z = w^T x_i
  2. Chain rule again: dL/dw = (dL/dz) · (dz/dw) = (p_i - y_i) · x_i
  3. Finite-difference check confirms the formula exactly

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Chain rule: dL/dz = -[y/p - (1-y)/(1-p)] · dp/dz.

11. Proving dL/dw = X^T(p − y)/n

Worked example

Differentiate L w.r.t. the linear score z = w^T x_i

Why: Chain rule: dL/dz = -[y/p - (1-y)/(1-p)] · dp/dz. Use dp/dz = p(1-p) (sigmoid derivative).

\[ \frac{\partial L}{\partial z_i} = -\left[\frac{y_i}{\hat{p}_i} - \frac{1-y_i}{1-\hat{p}_i}\right]\hat{p}_i(1-\hat{p}_i) = \hat{p}_i - y_i \]

Chain rule again: dL/dw = (dL/dz) · (dz/dw) = (p_i - y_i) · x_i

Why: dz/dw = x_i; summing over all n samples and dividing by n gives the compact matrix form.

\[ \frac{\partial L}{\partial w} = \frac{1}{n}X^\top(\hat{p} - y) \quad\text{(verified numerically, max error } {<} 10^{-11}\text{)} \]

Finite-difference check confirms the formula exactly

Why: Numerically computed gradient (central differences) matches the analytic formula to < 10^-11 on a 10-sample test.

12. Decode the notation: Proving dL/dw = X^T(p − y)/n

Notation

Annotate

From Proving dL/dw = X^T(p − y)/n — read this one piece at a time. What is each part doing?

On: \( \frac{\partial L}{\partial w} = \frac{1}{n}X^\top(\hat{p} - y) \quad\text{(verified numerically, max error } {<} 10^{-11}\text{)} \)

  • Chain rule: dL/dz = -[y/p - (1-y)/(1-p)] · dp/dz. Use dp/dz = p(1-p) (sigmoid derivative).
  • dz/dw = x_i; summing over all n samples and dividing by n gives the compact matrix form.
  • Numerically computed gradient (central differences) matches the analytic formula to < 10^-11 on a 10-sample test.

13. NumPy vs PyTorch implementation

Section

Part 2 of 3

14. Synthetic binary dataset

Concept

Generate 200 points from two Gaussian clusters: class 1 centered at (1,1), class 0 at (-1,-1). Both with covariance [[1, 0.3],[0.3, 1]]. Seed 42.

classcenternσ
1 (positive)(1, 1)1001.0 (correlated)
0 (negative)(-1, -1)1001.0 (correlated)

The clusters overlap moderately — a linear decision boundary won't achieve 100% accuracy, which is realistic and tests regularization properly.

15. Fill in: center for Synthetic binary dataset

Comparison

Comparison matrix

From Synthetic binary dataset: refill the center column from what you know. The rest of the table is as it appeared.

classcenternσ
1 (positive)(1, 1)1001.0 (correlated)
0 (negative)(-1, -1)1001.0 (correlated)

16. Guess the shape of the answer: NumPy manual backprop

Estimation

Predict first

Implement the gradient X^T(p-y)/n directly. No autograd — you own every computation.

Commit before you compute: what does NumPy manual backprop come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: w = [-0.1206, 1.3404, 1.8292], acc = 0.8950 after 250 epochs

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The gradient formula X^T(p-y)/n is the complete backprop for logistic regression — no library needed.

17. NumPy manual backprop

Worked example

Implement the gradient X^T(p-y)/n directly. No autograd — you own every computation.

import numpy as np
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1], [[1,.3],[.3,1]], 100)
X0 = rng.multivariate_normal([-1,-1], [[1,.3],[.3,1]], 100)
X = np.vstack([X1, X0]); y = np.array([1]*100 + [0]*100, dtype=float)
Xb = np.c_[np.ones(200), X]     # add bias column

def sigmoid(z): return 1/(1+np.exp(-z))

w = np.zeros(3); lr = 0.5
for epoch in range(250):
    p = sigmoid(Xb @ w)
    grad = Xb.T @ (p - y) / 200
    w -= lr * grad

acc = np.mean((sigmoid(Xb @ w) > 0.5) == y)
print(w.round(4), f'acc={acc:.4f}')

w = [-0.1206, 1.3404, 1.8292], acc = 0.8950 after 250 epochs

Why: The gradient formula X^T(p-y)/n is the complete backprop for logistic regression — no library needed.

epochlossw_biasw1w2
00.482030.00000.25940.2554
500.23777-0.06821.29471.5333
1000.23602-0.10371.34681.7214
1500.23582-0.11541.34801.7875
249 (final)0.23579-0.12021.34181.8245

18. What each one costs: NumPy manual backprop

Trade off

Comparison matrix

From NumPy manual backprop: every row here is a choice with a cost. Fill the w1 column, then say which row you would actually pick and what you give up for it.

epochlossw_biasw1w2
00.482030.00000.25940.2554
500.23777-0.06821.29471.5333
1000.23602-0.10371.34681.7214
1500.23582-0.11541.34801.7875
249 (final)0.23579-0.12021.34181.8245

19. What has to be given first: PyTorch: nn.Linear → sigmoid → BCELoss

Missing information

Discussion prompt

Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Autograd computes the same gradient — they're identical algorithms. The 300-epoch PyTorch run agrees with 250-epoch NumPy because lr and data match.

20. PyTorch: nn.Linear → sigmoid → BCELoss

Worked example

Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.

import torch; import torch.nn as nn
Xt = torch.tensor(X, dtype=torch.float32)
yt = torch.tensor(y.reshape(-1,1), dtype=torch.float32)
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(2, 1), nn.Sigmoid())
opt = torch.optim.SGD(model.parameters(), lr=0.5)
bce = nn.BCELoss()
for epoch in range(300):
    opt.zero_grad()
    loss = bce(model(Xt), yt)
    loss.backward(); opt.step()
preds = (model(Xt) > 0.5).float()
print(f'acc={(preds==yt).float().mean():.4f}')

PyTorch BCELoss converges to same weights as NumPy: [1.3400, 1.8301], bias −0.1207, acc 0.8950

Why: Autograd computes the same gradient — they're identical algorithms. The 300-epoch PyTorch run agrees with 250-epoch NumPy because lr and data match.

sourcew1w2biasacc
NumPy manual1.34041.8292-0.12060.895
PyTorch BCELoss1.34001.8301-0.12070.895
match✓✓✓✓

21. Work backwards from the answer: PyTorch: nn.Linear → sigmoid → BCELoss

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

PyTorch BCELoss converges to same weights as NumPy: [1.3400, 1.8301], bias −0.1207, acc 0.8950

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.

22. Something is wrong here: sigmoid + BCELoss on extreme logits

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Float32 saturates: sigmoid(50) = 1.0 to machine precision.

Use nn.Linear (raw logits) with nn.BCEWithLogitsLoss() — no sigmoid in the model.

Why: Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.

23. Trap: sigmoid + BCELoss on extreme logits

Trap

The trap

Use nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.

At large logit magnitudes (e.g. 50), sigmoid(50) rounds to exactly 1.0 in float32

Why: Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.

The fix

Use nn.Linear (raw logits) with nn.BCEWithLogitsLoss() — no sigmoid in the model.

BCEWithLogitsLoss uses the log-sum-exp trick: log(1 + exp(−z)) computed stably

Why: For z = 50: log(1 + exp(-50)) ≈ 0 exactly — no saturation. The numerics stay finite across the full logit range. Same final value: 0.2310 for our extreme test, but no NaN risk.

24. Break it on purpose: sigmoid + BCELoss on extreme logits

Break the constraint

Discussion prompt

The rule this trap just fixed:

For z = 50: log(1 + exp(-50)) ≈ 0 exactly — no saturation. The numerics stay finite across the full logit range. Same final value: 0.2310 for our extreme test, but no NaN risk.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.

25. BCEWithLogitsLoss: the log-sum-exp trick

Concept

The fix is computing the BCE loss directly from logits using a numerically equivalent form.

\[ \text{BCEWithLogitsLoss}(z, y) = \max(z,0) - yz + \log(1+e^{-|z|}) \]

For z = 50: max(50,0) − 1·50 + log(1+e^{−50}) = 50 − 50 + 0 = 0, finite. For z = −50 (class 0): 0 − 0 + log(1+e^{−50}) ≈ 0, also finite. Never divide by 0.

26. By analogy: BCEWithLogitsLoss: the log-sum-exp trick

Analogy

Discussion prompt

Explain BCEWithLogitsLoss: the log-sum-exp trick by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The fix is computing the BCE loss directly from logits using a numerically equivalent form.

27. Softmax regression and multi-class

Section

Part 3 of 3

28. From binary to K-class

Concept

Replace sigmoid with softmax: outputs a probability vector that sums to 1 over K classes. Still one linear layer — now outputting K logits.

\[ \hat{p}_k = \frac{e^{z_k}}{\sum_{j=1}^K e^{z_j}}, \quad z = Wx + b, \quad W \in \mathbb{R}^{K \times d} \]

The loss is cross-entropy over all K classes. nn.CrossEntropyLoss applies the log-sum-exp stable softmax internally — you pass raw logits, not probabilities.

29. Teach it back: From binary to K-class

Explain it

Discussion prompt

Explain From binary to K-class to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Replace sigmoid with softmax: outputs a probability vector that sums to 1 over K classes. Still one linear layer — now outputting K logits.

30. Guess the shape of the answer: Softmax CE by hand

Estimation

Predict first

3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.

Commit before you compute: what does Softmax CE by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Class 0 has logit 3.0, far above the others, so it captures 87.6% of the probability mass.

31. Softmax CE by hand

Worked example

3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.

import numpy as np
logits = np.array([3.0, 1.0, -2.0])
# Numerically stable: subtract max before exp
exp_z = np.exp(logits - np.max(logits))
probs = exp_z / exp_z.sum()
ce = -np.log(probs[0])      # correct class = 0
print(probs.round(4))       # [0.8756 0.1185 0.0059]
print(f'CE={ce:.4f}')       # 0.1328

Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328

Why: Class 0 has logit 3.0, far above the others, so it captures 87.6% of the probability mass. CE = −log(0.8756) = 0.1328.

classlogitsoftmax probCE contribution
0 (correct)3.00.8756−log(0.8756) = 0.1328
11.00.1185ignored (y≠1)
2−2.00.0059ignored (y≠2)

32. Work backwards from the answer: Softmax CE by hand

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.

33. Guess the shape of the answer: Softmax regression on Iris (3 classes)

Estimation

Predict first

Iris: 150 samples, 4 features, 3 classes — the canonical multi-class test. One nn.Linear(4, 3) + CrossEntropyLoss.

Commit before you compute: what does Softmax regression on Iris (3 classes) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Final accuracy 0.9733 on Iris; final CE loss 0.1734 — softmax regression with one linear layer

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Iris is nearly linearly separable; a single linear layer (no hidden units) reaches 97.3% in 500 epochs — same loop as L40, new loss.

34. Softmax regression on Iris (3 classes)

Worked example

Iris: 150 samples, 4 features, 3 classes — the canonical multi-class test. One nn.Linear(4, 3) + CrossEntropyLoss.

import torch; import torch.nn as nn
from sklearn.datasets import load_iris
iris = load_iris()
Xi = torch.tensor(iris.data, dtype=torch.float32)
yi = torch.tensor(iris.target, dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4, 3)         # raw logits, no softmax here
opt = torch.optim.SGD(model.parameters(), lr=0.1)
ce = nn.CrossEntropyLoss()      # softmax applied internally
for epoch in range(500):
    opt.zero_grad()
    loss = ce(model(Xi), yi)
    loss.backward(); opt.step()
preds = model(Xi).argmax(dim=1)
print(f'acc={(preds==yi).float().mean():.4f}')

Final accuracy 0.9733 on Iris; final CE loss 0.1734 — softmax regression with one linear layer

Why: Iris is nearly linearly separable; a single linear layer (no hidden units) reaches 97.3% in 500 epochs — same loop as L40, new loss.

epochCE loss
00.9524
1000.4592
2000.2656
3000.2207
500 (final)0.1734

35. Watch it run: Softmax regression on Iris (3 classes)

Pattern

Step through it

Step through Softmax regression on Iris (3 classes) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 100
  3. Step 3: epoch is 200
  4. Step 4: epoch is 300
  5. Step 5: epoch is 500 (final)

36. Decision boundary and regularization

Concept

Logistic regression produces a linear decision boundary: w^T x + b = 0. In 2D this is a line; in d dimensions, a hyperplane.

L2 regularization (weight_decay) shrinks the weight vector. A smaller ||w|| means a wider, smoother margin around the boundary — the boundary doesn't move, but the transition from p≈0 to p≈1 is more gradual.

setting||w||boundary sharpness
no regularization2.2683sharp sigmoid transition
weight_decay = 1.00.4113wide, smoother transition

37. By analogy: Decision boundary and regularization

Analogy

Discussion prompt

Explain Decision boundary and regularization by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Logistic regression produces a linear decision boundary: w^T x + b = 0. In 2D this is a line; in d dimensions, a hyperplane.

38. Something is wrong here: applying softmax before CrossEntropyLoss

Anomaly

Predict first

A student writes this, and it looks reasonable:

Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable).

Pass raw logits to CrossEntropyLoss. Add nn.Softmax only at inference for human-readable probabilities.

Why: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable). Feeding pre-softmax probabilities into it computes softmax(softmax(z)) — double-squashing the logits, producing badly distorted gradients.

39. Trap: applying softmax before CrossEntropyLoss

Trap

The trap

Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.

model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)); loss = CrossEntropyLoss()(model(x), y)

Why: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable). Feeding pre-softmax probabilities into it computes softmax(softmax(z)) — double-squashing the logits, producing badly distorted gradients.

The fix

Pass raw logits to CrossEntropyLoss. Add nn.Softmax only at inference for human-readable probabilities.

model = nn.Linear(4,3); loss = CrossEntropyLoss()(model(x), y)

Why: CrossEntropyLoss owns the softmax. At prediction time, model(x).argmax() doesn't need probabilities anyway — the argmax of logits equals argmax of probabilities.

40. Which of these survive contact with Lesson 41: Logistic Regression?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Logistic regression models the probability of class 1 as a sigmoid of a linear score. The key constraint: this must stay in (0, 1).; Generate 200 points from two Gaussian clusters: class 1 centered at (1,1), class 0 at (-1,-1). Both with covariance [[1, 0.3],[0.3, 1]]. Seed 42.; The fix is computing the BCE loss directly from logits using a numerically equivalent form.
Breaks
Use nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.; Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 41: Logistic Regression puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

41. Logistic regression vs linear SVM

Concept

Both are linear classifiers: same form of decision boundary (w^T x + b = 0). The difference is the loss function.

logistic regressionlinear SVM
losscross-entropy (log-loss)hinge loss
outputcalibrated probabilitiesmargin score (not a probability)
regularizationL2 built in via weight_decayC controls margin width
behaviour far from boundarypenalizes confidently wrong predictionszero loss once inside margin
sklearn accuracy (our data)0.90000.8950

In practice the accuracies are similar on linearly separable data; logistic regression is preferred when you need calibrated probabilities (e.g., ranking, cost-sensitive decisions).

42. Fill in: logistic regression for Logistic regression vs linear SVM

Comparison

Comparison matrix

From Logistic regression vs linear SVM: refill the logistic regression column from what you know. The rest of the table is as it appeared.

logistic regressionlinear SVM
losscross-entropy (log-loss)hinge loss
outputcalibrated probabilitiesmargin score (not a probability)
regularizationL2 built in via weight_decayC controls margin width
behaviour far from boundarypenalizes confidently wrong predictionszero loss once inside margin
sklearn accuracy (our data)0.90000.8950

43. Rebuild the recipe: The logistic regression recipe

Ranking

Put in order

These are the steps of The logistic regression recipe, scrambled. Put them back in order before the next slide shows you.

  1. Loss: negative log-likelihood of Bernoulli = BCE. Gradient = X^T(p−y)/n (proved from scratch).
  2. Binary: nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logits
  3. Multi-class: nn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logits
  4. Loop: identical to L40 — zero_grad → forward → loss → backward → step
  5. Regularize: weight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundary
  6. Interpret: predict class via logits.argmax(); probabilities via sigmoid (binary) or softmax (multi)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

44. The logistic regression recipe

Pattern

  1. Loss: negative log-likelihood of Bernoulli = BCE. Gradient = X^T(p−y)/n (proved from scratch).
  2. Binary: nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logits
  3. Multi-class: nn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logits
  4. Loop: identical to L40 — zero_grad → forward → loss → backward → step
  5. Regularize: weight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundary
  6. Interpret: predict class via logits.argmax(); probabilities via sigmoid (binary) or softmax (multi)

45. Where does it stop working: The logistic regression recipe

Edge cases

Discussion prompt

The logistic regression recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Loss: negative log-likelihood of Bernoulli = BCE. Gradient = X^T(p−y)/n (proved from scratch).
  2. Binary: nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logits
  3. Multi-class: nn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logits
  4. Loop: identical to L40 — zero_grad → forward → loss → backward → step
  5. Regularize: weight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundary
  6. Interpret: predict class via logits.argmax(); probabilities via sigmoid (binary) or softmax (multi)

46. Rule out three: Check yourself — the gradient

Elimination

Eliminate the wrong options

The gradient of binary cross-entropy loss w.r.t. w is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. X^T(p − y) / n
  • B. X^T(y − p) / n
  • C. X^T(p − y) · p(1−p)
  • D. 2 · X^T(Xw − y) / n

Survives elimination: A

Why: The chain rule yields dL/dz_i = p_i − y_i (the sigmoid derivative cancels cleanly with the log-likelihood). Summing over samples in matrix form gives X^T(p−y)/n. Verified numerically against finite differences with max error < 10^{-11}.

47. Check yourself — the gradient

Check

Cover the derivation and reproduce the result.

Check your understanding

The gradient of binary cross-entropy loss w.r.t. w is:

  • A. X^T(p − y) / n (correct)
  • B. X^T(y − p) / n
  • C. X^T(p − y) · p(1−p)
  • D. 2 · X^T(Xw − y) / n

Answer: A

Why: The chain rule yields dL/dz_i = p_i − y_i (the sigmoid derivative cancels cleanly with the log-likelihood). Summing over samples in matrix form gives X^T(p−y)/n. Verified numerically against finite differences with max error < 10^{-11}.

Why B tempts people
y − p has the wrong sign; that would be the gradient of the negative of the loss (i.e., the log-likelihood ascent direction).
Why C tempts people
Including p(1-p) double-counts the sigmoid derivative — the chain rule already absorbed it in the dL/dz step.
Why D tempts people
That's the MSE gradient (L40), not cross-entropy. Logistic regression uses BCE, not squared error.

48. Answer it before you see the options: Check yourself — BCEWithLogitsLoss

Prediction

Predict first

Why is BCEWithLogitsLoss preferred over nn.Sigmoid() + nn.BCELoss()?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Float32 sigmoid saturates at large logit magnitudes, making log(1−sigmoid(z)) = log(0) = −inf; BCEWithLogitsLoss avoids this with the log-sum-exp trick

Why: sigmoid(50) = 1.0 in float32. BCELoss then computes log(0) = −inf for a class-0 sample, giving NaN gradients. BCEWithLogitsLoss reformulates as max(z,0) − yz + log(1+exp(−|z|)), which is finite for any z. Both compute the same value for moderate logits (verified: 0.2310 on our test).

49. Check yourself — BCEWithLogitsLoss

Check

Numerical stability matters at competition time.

Check your understanding

Why is BCEWithLogitsLoss preferred over nn.Sigmoid() + nn.BCELoss()?

  • A. Float32 sigmoid saturates at large logit magnitudes, making log(1−sigmoid(z)) = log(0) = −inf; BCEWithLogitsLoss avoids this with the log-sum-exp trick (correct)
  • B. BCELoss produces a non-convex loss surface, BCEWithLogitsLoss is convex
  • C. BCEWithLogitsLoss trains faster because it skips the sigmoid computation
  • D. BCELoss requires class labels in {−1, +1} while BCEWithLogitsLoss accepts {0, 1}

Answer: A

Why: sigmoid(50) = 1.0 in float32. BCELoss then computes log(0) = −inf for a class-0 sample, giving NaN gradients. BCEWithLogitsLoss reformulates as max(z,0) − yz + log(1+exp(−|z|)), which is finite for any z. Both compute the same value for moderate logits (verified: 0.2310 on our test).

Why B tempts people
BCE is convex in z = w^Tx regardless of whether you split the sigmoid — both formulations are convex.
Why C tempts people
BCEWithLogitsLoss still computes the sigmoid internally — the speed difference is negligible; the motivation is numerical precision.
Why D tempts people
Both BCELoss and BCEWithLogitsLoss expect labels in {0, 1} for standard use; the label convention is the same.

50. Rule out three: Check yourself — CrossEntropyLoss usage

Elimination

Eliminate the wrong options

You build a 3-class classifier: model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)). You train with nn.CrossEntropyLoss()(model(x), y). What is wrong?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. CrossEntropyLoss applies softmax internally — passing pre-softmax output squashes the logits twice, distorting gradients
  • B. Nothing — Softmax normalizes to probabilities and CrossEntropyLoss then computes cross-entropy correctly
  • C. Softmax should be replaced by Sigmoid for multi-class
  • D. The output dimension should be 1, not 3, for 3 classes

Survives elimination: A

Why: nn.CrossEntropyLoss = log-softmax + NLLLoss. If you pass probabilities (already softmax'd) it computes softmax(softmax(z)), which compresses all activations toward uniform — training appears to work but converges much more slowly and to a suboptimal solution. Always pass raw logits.

51. Check yourself — CrossEntropyLoss usage

Check

A USAAIO classic trap question.

Check your understanding

You build a 3-class classifier: model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)). You train with nn.CrossEntropyLoss()(model(x), y). What is wrong?

  • A. CrossEntropyLoss applies softmax internally — passing pre-softmax output squashes the logits twice, distorting gradients (correct)
  • B. Nothing — Softmax normalizes to probabilities and CrossEntropyLoss then computes cross-entropy correctly
  • C. Softmax should be replaced by Sigmoid for multi-class
  • D. The output dimension should be 1, not 3, for 3 classes

Answer: A

Why: nn.CrossEntropyLoss = log-softmax + NLLLoss. If you pass probabilities (already softmax'd) it computes softmax(softmax(z)), which compresses all activations toward uniform — training appears to work but converges much more slowly and to a suboptimal solution. Always pass raw logits.

Why B tempts people
This is the misconception — CrossEntropyLoss does not 'receive probabilities and compute cross-entropy'; it expects logits and applies log-softmax itself.
Why C tempts people
Sigmoid is for binary output (each unit independent). For mutually exclusive K-class output, softmax is correct — but it belongs inside the loss, not the model.
Why D tempts people
The output size must match K (number of classes). For 3 classes, nn.Linear(4, 3) is correct.

52. Your turn: implement it

Section

Project

53. Project: logistic regression two ways

Concept

Implement binary logistic regression from scratch (NumPy) and in PyTorch, then extend to multi-class. Three milestones.

#taskkey call
1NumPy manual backpropX^T(p-y)/n gradient update
2PyTorch BCEWithLogitsLossnn.Linear(2,1) + BCEWithLogitsLoss
3Softmax regression on Irisnn.Linear(4,3) + CrossEntropyLoss, target acc ≥ 0.97

Build rule: no sigmoid in the model for milestones 2-3. Pass raw logits to the loss. Add weight_decay=1.0 to milestone 2 and confirm ||w|| shrinks.

54. Break it if you can: Project: logistic regression two ways

Counterexample

Discussion prompt

Implement binary logistic regression from scratch (NumPy) and in PyTorch, then extend to multi-class. Three milestones.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rule: no sigmoid in the model for milestones 2-3. Pass raw logits to the loss. Add weight_decay=1.0 to milestone 2 and confirm ||w|| shrinks.

55. Milestone 1 — NumPy manual backprop

Worked example

Your turn: implement the gradient update w -= lr * X^T(p-y)/n for 250 epochs on the synthetic dataset. Predict the final loss.

Hint: p = 1/(1+np.exp(-Xb @ w)), then grad = Xb.T @ (p - y) / n. Build Xb = np.c_[np.ones(n), X] first.

import numpy as np
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1], [[1,.3],[.3,1]], 100)
X0 = rng.multivariate_normal([-1,-1],[[1,.3],[.3,1]], 100)
X = np.vstack([X1, X0]); y = np.array([1]*100+[0]*100, dtype=float)
Xb = np.c_[np.ones(200), X]
def sigmoid(z): return 1/(1+np.exp(-z))
w = np.zeros(3); lr = 0.5
for _ in range(250):
    p = sigmoid(Xb @ w)
    w -= lr * Xb.T @ (p-y) / 200
print(w.round(4), np.mean((sigmoid(Xb@w)>0.5)==y))
paramvalue
w_bias-0.1202
w11.3418
w21.8245
accuracy0.895

56. Watch it run: Milestone 1 — NumPy manual backprop

Pattern

Step through it

Step through Milestone 1 — NumPy manual backprop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: param is w_bias
  2. Step 2: param is w1
  3. Step 3: param is w2
  4. Step 4: param is accuracy

57. Milestone 2 — PyTorch BCEWithLogitsLoss

Worked example

Your turn: replace sigmoid+BCELoss with the numerically stable version. Add weight_decay=1.0 and measure ||w|| before and after.

Hint: nn.Linear(2,1) (no sigmoid in model) + nn.BCEWithLogitsLoss(). Target in yt shape (200,1).

import torch; import torch.nn as nn
Xt = torch.tensor(X, dtype=torch.float32)
yt = torch.tensor(y.reshape(-1,1), dtype=torch.float32)
torch.manual_seed(0)
m = nn.Linear(2,1)
opt = torch.optim.SGD(m.parameters(), lr=0.5, weight_decay=1.0)
loss_fn = nn.BCEWithLogitsLoss()
for _ in range(300):
    opt.zero_grad()
    loss_fn(m(Xt), yt).backward(); opt.step()
w = m.weight.detach().numpy().flatten()
print(w.round(4), f'||w||={np.linalg.norm(w):.4f}')
setting||w||
no weight_decay2.2683
weight_decay = 1.00.4113

58. Milestone 3 — Iris softmax regression

Worked example

Your turn: use nn.Linear(4,3) + CrossEntropyLoss for 500 epochs. Predict the accuracy and confirm it exceeds 0.97.

Hint: target yi must be torch.long (integer class indices). loss = ce(model(Xi), yi). Predict with .argmax(dim=1).

import torch; import torch.nn as nn
from sklearn.datasets import load_iris
iris = load_iris()
Xi = torch.tensor(iris.data, dtype=torch.float32)
yi = torch.tensor(iris.target, dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4, 3)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
ce = nn.CrossEntropyLoss()
for _ in range(500):
    opt.zero_grad(); ce(model(Xi), yi).backward(); opt.step()
preds = model(Xi).argmax(dim=1)
print(f'acc={(preds==yi).float().mean():.4f}')
metricvalue
final CE loss0.1734
accuracy0.9733
target≥ 0.97 ✓

59. What each one costs: Milestone 3 — Iris softmax regression

Trade off

Comparison matrix

From Milestone 3 — Iris softmax regression: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

metricvalue
final CE loss0.1734
accuracy0.9733
target≥ 0.97 ✓

60. The full program

Concept

import numpy as np, torch, torch.nn as nn
from sklearn.datasets import load_iris

# --- Binary (NumPy) ---
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1],[[1,.3],[.3,1]],100)
X0 = rng.multivariate_normal([-1,-1],[[1,.3],[.3,1]],100)
X = np.vstack([X1,X0]); y = np.array([1]*100+[0]*100,dtype=float)
Xb = np.c_[np.ones(200),X]
sigmoid = lambda z: 1/(1+np.exp(-z))
w = np.zeros(3)
for _ in range(250):
    p = sigmoid(Xb@w); w -= 0.5*Xb.T@(p-y)/200
print('NumPy acc:', np.mean((sigmoid(Xb@w)>0.5)==y))

# --- Multi-class (PyTorch) ---
iris = load_iris()
Xi = torch.tensor(iris.data,dtype=torch.float32)
yi = torch.tensor(iris.target,dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4,3); opt=torch.optim.SGD(model.parameters(),lr=0.1)
ce = nn.CrossEntropyLoss()
for _ in range(500):
    opt.zero_grad(); ce(model(Xi),yi).backward(); opt.step()
print('Iris acc:', (model(Xi).argmax(1)==yi).float().mean().item())
taskresult
NumPy binary acc0.895
Iris multi-class acc0.9733

The training loop is identical to L40 — only the model output dim and the loss function change. That's the architecture of Phase 2: plug in a new head and loss, reuse everything else.

61. Fill in: result for The full program

Comparison

Comparison matrix

From The full program: refill the result column from what you know. The rest of the table is as it appeared.

taskresult
NumPy binary acc0.895
Iris multi-class acc0.9733

62. Show it off

Concept

Out loud, slides closed: (1) derive dL/dw = X^T(p-y)/n in 3 lines — sigmoid derivative, chain rule, sum; (2) explain why BCEWithLogitsLoss is preferred; (3) say what CrossEntropyLoss does internally and why you never pre-softmax the model output.

Stretch (homework): (1) prove the gradient fully in writing — all steps, no shortcuts; (2) implement both NumPy and PyTorch versions, compare weights; (3) show sigmoid+BCELoss fails on logits ± 100; (4) reach >90% on MNIST 10-class softmax regression in PyTorch.

63. Connect it up: Lesson 41: Logistic Regression

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Binary cross-entropy from scratch · NumPy vs PyTorch implementation · Softmax regression and multi-class · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

64. What you can do now

Recap

ideathe one thing to remember
gradient formulaX^T(p−y)/n — sigmoid derivative cancels
BCEWithLogitsLosslog-sum-exp trick — finite for any logit
CrossEntropyLossowns the softmax — never pre-apply it
regularizationweight_decay shrinks ||w||, widens margin
vs linear SVMsame boundary, different loss; logistic gives calibrated P

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 41 (Week 14 — Logistic Regression) — Barron · USAAIO Round 2 Preparation, 2026
  2. Gradient formula X^T(p-y)/n verified against finite differences, BCEWithLogitsLoss vs sigmoid+BCELoss stability, softmax CE verified by hand and with PyTorch, iris softmax accuracy 0.9733 with CrossEntropyLoss — torch 2.7.1+cpu, numpy 2.2.6, scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108