USAAIO Lesson 41, from Week 14 of Phase 2. You derive the logistic-regression loss from scratch and prove the gradient dL/dw = X^T(p−y), then implement binary logistic regression two ways: by manual backpropagation in NumPy, and in PyTorch with nn.Linear feeding a sigmoid and BCELoss. It covers BCEWithLogitsLoss and why it is numerically preferable, softmax regression for the multi-class case with CrossEntropyLoss, the intuition behind the decision boundary and how L2 regularization sharpens or softens it, and a comparison with the linear SVM. The lesson runs to 31 slides.
Subject: Machine Learning · 64 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 41 · Week 14 (Phase 2)
From BCE loss derivation to multi-class softmax — the same training loop as L40, a fundamentally different loss. Binary and K-class classifiers in NumPy and PyTorch.
Objectives
nn.Linear → sigmoid → BCELoss)BCEWithLogitsLoss is preferred and reproduce the instability with extreme logitsCrossEntropyLoss and interpret the decision boundaryWarm-up
Discussion prompt
Before we open Lesson 41: Logistic Regression: without looking back, what was the main idea of Linear Regression & the PyTorch Training Loop, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
closed-form OLS vs gradient descent, batch vs stochastic GD, the canonical PyTorch training loop (nn.Linear, MSELoss, SGD), polynomial features and overfitting, L2 regularization via weight_decay, and R²/RMSE evaluation. Build LinearRegressionPyTorch and verify it converges to the closed-form solution.
Section
Part 1 of 3
Concept
Logistic regression models the probability of class 1 as a sigmoid of a linear score. The key constraint: this must stay in (0, 1).
\[ \hat{p}(y=1 \mid x) = \sigma(w^\top x) = \frac{1}{1+e^{-w^\top x}} \]
The loss function is not MSE — squared error on probabilities creates a non-convex surface. The right choice is negative log-likelihood of the Bernoulli distribution.
Counterexample
Discussion prompt
Logistic regression models the probability of class 1 as a sigmoid of a linear score. The key constraint: this must stay in (0, 1).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The loss function is not MSE — squared error on probabilities creates a non-convex surface. The right choice is negative log-likelihood of the Bernoulli distribution.
Ranking
Put in order
Put the moves of Deriving BCE from Bernoulli log-likelihood into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. p(y|x) = p^y · (1-p)^(1-y) unifies both cases in one expression.
Worked example
Write the Bernoulli likelihood for one example (y ∈ {0,1})
Why: p(y|x) = p^y · (1-p)^(1-y) unifies both cases in one expression.
\[ \mathcal{L}(w) = \prod_{i=1}^{n} \hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \]
Take the negative log to convert product to sum
Why: log turns products into sums; negation converts maximizing likelihood to minimizing loss.
\[ L(w) = -\frac{1}{n}\sum_{i=1}^{n}\bigl[y_i \log \hat{p}_i + (1-y_i)\log(1-\hat{p}_i)\bigr] \]
Note that L is convex in w (unlike MSE on probabilities)
Why: The sigmoid composed with log is concave in the linear score; negation makes the overall objective convex — so gradient descent finds the global minimum.
Notation
Annotate
From Deriving BCE from Bernoulli log-likelihood — read this one piece at a time. What is each part doing?
On: \( \mathcal{L}(w) = \prod_{i=1}^{n} \hat{p}_i^{y_i}(1-\hat{p}_i)^{1-y_i} \)
Ranking
Put in order
Put the moves of Proving dL/dw = X^T(p − y)/n into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Chain rule: dL/dz = -[y/p - (1-y)/(1-p)] · dp/dz.
Worked example
Differentiate L w.r.t. the linear score z = w^T x_i
Why: Chain rule: dL/dz = -[y/p - (1-y)/(1-p)] · dp/dz. Use dp/dz = p(1-p) (sigmoid derivative).
\[ \frac{\partial L}{\partial z_i} = -\left[\frac{y_i}{\hat{p}_i} - \frac{1-y_i}{1-\hat{p}_i}\right]\hat{p}_i(1-\hat{p}_i) = \hat{p}_i - y_i \]
Chain rule again: dL/dw = (dL/dz) · (dz/dw) = (p_i - y_i) · x_i
Why: dz/dw = x_i; summing over all n samples and dividing by n gives the compact matrix form.
\[ \frac{\partial L}{\partial w} = \frac{1}{n}X^\top(\hat{p} - y) \quad\text{(verified numerically, max error } {<} 10^{-11}\text{)} \]
Finite-difference check confirms the formula exactly
Why: Numerically computed gradient (central differences) matches the analytic formula to < 10^-11 on a 10-sample test.
Notation
Annotate
From Proving dL/dw = X^T(p − y)/n — read this one piece at a time. What is each part doing?
On: \( \frac{\partial L}{\partial w} = \frac{1}{n}X^\top(\hat{p} - y) \quad\text{(verified numerically, max error } {<} 10^{-11}\text{)} \)
Section
Part 2 of 3
Concept
Generate 200 points from two Gaussian clusters: class 1 centered at (1,1), class 0 at (-1,-1). Both with covariance [[1, 0.3],[0.3, 1]]. Seed 42.
| class | center | n | σ |
|---|---|---|---|
| 1 (positive) | (1, 1) | 100 | 1.0 (correlated) |
| 0 (negative) | (-1, -1) | 100 | 1.0 (correlated) |
The clusters overlap moderately — a linear decision boundary won't achieve 100% accuracy, which is realistic and tests regularization properly.
Comparison
Comparison matrix
From Synthetic binary dataset: refill the center column from what you know. The rest of the table is as it appeared.
| class | center | n | σ |
|---|---|---|---|
| 1 (positive) | (1, 1) | 100 | 1.0 (correlated) |
| 0 (negative) | (-1, -1) | 100 | 1.0 (correlated) |
Estimation
Predict first
Implement the gradient X^T(p-y)/n directly. No autograd — you own every computation.
Commit before you compute: what does NumPy manual backprop come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: w = [-0.1206, 1.3404, 1.8292], acc = 0.8950 after 250 epochs
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The gradient formula X^T(p-y)/n is the complete backprop for logistic regression — no library needed.
Worked example
Implement the gradient X^T(p-y)/n directly. No autograd — you own every computation.
import numpy as np
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1], [[1,.3],[.3,1]], 100)
X0 = rng.multivariate_normal([-1,-1], [[1,.3],[.3,1]], 100)
X = np.vstack([X1, X0]); y = np.array([1]*100 + [0]*100, dtype=float)
Xb = np.c_[np.ones(200), X] # add bias column
def sigmoid(z): return 1/(1+np.exp(-z))
w = np.zeros(3); lr = 0.5
for epoch in range(250):
p = sigmoid(Xb @ w)
grad = Xb.T @ (p - y) / 200
w -= lr * grad
acc = np.mean((sigmoid(Xb @ w) > 0.5) == y)
print(w.round(4), f'acc={acc:.4f}')w = [-0.1206, 1.3404, 1.8292], acc = 0.8950 after 250 epochs
Why: The gradient formula X^T(p-y)/n is the complete backprop for logistic regression — no library needed.
| epoch | loss | w_bias | w1 | w2 |
|---|---|---|---|---|
| 0 | 0.48203 | 0.0000 | 0.2594 | 0.2554 |
| 50 | 0.23777 | -0.0682 | 1.2947 | 1.5333 |
| 100 | 0.23602 | -0.1037 | 1.3468 | 1.7214 |
| 150 | 0.23582 | -0.1154 | 1.3480 | 1.7875 |
| 249 (final) | 0.23579 | -0.1202 | 1.3418 | 1.8245 |
Trade off
Comparison matrix
From NumPy manual backprop: every row here is a choice with a cost. Fill the w1 column, then say which row you would actually pick and what you give up for it.
| epoch | loss | w_bias | w1 | w2 |
|---|---|---|---|---|
| 0 | 0.48203 | 0.0000 | 0.2594 | 0.2554 |
| 50 | 0.23777 | -0.0682 | 1.2947 | 1.5333 |
| 100 | 0.23602 | -0.1037 | 1.3468 | 1.7214 |
| 150 | 0.23582 | -0.1154 | 1.3480 | 1.7875 |
| 249 (final) | 0.23579 | -0.1202 | 1.3418 | 1.8245 |
Missing information
Discussion prompt
Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Autograd computes the same gradient — they're identical algorithms. The 300-epoch PyTorch run agrees with 250-epoch NumPy because lr and data match.
Worked example
Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.
import torch; import torch.nn as nn
Xt = torch.tensor(X, dtype=torch.float32)
yt = torch.tensor(y.reshape(-1,1), dtype=torch.float32)
torch.manual_seed(0)
model = nn.Sequential(nn.Linear(2, 1), nn.Sigmoid())
opt = torch.optim.SGD(model.parameters(), lr=0.5)
bce = nn.BCELoss()
for epoch in range(300):
opt.zero_grad()
loss = bce(model(Xt), yt)
loss.backward(); opt.step()
preds = (model(Xt) > 0.5).float()
print(f'acc={(preds==yt).float().mean():.4f}')PyTorch BCELoss converges to same weights as NumPy: [1.3400, 1.8301], bias −0.1207, acc 0.8950
Why: Autograd computes the same gradient — they're identical algorithms. The 300-epoch PyTorch run agrees with 250-epoch NumPy because lr and data match.
| source | w1 | w2 | bias | acc |
|---|---|---|---|---|
| NumPy manual | 1.3404 | 1.8292 | -0.1206 | 0.895 |
| PyTorch BCELoss | 1.3400 | 1.8301 | -0.1207 | 0.895 |
| match | ✓ | ✓ | ✓ | ✓ |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
PyTorch BCELoss converges to same weights as NumPy: [1.3400, 1.8301], bias −0.1207, acc 0.8950
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same dataset, same training loop rhythm from L40: zero_grad → forward → loss → backward → step.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Float32 saturates: sigmoid(50) = 1.0 to machine precision.
Use nn.Linear (raw logits) with nn.BCEWithLogitsLoss() — no sigmoid in the model.
Why: Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.
Trap
Use nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.
At large logit magnitudes (e.g. 50), sigmoid(50) rounds to exactly 1.0 in float32
Why: Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.
Use nn.Linear (raw logits) with nn.BCEWithLogitsLoss() — no sigmoid in the model.
BCEWithLogitsLoss uses the log-sum-exp trick: log(1 + exp(−z)) computed stably
Why: For z = 50: log(1 + exp(-50)) ≈ 0 exactly — no saturation. The numerics stay finite across the full logit range. Same final value: 0.2310 for our extreme test, but no NaN risk.
Break the constraint
Discussion prompt
The rule this trap just fixed:
For z = 50: log(1 + exp(-50)) ≈ 0 exactly — no saturation. The numerics stay finite across the full logit range. Same final value: 0.2310 for our extreme test, but no NaN risk.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Float32 saturates: sigmoid(50) = 1.0 to machine precision. Then BCELoss computes log(1 - 1.0) = log(0.0) = −inf, which propagates to NaN gradients and a broken model.
Concept
The fix is computing the BCE loss directly from logits using a numerically equivalent form.
\[ \text{BCEWithLogitsLoss}(z, y) = \max(z,0) - yz + \log(1+e^{-|z|}) \]
For z = 50: max(50,0) − 1·50 + log(1+e^{−50}) = 50 − 50 + 0 = 0, finite. For z = −50 (class 0): 0 − 0 + log(1+e^{−50}) ≈ 0, also finite. Never divide by 0.
Analogy
Discussion prompt
Explain BCEWithLogitsLoss: the log-sum-exp trick by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The fix is computing the BCE loss directly from logits using a numerically equivalent form.
Section
Part 3 of 3
Concept
Replace sigmoid with softmax: outputs a probability vector that sums to 1 over K classes. Still one linear layer — now outputting K logits.
\[ \hat{p}_k = \frac{e^{z_k}}{\sum_{j=1}^K e^{z_j}}, \quad z = Wx + b, \quad W \in \mathbb{R}^{K \times d} \]
The loss is cross-entropy over all K classes. nn.CrossEntropyLoss applies the log-sum-exp stable softmax internally — you pass raw logits, not probabilities.
Explain it
Discussion prompt
Explain From binary to K-class to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Replace sigmoid with softmax: outputs a probability vector that sums to 1 over K classes. Still one linear layer — now outputting K logits.
Estimation
Predict first
3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.
Commit before you compute: what does Softmax CE by hand come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Class 0 has logit 3.0, far above the others, so it captures 87.6% of the probability mass.
Worked example
3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.
import numpy as np
logits = np.array([3.0, 1.0, -2.0])
# Numerically stable: subtract max before exp
exp_z = np.exp(logits - np.max(logits))
probs = exp_z / exp_z.sum()
ce = -np.log(probs[0]) # correct class = 0
print(probs.round(4)) # [0.8756 0.1185 0.0059]
print(f'CE={ce:.4f}') # 0.1328Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328
Why: Class 0 has logit 3.0, far above the others, so it captures 87.6% of the probability mass. CE = −log(0.8756) = 0.1328.
| class | logit | softmax prob | CE contribution |
|---|---|---|---|
| 0 (correct) | 3.0 | 0.8756 | −log(0.8756) = 0.1328 |
| 1 | 1.0 | 0.1185 | ignored (y≠1) |
| 2 | −2.0 | 0.0059 | ignored (y≠2) |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Softmax probs: [0.8756, 0.1185, 0.0059]. CE loss = 0.1328
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
3-class example: logits = [3.0, 1.0, −2.0], true class = 0. Compute softmax probabilities and cross-entropy loss.
Estimation
Predict first
Iris: 150 samples, 4 features, 3 classes — the canonical multi-class test. One nn.Linear(4, 3) + CrossEntropyLoss.
Commit before you compute: what does Softmax regression on Iris (3 classes) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Final accuracy 0.9733 on Iris; final CE loss 0.1734 — softmax regression with one linear layer
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Iris is nearly linearly separable; a single linear layer (no hidden units) reaches 97.3% in 500 epochs — same loop as L40, new loss.
Worked example
Iris: 150 samples, 4 features, 3 classes — the canonical multi-class test. One nn.Linear(4, 3) + CrossEntropyLoss.
import torch; import torch.nn as nn
from sklearn.datasets import load_iris
iris = load_iris()
Xi = torch.tensor(iris.data, dtype=torch.float32)
yi = torch.tensor(iris.target, dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4, 3) # raw logits, no softmax here
opt = torch.optim.SGD(model.parameters(), lr=0.1)
ce = nn.CrossEntropyLoss() # softmax applied internally
for epoch in range(500):
opt.zero_grad()
loss = ce(model(Xi), yi)
loss.backward(); opt.step()
preds = model(Xi).argmax(dim=1)
print(f'acc={(preds==yi).float().mean():.4f}')Final accuracy 0.9733 on Iris; final CE loss 0.1734 — softmax regression with one linear layer
Why: Iris is nearly linearly separable; a single linear layer (no hidden units) reaches 97.3% in 500 epochs — same loop as L40, new loss.
| epoch | CE loss |
|---|---|
| 0 | 0.9524 |
| 100 | 0.4592 |
| 200 | 0.2656 |
| 300 | 0.2207 |
| 500 (final) | 0.1734 |
Pattern
Step through it
Step through Softmax regression on Iris (3 classes) one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Logistic regression produces a linear decision boundary: w^T x + b = 0. In 2D this is a line; in d dimensions, a hyperplane.
L2 regularization (weight_decay) shrinks the weight vector. A smaller ||w|| means a wider, smoother margin around the boundary — the boundary doesn't move, but the transition from p≈0 to p≈1 is more gradual.
| setting | ||w|| | boundary sharpness |
|---|---|---|
| no regularization | 2.2683 | sharp sigmoid transition |
| weight_decay = 1.0 | 0.4113 | wide, smoother transition |
Analogy
Discussion prompt
Explain Decision boundary and regularization by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Logistic regression produces a linear decision boundary: w^T x + b = 0. In 2D this is a line; in d dimensions, a hyperplane.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable).
Pass raw logits to CrossEntropyLoss. Add nn.Softmax only at inference for human-readable probabilities.
Why: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable). Feeding pre-softmax probabilities into it computes softmax(softmax(z)) — double-squashing the logits, producing badly distorted gradients.
Trap
Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.
model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)); loss = CrossEntropyLoss()(model(x), y)
Why: CrossEntropyLoss already applies log-softmax internally (log-sum-exp stable). Feeding pre-softmax probabilities into it computes softmax(softmax(z)) — double-squashing the logits, producing badly distorted gradients.
Pass raw logits to CrossEntropyLoss. Add nn.Softmax only at inference for human-readable probabilities.
model = nn.Linear(4,3); loss = CrossEntropyLoss()(model(x), y)
Why: CrossEntropyLoss owns the softmax. At prediction time, model(x).argmax() doesn't need probabilities anyway — the argmax of logits equals argmax of probabilities.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
nn.Sigmoid() in the model and pass output to nn.BCELoss() — this is cleaner because the model output is already a probability.; Since softmax outputs probabilities, apply it in the model and pass probabilities to CrossEntropyLoss — makes the model output interpretable.Concept
Both are linear classifiers: same form of decision boundary (w^T x + b = 0). The difference is the loss function.
| logistic regression | linear SVM | |
|---|---|---|
| loss | cross-entropy (log-loss) | hinge loss |
| output | calibrated probabilities | margin score (not a probability) |
| regularization | L2 built in via weight_decay | C controls margin width |
| behaviour far from boundary | penalizes confidently wrong predictions | zero loss once inside margin |
| sklearn accuracy (our data) | 0.9000 | 0.8950 |
In practice the accuracies are similar on linearly separable data; logistic regression is preferred when you need calibrated probabilities (e.g., ranking, cost-sensitive decisions).
Comparison
Comparison matrix
From Logistic regression vs linear SVM: refill the logistic regression column from what you know. The rest of the table is as it appeared.
| logistic regression | linear SVM | |
|---|---|---|
| loss | cross-entropy (log-loss) | hinge loss |
| output | calibrated probabilities | margin score (not a probability) |
| regularization | L2 built in via weight_decay | C controls margin width |
| behaviour far from boundary | penalizes confidently wrong predictions | zero loss once inside margin |
| sklearn accuracy (our data) | 0.9000 | 0.8950 |
Ranking
Put in order
These are the steps of The logistic regression recipe, scrambled. Put them back in order before the next slide shows you.
X^T(p−y)/n (proved from scratch).nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logitsnn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logitszero_grad → forward → loss → backward → stepweight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundarylogits.argmax(); probabilities via sigmoid (binary) or softmax (multi)Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
X^T(p−y)/n (proved from scratch).nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logitsnn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logitszero_grad → forward → loss → backward → stepweight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundarylogits.argmax(); probabilities via sigmoid (binary) or softmax (multi)Edge cases
Discussion prompt
The logistic regression recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
X^T(p−y)/n (proved from scratch).nn.Linear(d,1) + nn.BCEWithLogitsLoss() — never sigmoid → BCELoss on extreme logitsnn.Linear(d,K) + nn.CrossEntropyLoss() — softmax applied internally, pass raw logitszero_grad → forward → loss → backward → stepweight_decay=λ ↔ L2 penalty; shrinks ||w||, smooths the decision boundarylogits.argmax(); probabilities via sigmoid (binary) or softmax (multi)Elimination
Eliminate the wrong options
The gradient of binary cross-entropy loss w.r.t. w is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The chain rule yields dL/dz_i = p_i − y_i (the sigmoid derivative cancels cleanly with the log-likelihood). Summing over samples in matrix form gives X^T(p−y)/n. Verified numerically against finite differences with max error < 10^{-11}.
Check
Cover the derivation and reproduce the result.
Check your understanding
The gradient of binary cross-entropy loss w.r.t. w is:
Answer: A
Why: The chain rule yields dL/dz_i = p_i − y_i (the sigmoid derivative cancels cleanly with the log-likelihood). Summing over samples in matrix form gives X^T(p−y)/n. Verified numerically against finite differences with max error < 10^{-11}.
Prediction
Predict first
Why is BCEWithLogitsLoss preferred over nn.Sigmoid() + nn.BCELoss()?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Float32 sigmoid saturates at large logit magnitudes, making log(1−sigmoid(z)) = log(0) = −inf; BCEWithLogitsLoss avoids this with the log-sum-exp trick
Why: sigmoid(50) = 1.0 in float32. BCELoss then computes log(0) = −inf for a class-0 sample, giving NaN gradients. BCEWithLogitsLoss reformulates as max(z,0) − yz + log(1+exp(−|z|)), which is finite for any z. Both compute the same value for moderate logits (verified: 0.2310 on our test).
Check
Numerical stability matters at competition time.
Check your understanding
Why is BCEWithLogitsLoss preferred over nn.Sigmoid() + nn.BCELoss()?
Answer: A
Why: sigmoid(50) = 1.0 in float32. BCELoss then computes log(0) = −inf for a class-0 sample, giving NaN gradients. BCEWithLogitsLoss reformulates as max(z,0) − yz + log(1+exp(−|z|)), which is finite for any z. Both compute the same value for moderate logits (verified: 0.2310 on our test).
Elimination
Eliminate the wrong options
You build a 3-class classifier: model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)). You train with nn.CrossEntropyLoss()(model(x), y). What is wrong?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: nn.CrossEntropyLoss = log-softmax + NLLLoss. If you pass probabilities (already softmax'd) it computes softmax(softmax(z)), which compresses all activations toward uniform — training appears to work but converges much more slowly and to a suboptimal solution. Always pass raw logits.
Check
A USAAIO classic trap question.
Check your understanding
You build a 3-class classifier: model = nn.Sequential(nn.Linear(4,3), nn.Softmax(dim=1)). You train with nn.CrossEntropyLoss()(model(x), y). What is wrong?
Answer: A
Why: nn.CrossEntropyLoss = log-softmax + NLLLoss. If you pass probabilities (already softmax'd) it computes softmax(softmax(z)), which compresses all activations toward uniform — training appears to work but converges much more slowly and to a suboptimal solution. Always pass raw logits.
nn.Linear(4, 3) is correct.Section
Project
Concept
Implement binary logistic regression from scratch (NumPy) and in PyTorch, then extend to multi-class. Three milestones.
| # | task | key call |
|---|---|---|
| 1 | NumPy manual backprop | X^T(p-y)/n gradient update |
| 2 | PyTorch BCEWithLogitsLoss | nn.Linear(2,1) + BCEWithLogitsLoss |
| 3 | Softmax regression on Iris | nn.Linear(4,3) + CrossEntropyLoss, target acc ≥ 0.97 |
Build rule: no sigmoid in the model for milestones 2-3. Pass raw logits to the loss. Add weight_decay=1.0 to milestone 2 and confirm ||w|| shrinks.
Counterexample
Discussion prompt
Implement binary logistic regression from scratch (NumPy) and in PyTorch, then extend to multi-class. Three milestones.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rule: no sigmoid in the model for milestones 2-3. Pass raw logits to the loss. Add weight_decay=1.0 to milestone 2 and confirm ||w|| shrinks.
Worked example
Your turn: implement the gradient update w -= lr * X^T(p-y)/n for 250 epochs on the synthetic dataset. Predict the final loss.
Hint: p = 1/(1+np.exp(-Xb @ w)), then grad = Xb.T @ (p - y) / n. Build Xb = np.c_[np.ones(n), X] first.
import numpy as np
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1], [[1,.3],[.3,1]], 100)
X0 = rng.multivariate_normal([-1,-1],[[1,.3],[.3,1]], 100)
X = np.vstack([X1, X0]); y = np.array([1]*100+[0]*100, dtype=float)
Xb = np.c_[np.ones(200), X]
def sigmoid(z): return 1/(1+np.exp(-z))
w = np.zeros(3); lr = 0.5
for _ in range(250):
p = sigmoid(Xb @ w)
w -= lr * Xb.T @ (p-y) / 200
print(w.round(4), np.mean((sigmoid(Xb@w)>0.5)==y))| param | value |
|---|---|
| w_bias | -0.1202 |
| w1 | 1.3418 |
| w2 | 1.8245 |
| accuracy | 0.895 |
Pattern
Step through it
Step through Milestone 1 — NumPy manual backprop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: replace sigmoid+BCELoss with the numerically stable version. Add weight_decay=1.0 and measure ||w|| before and after.
Hint: nn.Linear(2,1) (no sigmoid in model) + nn.BCEWithLogitsLoss(). Target in yt shape (200,1).
import torch; import torch.nn as nn
Xt = torch.tensor(X, dtype=torch.float32)
yt = torch.tensor(y.reshape(-1,1), dtype=torch.float32)
torch.manual_seed(0)
m = nn.Linear(2,1)
opt = torch.optim.SGD(m.parameters(), lr=0.5, weight_decay=1.0)
loss_fn = nn.BCEWithLogitsLoss()
for _ in range(300):
opt.zero_grad()
loss_fn(m(Xt), yt).backward(); opt.step()
w = m.weight.detach().numpy().flatten()
print(w.round(4), f'||w||={np.linalg.norm(w):.4f}')| setting | ||w|| |
|---|---|
| no weight_decay | 2.2683 |
| weight_decay = 1.0 | 0.4113 |
Worked example
Your turn: use nn.Linear(4,3) + CrossEntropyLoss for 500 epochs. Predict the accuracy and confirm it exceeds 0.97.
Hint: target yi must be torch.long (integer class indices). loss = ce(model(Xi), yi). Predict with .argmax(dim=1).
import torch; import torch.nn as nn
from sklearn.datasets import load_iris
iris = load_iris()
Xi = torch.tensor(iris.data, dtype=torch.float32)
yi = torch.tensor(iris.target, dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4, 3)
opt = torch.optim.SGD(model.parameters(), lr=0.1)
ce = nn.CrossEntropyLoss()
for _ in range(500):
opt.zero_grad(); ce(model(Xi), yi).backward(); opt.step()
preds = model(Xi).argmax(dim=1)
print(f'acc={(preds==yi).float().mean():.4f}')| metric | value |
|---|---|
| final CE loss | 0.1734 |
| accuracy | 0.9733 |
| target | ≥ 0.97 ✓ |
Trade off
Comparison matrix
From Milestone 3 — Iris softmax regression: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| metric | value |
|---|---|
| final CE loss | 0.1734 |
| accuracy | 0.9733 |
| target | ≥ 0.97 ✓ |
Concept
import numpy as np, torch, torch.nn as nn
from sklearn.datasets import load_iris
# --- Binary (NumPy) ---
rng = np.random.default_rng(42)
X1 = rng.multivariate_normal([1,1],[[1,.3],[.3,1]],100)
X0 = rng.multivariate_normal([-1,-1],[[1,.3],[.3,1]],100)
X = np.vstack([X1,X0]); y = np.array([1]*100+[0]*100,dtype=float)
Xb = np.c_[np.ones(200),X]
sigmoid = lambda z: 1/(1+np.exp(-z))
w = np.zeros(3)
for _ in range(250):
p = sigmoid(Xb@w); w -= 0.5*Xb.T@(p-y)/200
print('NumPy acc:', np.mean((sigmoid(Xb@w)>0.5)==y))
# --- Multi-class (PyTorch) ---
iris = load_iris()
Xi = torch.tensor(iris.data,dtype=torch.float32)
yi = torch.tensor(iris.target,dtype=torch.long)
torch.manual_seed(0)
model = nn.Linear(4,3); opt=torch.optim.SGD(model.parameters(),lr=0.1)
ce = nn.CrossEntropyLoss()
for _ in range(500):
opt.zero_grad(); ce(model(Xi),yi).backward(); opt.step()
print('Iris acc:', (model(Xi).argmax(1)==yi).float().mean().item())| task | result |
|---|---|
| NumPy binary acc | 0.895 |
| Iris multi-class acc | 0.9733 |
The training loop is identical to L40 — only the model output dim and the loss function change. That's the architecture of Phase 2: plug in a new head and loss, reuse everything else.
Comparison
Comparison matrix
From The full program: refill the result column from what you know. The rest of the table is as it appeared.
| task | result |
|---|---|
| NumPy binary acc | 0.895 |
| Iris multi-class acc | 0.9733 |
Concept
Out loud, slides closed: (1) derive dL/dw = X^T(p-y)/n in 3 lines — sigmoid derivative, chain rule, sum; (2) explain why BCEWithLogitsLoss is preferred; (3) say what CrossEntropyLoss does internally and why you never pre-softmax the model output.
Stretch (homework): (1) prove the gradient fully in writing — all steps, no shortcuts; (2) implement both NumPy and PyTorch versions, compare weights; (3) show sigmoid+BCELoss fails on logits ± 100; (4) reach >90% on MNIST 10-class softmax regression in PyTorch.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Binary cross-entropy from scratch · NumPy vs PyTorch implementation · Softmax regression and multi-class · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
BCEWithLogitsLoss (not sigmoid + BCELoss) for numerical stabilityCrossEntropyLoss — pass raw logits onlyweight_decay) smooths the decision boundary| idea | the one thing to remember |
|---|---|
| gradient formula | X^T(p−y)/n — sigmoid derivative cancels |
| BCEWithLogitsLoss | log-sum-exp trick — finite for any logit |
| CrossEntropyLoss | owns the softmax — never pre-apply it |
| regularization | weight_decay shrinks ||w||, widens margin |
| vs linear SVM | same boundary, different loss; logistic gives calibrated P |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.