USAAIO Lesson 17, from Week 6, fully worked. It explains what a loss is, derives MSE and MAE term by term along with their gradients, and builds softmax and cross-entropy from scratch. It then proves the elegant result dL/dz = p - y TWICE, once by the softmax-Jacobian chain rule for the multi-class case and once directly for sigmoid with BCE, with no steps skipped. From there it explains why MSE stalls a classifier, through the vanishing p(1-p) factor, covers hinge loss and subgradients, gives a loss-to-gradient cheat sheet, and ends with a from-scratch build verified against torch autograd. One running example threads the whole deck - the regression y=[3,5,7] against f=[2,6,4], and the logits z=[1,2,3] with true class 2 - and every number was produced by real execution. The lesson runs to 60 slides.
Subject: Machine Learning · 113 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 17 · Week 6
Every loss is a choice, and every gradient has a shape. We build MSE, MAE, cross-entropy, and hinge from the ground up, derive the elegant dL/dz = p - y twice, and prove why MSE quietly stalls a classifier - skipping no step and checking every number against torch.
Objectives
dL/dz = p - y for softmax+CE via the softmax Jacobian, and independently for sigmoid+BCEtorch autograd to 1e-6Warm-up
Discussion prompt
Before we open Lesson 17: Loss Functions & Their Gradients: without looking back, what was the main idea of Singular Value Decomposition, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the SVD A=UΣVᵀ for every matrix, singular values as roots of eigenvalues of AᵀA, the rotate-scale-rotate geometry, low-rank approximation and Eckart-Young, and applications to PCA and the pseudoinverse. Build SVD compression and a pseudoinverse solver from scratch.
Section
Part 1 of 7
Concept
A model outputs a prediction f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.
Training means making L small. And the only thing an optimizer knows how to do with L is follow its gradient downhill. So a loss is only useful if its gradient points somewhere helpful.
Counterexample
Discussion prompt
A model outputs a prediction f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Training means making L small. And the only thing an optimizer knows how to do with L is follow its gradient downhill. So a loss is only useful if its gradient points somewhere helpful.
Intuition
You will never minimize a loss by hand. Gradient descent does it: nudge each parameter opposite the gradient, repeat. The loss value is a scoreboard; the loss gradient is the steering wheel.
So for every loss today we ask two questions: what is L (the number), and what is dL/d(prediction) (the direction). The second one is what actually trains the model.
That is why two losses that look similar can behave completely differently - a good-looking loss with a dead gradient in the wrong place will simply refuse to learn. We will watch exactly that happen to MSE.
Concept
One regression example and one classification example thread the whole deck, so every slide compounds. Regression: three targets and three model predictions.
| i | target y | prediction f | residual y - f |
|---|---|---|---|
| 0 | 3 | 2 | +1 |
| 1 | 5 | 6 | -1 |
| 2 | 7 | 4 | +3 |
Classification: raw logits z = [1, 2, 3] over three classes, with the true class = 2. We will keep returning to these exact numbers.
Comparison
Comparison matrix
From Our running examples: refill the target y column from what you know. The rest of the table is as it appeared.
| i | target y | prediction f | residual y - f |
|---|---|---|---|
| 0 | 3 | 2 | +1 |
| 1 | 5 | 6 | -1 |
| 2 | 7 | 4 | +3 |
Section
Part 2 of 7 - the regression staple
Concept
Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.
\[ \text{MSE}(y, f) = \frac{1}{n}\sum_{i=1}^{n}(y_i - f_i)^2 = \frac{1}{n}\lVert y - f \rVert^2 \]
It is smooth everywhere, so calculus finds its minimum with no trouble - the property that made it the backbone of Lesson 7's least squares.
Analogy
Discussion prompt
Explain MSE: average the squared misses by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.
Estimation
Predict first
Take y = [3, 5, 7], f = [2, 6, 4]. Residuals are [+1, -1, +3]. Square, sum, divide by n = 3:
Commit before you compute: what does MSE on our data, entry by entry come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Sum then divide: (1 + 1 + 9) / 3 = 11/3
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The mean keeps the loss on the same scale regardless of dataset size n, so learning rates transfer across datasets.
Worked example
Take y = [3, 5, 7], f = [2, 6, 4]. Residuals are [+1, -1, +3]. Square, sum, divide by n = 3:
Square each residual: 1² = 1, (-1)² = 1, 3² = 9
Why: Squaring kills the sign, so an overshoot and an undershoot of the same size cost the same. The big miss (+3) dominates: 9 out of a total of 11.
Sum then divide: (1 + 1 + 9) / 3 = 11/3
Why: The mean keeps the loss on the same scale regardless of dataset size n, so learning rates transfer across datasets.
\[ \text{MSE} = \frac{1 + 1 + 9}{3} = \frac{11}{3} \approx 3.6667 \]
| i | yᵢ | fᵢ | rᵢ = yᵢ - fᵢ | rᵢ² |
|---|---|---|---|---|
| 0 | 3 | 2 | +1 | 1 |
| 1 | 5 | 6 | -1 | 1 |
| 2 | 7 | 4 | +3 | 9 |
Trade off
Comparison matrix
From MSE on our data, entry by entry: every row here is a choice with a cost. Fill the rᵢ = yᵢ - fᵢ column, then say which row you would actually pick and what you give up for it.
| i | yᵢ | fᵢ | rᵢ = yᵢ - fᵢ | rᵢ² |
|---|---|---|---|---|
| 0 | 3 | 2 | +1 | 1 |
| 1 | 5 | 6 | -1 | 1 |
| 2 | 7 | 4 | +3 | 9 |
Ranking
Put in order
Put the moves of Differentiate MSE - one prediction first into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Chain rule: d(u²)/df = 2u · du/df.
Worked example
The gradient is per-prediction: how does L change if we wiggle a single fᵢ? Isolate one term L = (yᵢ - fᵢ)²/n and differentiate with respect to fᵢ.
Let u = yᵢ - fᵢ, so the term is u²/n
Why: Chain rule: d(u²)/df = 2u · du/df. We differentiate the outer square, then the inside.
\[ \frac{\partial}{\partial f_i}\,\frac{(y_i - f_i)^2}{n} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot\frac{\partial(y_i - f_i)}{\partial f_i} \]
Inside derivative: d(yᵢ - fᵢ)/dfᵢ = -1
Why: yᵢ is a constant target; only -fᵢ depends on fᵢ. That minus sign is the whole reason the gradient has a leading minus.
\[ \frac{\partial L}{\partial f_i} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot(-1) = -\frac{2}{n}(y_i - f_i) \]
Stack all i into a vector
Why: Every coordinate has the same form, so the full gradient is one clean vector expression.
\[ \boxed{\;\nabla_f\,\text{MSE} = -\frac{2}{n}(y - f)\;} \]
Notation
Annotate
From Differentiate MSE - one prediction first — read this one piece at a time. What is each part doing?
On: \( \frac{\partial}{\partial f_i}\,\frac{(y_i - f_i)^2}{n} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot\frac{\partial(y_i - f_i)}{\partial f_i} \)
Missing information
Discussion prompt
Compute the value and the gradient -2(y - f)/n directly. Self-contained - re-imports NumPy and re-defines y, f:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
MSE's gradient GROWS with the error, so the worst-fit point (i=2) gets the largest correction. Great for clean data, dangerous with outliers.
Worked example
Compute the value and the gradient -2(y - f)/n directly. Self-contained - re-imports NumPy and re-defines y, f:
import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mse = np.mean((y - f)**2)
grad = -2*(y - f)/n
print(round(mse, 4))
print(grad.round(4))grad[2] = -2/3 · (7 - 4) = -2 is the biggest push
Why: MSE's gradient GROWS with the error, so the worst-fit point (i=2) gets the largest correction. Great for clean data, dangerous with outliers.
| quantity | formula | value (verified) |
|---|---|---|
| MSE | (1 + 1 + 9)/3 | 3.6667 |
| grad[0] | -2/3 · (+1) | -0.6667 |
| grad[1] | -2/3 · (-1) | +0.6667 |
| grad[2] | -2/3 · (+3) | -2.0 |
Pattern
Step through it
Step through MSE gradient, verified in code one row at a time. What is driving the change, and what would the row after the last one be?
Intuition
At i = 2 we under-predicted (f = 4 < y = 7), so the residual is positive and the gradient is negative. Gradient descent steps opposite the gradient - so it raises f. Exactly what we want: predict higher.
At i = 1 we over-predicted (f = 6 > y = 5), the gradient is positive, descent lowers f. The sign of the residual, flipped by descent, always pushes the prediction toward the target.
Explain it
Discussion prompt
Explain Why the sign points the right way to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
At i = 1 we over-predicted (f = 6 > y = 5), the gradient is positive, descent lowers f. The sign of the residual, flipped by descent, always pushes the prediction toward the target.
Intuition
Imagine point 2's target were a wild outlier - say 20 instead of 7, so the residual is +16. MSE squares it to 256, and its gradient becomes -2/3 · 16 = -10.67: one bad point hijacks the whole update.
MAE would give that same outlier a gradient of just -1/3 - the same as every other point. That capped, democratic gradient is why MAE is called robust, and it's the reason we bother with a second regression loss at all.
So the choice is a trade-off: MSE for clean data and smooth optimization, MAE when outliers would otherwise dominate. The gradient shape - growing vs flat - is the whole difference.
Section
Part 3 of 7 - robust, but with a corner
Concept
Mean absolute error uses the absolute value instead of the square, so a miss of 3 costs 3, not 9 - big errors are not amplified.
\[ \text{MAE}(y, f) = \frac{1}{n}\sum_{i=1}^{n}\lvert y_i - f_i\rvert = \frac{1}{n}\lVert y - f\rVert_1 \]
That makes MAE robust to outliers - one wild point can't dominate the sum. The price is a gradient that behaves very differently, as we'll see.
Pattern
Predict first
The table runs: 0 | +1 | 1 | +1 · 1 | -1 | 1 | -1
In MAE on our data, given the rows so far: what is the next one — the row where i is 2?
Correct: 2 | +3 | 3 | +1
| i | rᵢ = yᵢ - fᵢ | |rᵢ| | sign(rᵢ) |
|---|---|---|---|
| 0 | +1 | 1 | +1 |
| 1 | -1 | 1 | -1 |
| 2 | +3 | 3 | +1 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Compare to MSE's numerator of 11: the big miss (+3) contributes 3 here versus 9 there.
Worked example
Same residuals [+1, -1, +3], now take absolute values [1, 1, 3], average over n = 3:
|+1| + |-1| + |+3| = 1 + 1 + 3 = 5
Why: Compare to MSE's numerator of 11: the big miss (+3) contributes 3 here versus 9 there. MAE refuses to let one point take over.
\[ \text{MAE} = \frac{1 + 1 + 3}{3} = \frac{5}{3} \approx 1.6667 \]
| i | rᵢ = yᵢ - fᵢ | |rᵢ| | sign(rᵢ) |
|---|---|---|---|
| 0 | +1 | 1 | +1 |
| 1 | -1 | 1 | -1 |
| 2 | +3 | 3 | +1 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
|+1| + |-1| + |+3| = 1 + 1 + 3 = 5
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Same residuals [+1, -1, +3], now take absolute values [1, 1, 3], average over n = 3:
Step zero
Discussion prompt
Differentiate MAE - the sign appears — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: d|u|/du = sign(u) for u ≠ 0
Answer:
Worked example
One term is |yᵢ - fᵢ|/n. The derivative of |u| is sign(u), then chain through u = yᵢ - fᵢ whose fᵢ-derivative is -1:
d|u|/du = sign(u) for u ≠ 0
Why: |u| has slope +1 when u > 0 and -1 when u < 0. That is exactly the sign function - as long as u is not 0.
\[ \frac{\partial}{\partial f_i}\,\frac{\lvert y_i - f_i\rvert}{n} = \frac{1}{n}\,\text{sign}(y_i - f_i)\cdot(-1) \]
Fold in the -1
Why: Same inside-derivative -1 as MSE. The result is constant magnitude 1/n, only its sign varies.
\[ \boxed{\;\nabla_f\,\text{MAE} = -\frac{1}{n}\,\text{sign}(y - f)\;} \]
Contrast with MSE: magnitude is FLAT
Why: MSE grad scales with the error; MAE grad is always ±1/n. A point that's off by 100 gets the same push as one off by 0.1 - robust, but slow to correct large errors.
Blank canvas
Draw it
Draw what Differentiate MAE - the sign appears just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Fill the middle
Fill in the blanks
From MAE gradient, verified in code — one line has had its right-hand side removed. Put it back.
import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mae = np.mean(np.abs(y - f))
grad = -np.sign(y - f)/n
print(round(mae, 4))
print(grad.round(4))
Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. Every gradient entry is ±1/3. The big miss at i=2 gets the SAME magnitude as the small misses - unlike MSE, where it was -2.0.
Worked example
Value and gradient -sign(y - f)/n. Self-contained:
import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mae = np.mean(np.abs(y - f))
grad = -np.sign(y - f)/n
print(round(mae, 4))
print(grad.round(4))grad = [-0.333, +0.333, -0.333] - all the same size
Why: Every gradient entry is ±1/3. The big miss at i=2 gets the SAME magnitude as the small misses - unlike MSE, where it was -2.0.
| i | sign(y - f) | grad = -sign/n |
|---|---|---|
| 0 | +1 | -0.3333 |
| 1 | -1 | +0.3333 |
| 2 | +1 | -0.3333 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
MAE is a simple linear-looking function of the error, so its gradient is just ±1/n everywhere, including when the prediction is exactly right.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right).
MAE is non-differentiable at 0. There, use a subgradient - any value in the interval [-1/n, +1/n].
Why: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.
Trap
MAE is a simple linear-looking function of the error, so its gradient is just ±1/n everywhere, including when the prediction is exactly right.
Report sign(0) as some fixed value at the kink
Why: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.
MAE is non-differentiable at 0. There, use a subgradient - any value in the interval [-1/n, +1/n].
Subgradient (usually 0) at the kink, sign elsewhere
Why: Optimizers pick a value in [-1/n, 1/n] (numpy's sign(0) = 0) and march on. This same subgradient idea handles the hinge kink and L1 (Lesson 12) - a corner is fine, you just choose a slope.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Optimizers pick a value in [-1/n, 1/n] (numpy's sign(0) = 0) and march on. This same subgradient idea handles the hinge kink and L1 (Lesson 12) - a corner is fine, you just choose a slope.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.
Section
Part 4 of 7 - the classification loss
Concept
A classifier outputs raw scores called logits z. Softmax turns them into a probability vector p - all positive, summing to 1 - by exponentiating and normalizing.
\[ p_k = \operatorname{softmax}(z)_k = \frac{e^{z_k}}{\sum_j e^{z_j}} \]
logit — A raw, unbounded real-valued score before the softmax/sigmoid. Positive logits become high probabilities, negative become low. The gradient dL/dz is taken with respect to THESE, not the probabilities.
Worked example
e^z overflows for large z. Softmax is unchanged if we subtract any constant c from every logit (it cancels top and bottom), so subtract c = max(z). For z = [1, 2, 3], c = 3:
\[ \frac{e^{z_k}}{\sum_j e^{z_j}} = \frac{e^{z_k - c}}{\sum_j e^{z_j - c}}\quad\text{(multiply top and bottom by } e^{-c}) \]
Shifted exponentials: e^(1-3), e^(2-3), e^(3-3)
Why: = e^-2, e^-1, e^0 = 0.1353, 0.3679, 1.0. The largest is now exactly 1, so nothing overflows - and the ratios are identical to the unshifted version.
| k | zₖ | zₖ - 3 | e^(zₖ - 3) |
|---|---|---|---|
| 0 | 1 | -2 | 0.1353 |
| 1 | 2 | -1 | 0.3679 |
| 2 | 3 | 0 | 1.0000 |
Pattern
Step through it
Step through The stability trick: subtract the max one row at a time. What is driving the change, and what would the row after the last one be?
Faded example
Fill in the blanks
Finish softmax: normalize, with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z = np.array([1., 2., 3.])
p = softmax(z)
print(p.round(4))
onehot = np.array([0., 0., 1.])
print((p - onehot).round(4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Class 2 (the largest logit) gets the largest probability, 0.665.
Worked example
Divide each shifted exponential by their sum 0.1353 + 0.3679 + 1.0 = 1.5032:
import numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z = np.array([1., 2., 3.])
p = softmax(z)
print(p.round(4))
onehot = np.array([0., 0., 1.])
print((p - onehot).round(4))p = [0.0900, 0.2447, 0.6652], sums to 1
Why: Class 2 (the largest logit) gets the largest probability, 0.665. The model is only moderately confident - which is exactly why there's a real gradient to learn from.
| k | e^(zₖ-3) | pₖ = e / 1.5032 |
|---|---|---|
| 0 | 0.1353 | 0.0900 |
| 1 | 0.3679 | 0.2447 |
| 2 | 1.0000 | 0.6652 |
Fill the middle
Fill in the blanks
From Prove the max-subtraction is exact — one line has had its right-hand side removed. Put it back.
import numpy as np
def softmax(z):
e = np.exp(z - z.max()); return e / e.sum()
def naive(z):
e = np.exp(z); return e / e.sum()
z = np.array([1., 2., 3.])
print(softmax(z).round(6))
print(naive(z).round(6))
print(np.allclose(softmax(z), naive(z)))
Why: z is what everything below it consumes, so the wrong expression here fails later and somewhere else. Subtracting max(z) multiplies numerator and denominator by the same e^{-c}, which cancels - so the probabilities are identical, just computed without overflow risk.
Worked example
The stability trick must not change the answer. Compare the stable softmax (subtract the max) against the naive e^z / Σe^z on our logits - they should agree to machine precision. Self-contained:
import numpy as np
def softmax(z):
e = np.exp(z - z.max()); return e / e.sum()
def naive(z):
e = np.exp(z); return e / e.sum()
z = np.array([1., 2., 3.])
print(softmax(z).round(6))
print(naive(z).round(6))
print(np.allclose(softmax(z), naive(z)))Both give [0.090031, 0.244728, 0.665241]; allclose = True
Why: Subtracting max(z) multiplies numerator and denominator by the same e^{-c}, which cancels - so the probabilities are identical, just computed without overflow risk. Always subtract the max in real code.
| method | p₀ | p₁ | p₂ |
|---|---|---|---|
| stable (z - max) | 0.090031 | 0.244728 | 0.665241 |
| naive | 0.090031 | 0.244728 | 0.665241 |
| np.allclose | - | - | True |
Concept
Cross-entropy looks at the probability the model assigned to the true class and charges -log of it. Confident and right → near 0; confident and wrong → huge.
\[ \text{CE} = -\sum_{k} y_k \log p_k \;\overset{\text{one-hot } y}{=}\; -\log p_{\text{true}} \]
With a one-hot y, every term is zero except the true class, so the sum collapses to a single -log p_true. The binary special case is BCE: -y log p - (1-y) log(1-p).
Estimation
Predict first
True class is 2, and p₂ = 0.6652. Cross-entropy is just -log(0.6652):
Commit before you compute: what does CE value on our logits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Sanity: if p_true were 1.0, CE = -log 1 = 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A perfect confident prediction has zero loss; as p_true → 0 the loss → +infinity.
Worked example
True class is 2, and p₂ = 0.6652. Cross-entropy is just -log(0.6652):
CE = -log(0.6652) = 0.4076
Why: The other two probabilities (0.09, 0.2447) never enter the loss VALUE - only p_true does. They will, however, enter the gradient.
\[ \text{CE} = -\log(0.6652) \approx 0.4076 \]
Sanity: if p_true were 1.0, CE = -log 1 = 0
Why: A perfect confident prediction has zero loss; as p_true → 0 the loss → +infinity. That steepness near 0 is what makes CE punish confident mistakes hard.
| quantity | value (verified) |
|---|---|
| p_true = p₂ | 0.6652 |
| -log(p₂) | 0.4076 |
| torch cross_entropy | 0.4076 |
Section
Part 5 of 7 - the derivation, every step
Concept
The gradient dL/dp of cross-entropy alone is messy (it has a 1/p). The magic happens when we go one step further back and differentiate w.r.t. the logits z - the softmax Jacobian cancels that 1/p.
\[ \frac{\partial L}{\partial z_j} = \sum_k \frac{\partial L}{\partial p_k}\,\frac{\partial p_k}{\partial z_j} \quad\text{(chain rule through the softmax)} \]
We need two pieces: the loss-to-probability gradient dL/dp, and the softmax Jacobian dp/dz. Then we multiply and watch everything collapse.
Intuition
The network computes logits z → softmax → p → loss. Backprop hands each layer the gradient w.r.t. its output. The softmax layer's job is to convert dL/dp (messy, has a 1/p) into dL/dz (clean) before passing it on.
So stopping at dL/dp is the wrong altitude - it's an intermediate quantity. The number that actually flows into the weights is dL/dz. We chase it back one link, through the softmax, and the 1/p disappears.
Step zero
Discussion prompt
Piece 1: dL/dp for cross-entropy — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: L = -Σ y_k log p_k
Answer:
Worked example
L = -Σ y_k log p_k
Why: Cross-entropy written over all classes, before assuming one-hot.
\[ L = -\sum_k y_k \log p_k \]
Differentiate w.r.t. a single p_k: d(-y_k log p_k)/dp_k = -y_k / p_k
Why: Only the k-th term depends on p_k; d(log p)/dp = 1/p. So the loss-to-probability gradient is -y_k / p_k.
\[ \frac{\partial L}{\partial p_k} = -\frac{y_k}{p_k} \]
One-hot y: only the true class survives
Why: y_k = 0 except at the true class, so dL/dp is a single nonzero entry -1/p_true. For our data that's -1/0.6652 = -1.5032.
Ranking
Put in order
Put the moves of Piece 2: the softmax Jacobian into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The numerator e^{zₖ} depends on zₖ too, so you get pₖ minus pₖ² from the quotient rule.
Worked example
Differentiate pₖ = e^{zₖ}/Σⱼe^{zⱼ} w.r.t. z_j. Quotient rule splits into the diagonal case (k = j) and off-diagonal (k ≠ j):
Diagonal k = j: ∂pₖ/∂zₖ = pₖ(1 - pₖ)
Why: The numerator e^{zₖ} depends on zₖ too, so you get pₖ minus pₖ² from the quotient rule.
\[ \frac{\partial p_k}{\partial z_k} = p_k(1 - p_k) \]
Off-diagonal k ≠ j: ∂pₖ/∂z_j = -pₖ p_j
Why: Here only the denominator sees z_j, giving a negative coupling. Raising one logit steals probability from the others.
\[ \frac{\partial p_k}{\partial z_j} = -p_k p_j \]
Both cases in one line
Why: δ is the Kronecker delta (1 if k=j else 0). This single formula is the entire softmax Jacobian.
\[ \frac{\partial p_k}{\partial z_j} = p_k(\delta_{kj} - p_j) \]
Translation
\( \frac{\partial p_k}{\partial z_j} = p_k(\delta_{kj} - p_j) \)
Draw it
Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.
Step zero
Discussion prompt
Multiply the pieces - the cancellation — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Chain: ∂L/∂z_j = Σ_k (-y_k/p_k) · p_k(δ_kj - p_j)
Answer:
Worked example
Chain: ∂L/∂z_j = Σ_k (-y_k/p_k) · p_k(δ_kj - p_j)
Why: Substitute Piece 1 and Piece 2 into the chain-rule sum over k.
\[ \frac{\partial L}{\partial z_j} = \sum_k \Big(-\frac{y_k}{p_k}\Big)\,p_k(\delta_{kj} - p_j) \]
The pₖ cancels
Why: (-y_k/p_k)·p_k = -y_k. The messy 1/p is gone - this is the cancellation that makes the result clean.
\[ = -\sum_k y_k(\delta_{kj} - p_j) = -\sum_k y_k \delta_{kj} + p_j\sum_k y_k \]
Use Σy_k = 1 and Σy_k δ_kj = y_j
Why: Probabilities of the one-hot target sum to 1, and the delta picks out the j-th entry. So the first term is -y_j and the second is p_j·1.
\[ \frac{\partial L}{\partial z_j} = p_j - y_j \;\Longrightarrow\; \boxed{\;\frac{\partial L}{\partial z} = p - y\;} \]
Intuition
The gradient of the loss w.r.t. the logits is literally the prediction error: how far each predicted probability sits from its target. No 1/p, no Jacobian left over - just p - y.
The true class has p - y < 0, so descent raises its logit; every wrong class has p - y > 0, so descent lowers theirs. The size of each push is exactly how wrong that class was. That is why classifiers train cleanly.
Faded example
Fill in the blanks
Verify p - y numerically (Jacobian route), with the scaffolding fading: two lines are gone now — fill both.
import numpy as np
z = np.array([1., 2., 3.])
e = np.exp(z - z.max())
p = e / e.sum()
J = np.diag(p) - np.outer(p, p)
onehot = np.array([0., 0., 1.])
dLdp = -onehot / p
dLdz = J.T @ dLdp
print(dLdz.round(4))
print((p - onehot).round(4))
Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The full chain-rule product equals p - y exactly, confirming the algebra.
Worked example
Build the Jacobian J = diag(p) - p pᵀ, form dL/dp = -y/p, and confirm Jᵀ (dL/dp) = p - y. Self-contained:
import numpy as np
z = np.array([1., 2., 3.])
e = np.exp(z - z.max())
p = e / e.sum()
J = np.diag(p) - np.outer(p, p)
onehot = np.array([0., 0., 1.])
dLdp = -onehot / p
dLdz = J.T @ dLdp
print(dLdz.round(4))
print((p - onehot).round(4))Jᵀ (dL/dp) = [0.0900, 0.2447, -0.3348] = p - onehot
Why: The full chain-rule product equals p - y exactly, confirming the algebra. The one-hot has its 1 at class 2, so only that entry is negative.
| k | via Jacobian | p - y | match |
|---|---|---|---|
| 0 | 0.0900 | 0.0900 | yes |
| 1 | 0.2447 | 0.2447 | yes |
| 2 | -0.3348 | -0.3348 | yes |
Invariant
Step through it
Step through Verify p - y numerically (Jacobian route) one row at a time. One of these columns never changes — find it, and say why it cannot.
Worked example
The ultimate check: does real backprop agree? Build z with requires_grad, run cross_entropy, and read z.grad. Self-contained:
import torch, numpy as np
z = torch.tensor([1., 2., 3.], requires_grad=True)
loss = torch.nn.functional.cross_entropy(z.unsqueeze(0), torch.tensor([2]))
loss.backward()
p = torch.softmax(z, 0).detach().numpy()
print(p.round(4))
print(z.grad.numpy().round(4))
print((p - np.array([0., 0., 1.])).round(4))z.grad = [0.0900, 0.2447, -0.3348] = p - onehot
Why: torch's autograd produces exactly p - y. The true class (2) gets the negative push; the impostors get positive pushes proportional to their probability.
| class | p | y (one-hot) | z.grad = p - y |
|---|---|---|---|
| 0 | 0.0900 | 0 | +0.0900 |
| 1 | 0.2447 | 0 | +0.2447 |
| 2 (true) | 0.6652 | 1 | -0.3348 |
Comparison
Comparison matrix
From Verify against torch autograd: refill the y (one-hot) column from what you know. The rest of the table is as it appeared.
| class | p | y (one-hot) | z.grad = p - y |
|---|---|---|---|
| 0 | 0.0900 | 0 | +0.0900 |
| 1 | 0.2447 | 0 | +0.2447 |
| 2 (true) | 0.6652 | 1 | -0.3348 |
Fill the middle
Fill in the blanks
From The binary case: sigmoid + BCE gives p - y too — one line has had its right-hand side removed. Put it back.
import torch
z = torch.tensor([0.5], requires_grad=True)
loss = torch.nn.functional.binary_cross_entropy_with_logits(z, torch.tensor([1.]))
loss.backward()
p = torch.sigmoid(z).item()
print(round(p, 4), round(z.grad.item(), 4), round(p - 1, 4))
Why: p is what everything below it consumes, so the wrong expression here fails later and somewhere else. Same clean form. The 0.3775 gap is exactly how far the prediction (0.6225) sits below the target (1.0).
Worked example
For a single logit, softmax becomes the sigmoid σ(z) = 1/(1 + e^{-z}) and CE becomes BCE. The identity dL/dz = p - y still holds. Take z = 0.5, label y = 1:
import torch
z = torch.tensor([0.5], requires_grad=True)
loss = torch.nn.functional.binary_cross_entropy_with_logits(z, torch.tensor([1.]))
loss.backward()
p = torch.sigmoid(z).item()
print(round(p, 4), round(z.grad.item(), 4), round(p - 1, 4))p = 0.6225, dL/dz = -0.3775 = p - 1
Why: Same clean form. The 0.3775 gap is exactly how far the prediction (0.6225) sits below the target (1.0).
| quantity | value (verified) |
|---|---|
| p = σ(0.5) | 0.6225 |
| dL/dz (autograd) | -0.3775 |
| p - y | -0.3775 |
Step zero
Discussion prompt
Prove sigmoid + BCE = p - y by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: L = -y log p - (1-y) log(1-p), p = σ(z)
Answer:
Worked example
L = -y log p - (1-y) log(1-p), p = σ(z)
Why: BCE in terms of the logit z through the sigmoid p.
\[ \frac{\partial L}{\partial p} = -\frac{y}{p} + \frac{1-y}{1-p} \]
Sigmoid derivative: dp/dz = p(1 - p)
Why: The 1-D version of the softmax diagonal. This is the factor that will cancel the denominators.
\[ \frac{\partial L}{\partial z} = \Big(-\frac{y}{p} + \frac{1-y}{1-p}\Big)\,p(1-p) \]
Distribute p(1-p): the denominators cancel
Why: -y/p · p(1-p) = -y(1-p); (1-y)/(1-p) · p(1-p) = (1-y)p. Add them.
\[ = -y(1-p) + (1-y)p = -y + yp + p - yp = p - y \]
Check: p - y = 0.6225 - 1 = -0.3775
Why: Matches the autograd value exactly. Both softmax+CE and sigmoid+BCE collapse to p - y - it is not a coincidence, it's the structure of the exponential-family loss paired with its natural link.
Blank canvas
Draw it
Draw what Prove sigmoid + BCE = p - y by hand just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.
Pattern
Predict first
The table runs: -2.0 | 0.119203 | 0.119203 · 0.0 | 0.500000 | 0.500000 · 0.5 | 0.622459 | 0.622459
In Sigmoid IS the two-class softmax, given the rows so far: what is the next one — the row where z is 3.0?
Correct: 3.0 | 0.952574 | 0.952574
| z | softmax([0,z])[1] | sigmoid(z) |
|---|---|---|
| -2.0 | 0.119203 | 0.119203 |
| 0.0 | 0.500000 | 0.500000 |
| 0.5 | 0.622459 | 0.622459 |
| 3.0 | 0.952574 | 0.952574 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two are algebraically identical, not merely close.
Worked example
The binary and multi-class cases are the same object. Softmax over logits [0, z] gives class-1 probability e^z/(1 + e^z) = σ(z). So sigmoid+BCE is just softmax+CE with two classes - which is why they share dL/dz = p - y. Check numerically:
import numpy as np
def softmax(z):
e = np.exp(z - z.max()); return e / e.sum()
def sigmoid(z): return 1/(1 + np.exp(-z))
for z in [-2., 0., 0.5, 3.]:
two_class = softmax(np.array([0., z]))[1]
print(z, round(two_class, 6), round(sigmoid(z), 6))softmax([0, z])[1] equals sigmoid(z) for every z
Why: The two are algebraically identical, not merely close. That is the deep reason the p - y result is one result, not two coincidences.
| z | softmax([0,z])[1] | sigmoid(z) |
|---|---|---|
| -2.0 | 0.119203 | 0.119203 |
| 0.0 | 0.500000 | 0.500000 |
| 0.5 | 0.622459 | 0.622459 |
| 3.0 | 0.952574 | 0.952574 |
Scale up
Step through it
Step through Sigmoid IS the two-class softmax and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?
Section
Part 6 of 7 - the vanishing gradient
Concept
Suppose you (wrongly) train a classifier with sigmoid output and MSE loss L = ½(p - y)². The chain rule to the logit picks up the sigmoid derivative σ'(z) = p(1 - p):
\[ \frac{\partial L}{\partial z} = (p - y)\cdot\underbrace{p(1-p)}_{\sigma'(z)} \]
Compare with cross-entropy's clean p - y. MSE carries an extra p(1-p) multiplier - and that multiplier is exactly what breaks training.
Intuition
p(1-p) is largest at p = 0.5 (value 0.25) and shrinks to 0 as p approaches 0 or 1. So when the model is confident - p ≈ 0 or p ≈ 1 - the multiplier is nearly zero.
Now picture a confident mistake: true label y = 1 but the model says p ≈ 0. Cross-entropy screams (p - y ≈ -1). MSE multiplies that by p(1-p) ≈ 0 and whispers - almost no gradient, right when you need the most.
Explain it
Discussion prompt
Explain The multiplier dies when you're confidently wrong to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
p(1-p) is largest at p = 0.5 (value 0.25) and shrinks to 0 as p approaches 0 or 1. So when the model is confident - p ≈ 0 or p ≈ 1 - the multiplier is nearly zero.
Missing information
Discussion prompt
True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
CE keeps pushing hard on the confident mistake; MSE's gradient is 57x smaller because p(1-p) = 0.0177 crushed it. An MSE classifier gets STUCK on its worst errors.
Worked example
True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:
import numpy as np
def sigmoid(z): return 1/(1 + np.exp(-z))
for z in [-4., 0.]:
p = sigmoid(z)
sg = p*(1 - p)
ce = p - 1 # cross-entropy dL/dz
mse = (p - 1)*sg # MSE-through-sigmoid dL/dz
print(z, round(p, 4), round(ce, 4), round(mse, 6))At z = -4: CE grad -0.982 (strong) vs MSE grad -0.017 (≈ 0)
Why: CE keeps pushing hard on the confident mistake; MSE's gradient is 57x smaller because p(1-p) = 0.0177 crushed it. An MSE classifier gets STUCK on its worst errors.
| logit z | p | p(1-p) | CE dL/dz | MSE dL/dz |
|---|---|---|---|---|
| -4.0 | 0.0180 | 0.0177 | -0.9820 | -0.017345 |
| 0.0 | 0.5000 | 0.2500 | -0.5000 | -0.125000 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
At z = -4: CE grad -0.982 (strong) vs MSE grad -0.017 (≈ 0)
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:
Anomaly
Predict first
A student writes this, and it looks reasonable:
A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix.
Match the loss to the output: cross-entropy for classification, MSE for regression.
Why: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix. Verified: gradient -0.017 vs CE's -0.982 at a confident error.
Trap
A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.
Train with sigmoid + MSE, watch it stall
Why: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix. Verified: gradient -0.017 vs CE's -0.982 at a confident error.
Match the loss to the output: cross-entropy for classification, MSE for regression.
Use softmax/sigmoid + cross-entropy
Why: Then dL/dz = p - y - a full-strength gradient on every error, with no vanishing multiplier. The loss is deliberately chosen so its gradient stays healthy everywhere.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.; One regression example and one classification example thread the whole deck, so every slide compounds. Regression: three targets and three model predictions.; Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.±1/n everywhere, including when the prediction is exactly right.; A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.Section
Part 7 of 7 - margins & subgradients
Concept
The SVM loss uses labels y ∈ {-1, +1} and a raw score f(x). It penalizes a point only until it is correctly classified with margin (y·f ≥ 1); beyond that, zero.
\[ L = \max\!\big(0,\; 1 - y\,f(x)\big) \]
The quantity y·f is the signed margin: positive when the sign of f matches the label, and its size says how confidently. Hinge cares only when that margin is below 1.
Analogy
Discussion prompt
Explain Hinge: only punish inside the margin by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The SVM loss uses labels y ∈ {-1, +1} and a raw score f(x). It penalizes a point only until it is correctly classified with margin (y·f ≥ 1); beyond that, zero.
Ranking
Put in order
Put the moves of Hinge gradient - three regimes into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The point is safely beyond the margin - already correct with room to spare - so it contributes no push at all.
Worked example
Differentiate max(0, 1 - y·f) w.r.t. f. Two smooth pieces meet at the kink y·f = 1:
If y·f > 1: L = 0, gradient 0
Why: The point is safely beyond the margin - already correct with room to spare - so it contributes no push at all.
\[ \frac{\partial L}{\partial f} = \begin{cases} 0 & y f > 1 \\ -y & y f < 1 \end{cases} \]
If y·f < 1: L = 1 - y·f, gradient -y
Why: Inside the margin the loss is linear with slope -y in f. Descent (opposite -y) moves f in the +y direction, pushing the point toward the correct side.
At y·f = 1 exactly: subgradient in [-y, 0]
Why: Same corner story as MAE - the two slopes (0 and -y) disagree at the kink, so any value between them is a valid subgradient. Optimizers just pick one.
Notation
Annotate
From Hinge gradient - three regimes — read this one piece at a time. What is each part doing?
On: \( \frac{\partial L}{\partial f} = \begin{cases} 0 & y f > 1 \\ -y & y f < 1 \end{cases} \)
Estimation
Predict first
Evaluate loss and gradient at three points. Self-contained:
Commit before you compute: what does Hinge on three points come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: (y=+1, f=1.5): y·f = 1.5 > 1 → loss 0, grad 0
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Correct and beyond the margin. The SVM ignores it entirely - only support vectors near the boundary matter.
Worked example
Evaluate loss and gradient at three points. Self-contained:
import numpy as np
def hinge(y, f): return max(0., 1 - y*f)
def hinge_grad(y, f): return 0. if y*f > 1 else -y
for y, f in [(1, 1.5), (1, 0.3), (-1, 0.3)]:
print(y, f, round(y*f, 2), round(hinge(y, f), 2), hinge_grad(y, f))(y=+1, f=1.5): y·f = 1.5 > 1 → loss 0, grad 0
Why: Correct and beyond the margin. The SVM ignores it entirely - only support vectors near the boundary matter.
| y | f | y·f | loss | gradient ∂L/∂f |
|---|---|---|---|---|
| +1 | 1.5 | +1.5 | 0.00 | 0 |
| +1 | 0.3 | +0.3 | 0.70 | -1 |
| -1 | 0.3 | -0.3 | 1.30 | +1 |
Pattern
The whole lesson in one table. Notice how the two classification rows share the same clean p - y, and how every kink comes with a subgradient.
| loss | value | gradient | use / note |
|---|---|---|---|
| MSE | ‖y - f‖²/n | -2(y - f)/n | regression; grows with error, outlier-sensitive |
| MAE | ‖y - f‖₁/n | -sign(y - f)/n | regression; robust; subgradient at 0 |
| softmax + CE | -log p_true | p - y (w.r.t. logits) | multi-class; clean, full-strength gradient |
| sigmoid + BCE | -log p_true | p - y (w.r.t. logit) | binary; same clean form |
| hinge | max(0, 1 - yf) | 0 if yf>1 else -y | SVM margin; subgradient at yf=1 |
Edge cases
Discussion prompt
Loss → gradient cheat sheet works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
The whole lesson in one table. Notice how the two classification rows share the same clean p - y, and how every kink comes with a subgradient.
Elimination
Eliminate the wrong options
For softmax + cross-entropy, the gradient of the loss with respect to the pre-softmax logits z is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: In the chain rule dL/dz = Jᵀ(dL/dp), the softmax Jacobian's pₖ cancels cross-entropy's -yₖ/pₖ, leaving p - y. Verified: z = [1,2,3], true class 2 gives [0.0900, 0.2447, -0.3348].
Check
Differentiate with respect to the pre-softmax logits, not the probabilities.
Check your understanding
For softmax + cross-entropy, the gradient of the loss with respect to the pre-softmax logits z is:
Answer: A
Why: In the chain rule dL/dz = Jᵀ(dL/dp), the softmax Jacobian's pₖ cancels cross-entropy's -yₖ/pₖ, leaving p - y. Verified: z = [1,2,3], true class 2 gives [0.0900, 0.2447, -0.3348].
Prediction
Predict first
Why is sigmoid + MSE a poor loss for classification?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Its gradient (p - y)·p(1 - p) vanishes when the model is confidently wrong
Why: The chain rule adds a p(1 - p) factor. For a confident mistake (p ≈ 0, y = 1) that factor ≈ 0, so the gradient ≈ 0 - learning stalls exactly where it matters most (-0.017 vs cross-entropy's -0.982).
Check
Reason about the gradient magnitude on a confident mistake.
Check your understanding
Why is sigmoid + MSE a poor loss for classification?
Answer: A
Why: The chain rule adds a p(1 - p) factor. For a confident mistake (p ≈ 0, y = 1) that factor ≈ 0, so the gradient ≈ 0 - learning stalls exactly where it matters most (-0.017 vs cross-entropy's -0.982).
Prediction
Predict first
For hinge loss max(0, 1 - y·f), a correctly classified point with y·f = 1.5 contributes what gradient ∂L/∂f?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 0 - it is beyond the margin
Why: Since y·f = 1.5 > 1, the max selects 0 and the loss is flat there, so the gradient is 0. The SVM only pushes on points inside the margin (y·f < 1) - the support vectors.
Check
When does the SVM stop caring about a point?
Check your understanding
For hinge loss max(0, 1 - y·f), a correctly classified point with y·f = 1.5 contributes what gradient ∂L/∂f?
Answer: A
Why: Since y·f = 1.5 > 1, the max selects 0 and the loss is flat there, so the gradient is 0. The SVM only pushes on points inside the margin (y·f < 1) - the support vectors.
Elimination
Eliminate the wrong options
Where is MAE = (1/n)Σ|yᵢ - fᵢ| non-differentiable, and what do we use there?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: |yᵢ - fᵢ| has a corner at fᵢ = yᵢ: the slope jumps from -1/n to +1/n, so no single derivative exists. Optimizers use any subgradient in [-1/n, 1/n] (numpy's sign(0) = 0 picks the middle).
Check
Think about the shape of |r| right at r = 0.
Check your understanding
Where is MAE = (1/n)Σ|yᵢ - fᵢ| non-differentiable, and what do we use there?
Answer: A
Why: |yᵢ - fᵢ| has a corner at fᵢ = yᵢ: the slope jumps from -1/n to +1/n, so no single derivative exists. Optimizers use any subgradient in [-1/n, 1/n] (numpy's sign(0) = 0 picks the middle).
Section
The project
Concept
Implement MSE (value + gradient) and softmax+CE (gradient p - y) in pure NumPy, then verify your cross-entropy gradient against torch's autograd. You derived every piece - now assemble it yourself.
| # | requirement | tool |
|---|---|---|
| 1 | MSE value + gradient -2(y - f)/n | numpy |
| 2 | Stable softmax, then gradient p - y | np.exp(z - z.max()) |
| 3 | Verify dL/dz vs torch autograd | torch + np.allclose |
Build rules: type every line yourself, use a numerically stable softmax (subtract the max), and compare to autograd with np.allclose(..., atol=1e-6). When something errors, read the shapes - don't delete the error.
Worked example
Your turn: write the MSE value and gradient for y = [3,5,7], f = [2,6,4]. Predict the sign of grad[2] before you print (we under-predicted there).
Hint: value is np.mean((y - f)**2); gradient is -2*(y - f)/len(y). The big miss at i = 2 should get the largest-magnitude push.
import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
mse = np.mean((y - f)**2)
grad = -2*(y - f)/len(y)
print(round(mse, 4))
print(grad.round(4))| check | value |
|---|---|
| MSE | 3.6667 |
| gradient | [-0.6667, +0.6667, -2.0] |
| biggest push | grad[2] = -2.0 (worst-fit point) |
Worked example
Your turn: write a stable softmax, then return p - onehot for logits [1,2,3], true class 2. Say aloud which entry will be negative before you print.
Hint: e = np.exp(z - z.max()); p = e/e.sum(); the gradient is p - onehot. Only the true class (2) should come out negative.
import numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z = np.array([1., 2., 3.])
onehot = np.array([0., 0., 1.])
p = softmax(z)
print(p.round(4))
print((p - onehot).round(4))| class | p | p - y |
|---|---|---|
| 0 | 0.0900 | +0.0900 |
| 1 | 0.2447 | +0.2447 |
| 2 (true) | 0.6652 | -0.3348 |
Pattern
Step through it
Step through Milestone 2 - softmax + CE → p - y one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: confirm your p - y matches torch's autograd gradient of cross-entropy. Predict: exact match, or only close?
Hint: build z with requires_grad=True, call cross_entropy(z.unsqueeze(0), tensor([2])), .backward(), then compare z.grad with np.allclose(..., atol=1e-6).
import numpy as np, torch
def softmax(z):
e = np.exp(z - z.max()); return e/e.sum()
z = np.array([1., 2., 3.]); onehot = np.array([0., 0., 1.])
print((softmax(z) - onehot).round(4))
zt = torch.tensor([1., 2., 3.], requires_grad=True)
torch.nn.functional.cross_entropy(zt.unsqueeze(0), torch.tensor([2])).backward()
print(np.allclose(softmax(z) - onehot, zt.grad.numpy(), atol=1e-6))| source | dL/dz class 2 | match |
|---|---|---|
| your p - y | -0.3348 | - |
| torch autograd | -0.3348 | - |
| np.allclose | - | True |
Trade off
Comparison matrix
From Milestone 3 - verify against torch: every row here is a choice with a cost. Fill the dL/dz class 2 column, then say which row you would actually pick and what you give up for it.
| source | dL/dz class 2 | match |
|---|---|---|
| your p - y | -0.3348 | - |
| torch autograd | -0.3348 | - |
| np.allclose | - | True |
Concept
import numpy as np, torch
def mse(y, f): return np.mean((y - f)**2)
def mse_grad(y, f): return -2*(y - f)/len(y)
def softmax(z):
e = np.exp(z - z.max()); return e/e.sum()
y = np.array([3., 5., 7.]); f = np.array([2., 6., 4.])
print('MSE =', round(mse(y, f), 4), ' grad =', mse_grad(y, f).round(4))
z = np.array([1., 2., 3.]); onehot = np.array([0., 0., 1.])
ce_grad = softmax(z) - onehot # dL/dz = p - y
zt = torch.tensor([1., 2., 3.], requires_grad=True)
torch.nn.functional.cross_entropy(zt.unsqueeze(0), torch.tensor([2])).backward()
print('CE grad =', ce_grad.round(4))
print('matches torch:', np.allclose(ce_grad, zt.grad.numpy(), atol=1e-6))| printed line | value |
|---|---|
| MSE = | 3.6667 grad = [-0.6667 0.6667 -2.] |
| CE grad = | [0.09 0.2447 -0.3348] |
| matches torch: | True |
If your MSE reads 3.6667, your CE gradient reads [0.09, 0.2447, -0.3348], and matches torch prints True - you built the training signal of every classifier from the math up.
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| printed line | value |
|---|---|
| MSE = | 3.6667 grad = [-0.6667 0.6667 -2.] |
| CE grad = | [0.09 0.2447 -0.3348] |
| matches torch: | True |
Concept
Slides closed, out loud: explain (1) why softmax+CE collapses to p - y (which pₖ cancels what), (2) why MSE stalls on confident mistakes (name the vanishing factor), and (3) where MAE and hinge are non-differentiable and what you use there.
Stretch (homework): re-derive dL/dz = p - y for softmax+CE from the Jacobian by hand, then prove sigmoid+BCE gives the same. These exact gradients are the seed of Lesson 18's full backward pass.
Counterexample
Discussion prompt
Stretch (homework): re-derive dL/dz = p - y for softmax+CE from the Jacobian by hand, then prove sigmoid+BCE gives the same. These exact gradients are the seed of Lesson 18's full backward pass.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — What a loss even is · MSE - mean squared error · MAE - mean absolute error · Cross-entropy & softmax · The elegant dL/dz = p - y · Why MSE fails for classification. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
-2(y-f)/n and MAE -sign(y-f)/n, and place MAE's kink at f = ydL/dz = p - y via the softmax Jacobianp - y, and explain why MSE stalls a classifier via the p(1-p) factory·f = 1 and use subgradients, and match every result to torch autograd| move | the one thing to remember |
|---|---|
| softmax/sigmoid + CE | dL/dz = p - y (the Jacobian cancels the 1/p) |
| MSE for classification | gradient vanishes when confidently wrong: p(1-p) → 0 |
| MSE vs MAE | MSE grows with error; MAE is flat ±1/n and robust |
| kinks (MAE, hinge) | non-differentiable → pick a subgradient |
| choosing a loss | match the loss to the output so the gradient stays healthy |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.