Lesson 17: Loss Functions & Their Gradients

USAAIO Lesson 17, from Week 6, fully worked. It explains what a loss is, derives MSE and MAE term by term along with their gradients, and builds softmax and cross-entropy from scratch. It then proves the elegant result dL/dz = p - y TWICE, once by the softmax-Jacobian chain rule for the multi-class case and once directly for sigmoid with BCE, with no steps skipped. From there it explains why MSE stalls a classifier, through the vanishing p(1-p) factor, covers hinge loss and subgradients, gives a loss-to-gradient cheat sheet, and ends with a from-scratch build verified against torch autograd. One running example threads the whole deck - the regression y=[3,5,7] against f=[2,6,4], and the logits z=[1,2,3] with true class 2 - and every number was produced by real execution. The lesson runs to 60 slides.

Subject: Machine Learning · 113 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Loss Functions & Their Gradients

Title

USAAIO · Lesson 17 · Week 6

Every loss is a choice, and every gradient has a shape. We build MSE, MAE, cross-entropy, and hinge from the ground up, derive the elegant dL/dz = p - y twice, and prove why MSE quietly stalls a classifier - skipping no step and checking every number against torch.

2. By the end of this lesson you can

Objectives

  1. Say exactly what a loss is and why we differentiate it - the gradient is the training signal
  2. Derive MSE and MAE and their gradients term by term, and name where MAE is non-differentiable
  3. Build softmax and cross-entropy from scratch, numerically stably
  4. Derive dL/dz = p - y for softmax+CE via the softmax Jacobian, and independently for sigmoid+BCE
  5. Explain from the gradient shape why MSE is a poor classification loss, and place the hinge kink
  6. Implement every loss + gradient in NumPy and match torch autograd to 1e-6

3. What survived from Singular Value Decomposition?

Warm-up

Discussion prompt

Before we open Lesson 17: Loss Functions & Their Gradients: without looking back, what was the main idea of Singular Value Decomposition, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the SVD A=UΣVᵀ for every matrix, singular values as roots of eigenvalues of AᵀA, the rotate-scale-rotate geometry, low-rank approximation and Eckart-Young, and applications to PCA and the pseudoinverse. Build SVD compression and a pseudoinverse solver from scratch.

4. What a loss even is

Section

Part 1 of 7

5. A loss turns 'how wrong' into one number

Concept

A model outputs a prediction f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.

Training means making L small. And the only thing an optimizer knows how to do with L is follow its gradient downhill. So a loss is only useful if its gradient points somewhere helpful.

6. Break it if you can: A loss turns 'how wrong' into one number

Counterexample

Discussion prompt

A model outputs a prediction f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Training means making L small. And the only thing an optimizer knows how to do with L is follow its gradient downhill. So a loss is only useful if its gradient points somewhere helpful.

7. The gradient is the whole point

Intuition

You will never minimize a loss by hand. Gradient descent does it: nudge each parameter opposite the gradient, repeat. The loss value is a scoreboard; the loss gradient is the steering wheel.

So for every loss today we ask two questions: what is L (the number), and what is dL/d(prediction) (the direction). The second one is what actually trains the model.

That is why two losses that look similar can behave completely differently - a good-looking loss with a dead gradient in the wrong place will simply refuse to learn. We will watch exactly that happen to MSE.

8. Our running examples

Concept

One regression example and one classification example thread the whole deck, so every slide compounds. Regression: three targets and three model predictions.

itarget yprediction fresidual y - f
032+1
156-1
274+3

Classification: raw logits z = [1, 2, 3] over three classes, with the true class = 2. We will keep returning to these exact numbers.

9. Fill in: target y for Our running examples

Comparison

Comparison matrix

From Our running examples: refill the target y column from what you know. The rest of the table is as it appeared.

itarget yprediction fresidual y - f
032+1
156-1
274+3

10. MSE - mean squared error

Section

Part 2 of 7 - the regression staple

11. MSE: average the squared misses

Concept

Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.

\[ \text{MSE}(y, f) = \frac{1}{n}\sum_{i=1}^{n}(y_i - f_i)^2 = \frac{1}{n}\lVert y - f \rVert^2 \]

It is smooth everywhere, so calculus finds its minimum with no trouble - the property that made it the backbone of Lesson 7's least squares.

12. By analogy: MSE: average the squared misses

Analogy

Discussion prompt

Explain MSE: average the squared misses by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.

13. Guess the shape of the answer: MSE on our data, entry by entry

Estimation

Predict first

Take y = [3, 5, 7], f = [2, 6, 4]. Residuals are [+1, -1, +3]. Square, sum, divide by n = 3:

Commit before you compute: what does MSE on our data, entry by entry come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Sum then divide: (1 + 1 + 9) / 3 = 11/3

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The mean keeps the loss on the same scale regardless of dataset size n, so learning rates transfer across datasets.

14. MSE on our data, entry by entry

Worked example

Take y = [3, 5, 7], f = [2, 6, 4]. Residuals are [+1, -1, +3]. Square, sum, divide by n = 3:

Square each residual: 1² = 1, (-1)² = 1, 3² = 9

Why: Squaring kills the sign, so an overshoot and an undershoot of the same size cost the same. The big miss (+3) dominates: 9 out of a total of 11.

Sum then divide: (1 + 1 + 9) / 3 = 11/3

Why: The mean keeps the loss on the same scale regardless of dataset size n, so learning rates transfer across datasets.

\[ \text{MSE} = \frac{1 + 1 + 9}{3} = \frac{11}{3} \approx 3.6667 \]

iyᵢfᵢrᵢ = yᵢ - fᵢrᵢ²
032+11
156-11
274+39

15. What each one costs: MSE on our data, entry by entry

Trade off

Comparison matrix

From MSE on our data, entry by entry: every row here is a choice with a cost. Fill the rᵢ = yᵢ - fᵢ column, then say which row you would actually pick and what you give up for it.

iyᵢfᵢrᵢ = yᵢ - fᵢrᵢ²
032+11
156-11
274+39

16. What has to happen first: Differentiate MSE - one prediction first

Ranking

Put in order

Put the moves of Differentiate MSE - one prediction first into the order they have to happen.

  1. Let u = yᵢ - fᵢ, so the term is u²/n
  2. Inside derivative: d(yᵢ - fᵢ)/dfᵢ = -1
  3. Stack all i into a vector

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Chain rule: d(u²)/df = 2u · du/df.

17. Differentiate MSE - one prediction first

Worked example

The gradient is per-prediction: how does L change if we wiggle a single fᵢ? Isolate one term L = (yᵢ - fᵢ)²/n and differentiate with respect to fᵢ.

Let u = yᵢ - fᵢ, so the term is u²/n

Why: Chain rule: d(u²)/df = 2u · du/df. We differentiate the outer square, then the inside.

\[ \frac{\partial}{\partial f_i}\,\frac{(y_i - f_i)^2}{n} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot\frac{\partial(y_i - f_i)}{\partial f_i} \]

Inside derivative: d(yᵢ - fᵢ)/dfᵢ = -1

Why: yᵢ is a constant target; only -fᵢ depends on fᵢ. That minus sign is the whole reason the gradient has a leading minus.

\[ \frac{\partial L}{\partial f_i} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot(-1) = -\frac{2}{n}(y_i - f_i) \]

Stack all i into a vector

Why: Every coordinate has the same form, so the full gradient is one clean vector expression.

\[ \boxed{\;\nabla_f\,\text{MSE} = -\frac{2}{n}(y - f)\;} \]

18. Decode the notation: Differentiate MSE - one prediction first

Notation

Annotate

From Differentiate MSE - one prediction first — read this one piece at a time. What is each part doing?

On: \( \frac{\partial}{\partial f_i}\,\frac{(y_i - f_i)^2}{n} = \frac{1}{n}\cdot 2(y_i - f_i)\cdot\frac{\partial(y_i - f_i)}{\partial f_i} \)

  • Chain rule: d(u²)/df = 2u · du/df. We differentiate the outer square, then the inside.
  • yᵢ is a constant target; only -fᵢ depends on fᵢ. That minus sign is the whole reason the gradient has a leading minus.
  • Every coordinate has the same form, so the full gradient is one clean vector expression.

19. What has to be given first: MSE gradient, verified in code

Missing information

Discussion prompt

Compute the value and the gradient -2(y - f)/n directly. Self-contained - re-imports NumPy and re-defines y, f:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

MSE's gradient GROWS with the error, so the worst-fit point (i=2) gets the largest correction. Great for clean data, dangerous with outliers.

20. MSE gradient, verified in code

Worked example

Compute the value and the gradient -2(y - f)/n directly. Self-contained - re-imports NumPy and re-defines y, f:

import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mse = np.mean((y - f)**2)
grad = -2*(y - f)/n
print(round(mse, 4))
print(grad.round(4))

grad[2] = -2/3 · (7 - 4) = -2 is the biggest push

Why: MSE's gradient GROWS with the error, so the worst-fit point (i=2) gets the largest correction. Great for clean data, dangerous with outliers.

quantityformulavalue (verified)
MSE(1 + 1 + 9)/33.6667
grad[0]-2/3 · (+1)-0.6667
grad[1]-2/3 · (-1)+0.6667
grad[2]-2/3 · (+3)-2.0

21. Watch it run: MSE gradient, verified in code

Pattern

Step through it

Step through MSE gradient, verified in code one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: quantity is MSE
  2. Step 2: quantity is grad[0]
  3. Step 3: quantity is grad[1]
  4. Step 4: quantity is grad[2]

22. Why the sign points the right way

Intuition

At i = 2 we under-predicted (f = 4 < y = 7), so the residual is positive and the gradient is negative. Gradient descent steps opposite the gradient - so it raises f. Exactly what we want: predict higher.

At i = 1 we over-predicted (f = 6 > y = 5), the gradient is positive, descent lowers f. The sign of the residual, flipped by descent, always pushes the prediction toward the target.

23. Teach it back: Why the sign points the right way

Explain it

Discussion prompt

Explain Why the sign points the right way to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

At i = 1 we over-predicted (f = 6 > y = 5), the gradient is positive, descent lowers f. The sign of the residual, flipped by descent, always pushes the prediction toward the target.

24. The outlier problem, previewed

Intuition

Imagine point 2's target were a wild outlier - say 20 instead of 7, so the residual is +16. MSE squares it to 256, and its gradient becomes -2/3 · 16 = -10.67: one bad point hijacks the whole update.

MAE would give that same outlier a gradient of just -1/3 - the same as every other point. That capped, democratic gradient is why MAE is called robust, and it's the reason we bother with a second regression loss at all.

So the choice is a trade-off: MSE for clean data and smooth optimization, MAE when outliers would otherwise dominate. The gradient shape - growing vs flat - is the whole difference.

25. MAE - mean absolute error

Section

Part 3 of 7 - robust, but with a corner

26. MAE: average the absolute misses

Concept

Mean absolute error uses the absolute value instead of the square, so a miss of 3 costs 3, not 9 - big errors are not amplified.

\[ \text{MAE}(y, f) = \frac{1}{n}\sum_{i=1}^{n}\lvert y_i - f_i\rvert = \frac{1}{n}\lVert y - f\rVert_1 \]

That makes MAE robust to outliers - one wild point can't dominate the sum. The price is a gradient that behaves very differently, as we'll see.

27. Predict the next row: MAE on our data

Pattern

Predict first

The table runs: 0 | +1 | 1 | +1 · 1 | -1 | 1 | -1

In MAE on our data, given the rows so far: what is the next one — the row where i is 2?

Correct: 2 | +3 | 3 | +1

irᵢ = yᵢ - fᵢ|rᵢ|sign(rᵢ)
0+11+1
1-11-1
2+33+1

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Compare to MSE's numerator of 11: the big miss (+3) contributes 3 here versus 9 there.

28. MAE on our data

Worked example

Same residuals [+1, -1, +3], now take absolute values [1, 1, 3], average over n = 3:

|+1| + |-1| + |+3| = 1 + 1 + 3 = 5

Why: Compare to MSE's numerator of 11: the big miss (+3) contributes 3 here versus 9 there. MAE refuses to let one point take over.

\[ \text{MAE} = \frac{1 + 1 + 3}{3} = \frac{5}{3} \approx 1.6667 \]

irᵢ = yᵢ - fᵢ|rᵢ|sign(rᵢ)
0+11+1
1-11-1
2+33+1

29. Work backwards from the answer: MAE on our data

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

|+1| + |-1| + |+3| = 1 + 1 + 3 = 5

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Same residuals [+1, -1, +3], now take absolute values [1, 1, 3], average over n = 3:

30. Plan first: Differentiate MAE - the sign appears

Step zero

Discussion prompt

Differentiate MAE - the sign appears — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: d|u|/du = sign(u) for u ≠ 0

Answer:

  1. d|u|/du = sign(u) for u ≠ 0
  2. Fold in the -1
  3. Contrast with MSE: magnitude is FLAT

31. Differentiate MAE - the sign appears

Worked example

One term is |yᵢ - fᵢ|/n. The derivative of |u| is sign(u), then chain through u = yᵢ - fᵢ whose fᵢ-derivative is -1:

d|u|/du = sign(u) for u ≠ 0

Why: |u| has slope +1 when u > 0 and -1 when u < 0. That is exactly the sign function - as long as u is not 0.

\[ \frac{\partial}{\partial f_i}\,\frac{\lvert y_i - f_i\rvert}{n} = \frac{1}{n}\,\text{sign}(y_i - f_i)\cdot(-1) \]

Fold in the -1

Why: Same inside-derivative -1 as MSE. The result is constant magnitude 1/n, only its sign varies.

\[ \boxed{\;\nabla_f\,\text{MAE} = -\frac{1}{n}\,\text{sign}(y - f)\;} \]

Contrast with MSE: magnitude is FLAT

Why: MSE grad scales with the error; MAE grad is always ±1/n. A point that's off by 100 gets the same push as one off by 0.1 - robust, but slow to correct large errors.

32. Draw the shape of it: Differentiate MAE - the sign appears

Blank canvas

Draw it

Draw what Differentiate MAE - the sign appears just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

33. Restore the missing line: MAE gradient, verified in code

Fill the middle

Fill in the blanks

From MAE gradient, verified in code — one line has had its right-hand side removed. Put it back.

import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mae = np.mean(np.abs(y - f))
grad = -np.sign(y - f)/n
print(round(mae, 4))
print(grad.round(4))

Why: n is what everything below it consumes, so the wrong expression here fails later and somewhere else. Every gradient entry is ±1/3. The big miss at i=2 gets the SAME magnitude as the small misses - unlike MSE, where it was -2.0.

34. MAE gradient, verified in code

Worked example

Value and gradient -sign(y - f)/n. Self-contained:

import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
n = len(y)
mae = np.mean(np.abs(y - f))
grad = -np.sign(y - f)/n
print(round(mae, 4))
print(grad.round(4))

grad = [-0.333, +0.333, -0.333] - all the same size

Why: Every gradient entry is ±1/3. The big miss at i=2 gets the SAME magnitude as the small misses - unlike MSE, where it was -2.0.

isign(y - f)grad = -sign/n
0+1-0.3333
1-1+0.3333
2+1-0.3333

35. Something is wrong here: is MAE differentiable everywhere?

Anomaly

Predict first

A student writes this, and it looks reasonable:

MAE is a simple linear-looking function of the error, so its gradient is just ±1/n everywhere, including when the prediction is exactly right.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right).

MAE is non-differentiable at 0. There, use a subgradient - any value in the interval [-1/n, +1/n].

Why: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.

36. Trap: is MAE differentiable everywhere?

Trap

The trap

MAE is a simple linear-looking function of the error, so its gradient is just ±1/n everywhere, including when the prediction is exactly right.

Report sign(0) as some fixed value at the kink

Why: At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.

The fix

MAE is non-differentiable at 0. There, use a subgradient - any value in the interval [-1/n, +1/n].

Subgradient (usually 0) at the kink, sign elsewhere

Why: Optimizers pick a value in [-1/n, 1/n] (numpy's sign(0) = 0) and march on. This same subgradient idea handles the hinge kink and L1 (Lesson 12) - a corner is fine, you just choose a slope.

37. Break it on purpose: is MAE differentiable everywhere?

Break the constraint

Discussion prompt

The rule this trap just fixed:

Optimizers pick a value in [-1/n, 1/n] (numpy's sign(0) = 0) and march on. This same subgradient idea handles the hinge kink and L1 (Lesson 12) - a corner is fine, you just choose a slope.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

At yᵢ = fᵢ the absolute value has a CORNER: the slope jumps from -1/n (left) to +1/n (right). No single number is 'the' derivative there - the two-sided limits disagree.

38. Cross-entropy & softmax

Section

Part 4 of 7 - the classification loss

39. From logits to probabilities: softmax

Concept

A classifier outputs raw scores called logits z. Softmax turns them into a probability vector p - all positive, summing to 1 - by exponentiating and normalizing.

\[ p_k = \operatorname{softmax}(z)_k = \frac{e^{z_k}}{\sum_j e^{z_j}} \]

logit — A raw, unbounded real-valued score before the softmax/sigmoid. Positive logits become high probabilities, negative become low. The gradient dL/dz is taken with respect to THESE, not the probabilities.

40. The stability trick: subtract the max

Worked example

e^z overflows for large z. Softmax is unchanged if we subtract any constant c from every logit (it cancels top and bottom), so subtract c = max(z). For z = [1, 2, 3], c = 3:

\[ \frac{e^{z_k}}{\sum_j e^{z_j}} = \frac{e^{z_k - c}}{\sum_j e^{z_j - c}}\quad\text{(multiply top and bottom by } e^{-c}) \]

Shifted exponentials: e^(1-3), e^(2-3), e^(3-3)

Why: = e^-2, e^-1, e^0 = 0.1353, 0.3679, 1.0. The largest is now exactly 1, so nothing overflows - and the ratios are identical to the unshifted version.

kzₖzₖ - 3e^(zₖ - 3)
01-20.1353
12-10.3679
2301.0000

41. Watch it run: The stability trick: subtract the max

Pattern

Step through it

Step through The stability trick: subtract the max one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: k is 0
  2. Step 2: k is 1
  3. Step 3: k is 2

42. Finish it with less help: Finish softmax: normalize

Faded example

Fill in the blanks

Finish softmax: normalize, with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
def softmax(z):
e = np.exp(z - z.max())
return e / e.sum()
z = np.array([1., 2., 3.])
p = softmax(z)
print(p.round(4))
onehot = np.array([0., 0., 1.])
print((p - onehot).round(4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. Class 2 (the largest logit) gets the largest probability, 0.665.

43. Finish softmax: normalize

Worked example

Divide each shifted exponential by their sum 0.1353 + 0.3679 + 1.0 = 1.5032:

import numpy as np
def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()
z = np.array([1., 2., 3.])
p = softmax(z)
print(p.round(4))
onehot = np.array([0., 0., 1.])
print((p - onehot).round(4))

p = [0.0900, 0.2447, 0.6652], sums to 1

Why: Class 2 (the largest logit) gets the largest probability, 0.665. The model is only moderately confident - which is exactly why there's a real gradient to learn from.

ke^(zₖ-3)pₖ = e / 1.5032
00.13530.0900
10.36790.2447
21.00000.6652

44. Restore the missing line: Prove the max-subtraction is exact

Fill the middle

Fill in the blanks

From Prove the max-subtraction is exact — one line has had its right-hand side removed. Put it back.

import numpy as np
def softmax(z):
e = np.exp(z - z.max()); return e / e.sum()
def naive(z):
e = np.exp(z); return e / e.sum()
z = np.array([1., 2., 3.])
print(softmax(z).round(6))
print(naive(z).round(6))
print(np.allclose(softmax(z), naive(z)))

Why: z is what everything below it consumes, so the wrong expression here fails later and somewhere else. Subtracting max(z) multiplies numerator and denominator by the same e^{-c}, which cancels - so the probabilities are identical, just computed without overflow risk.

45. Prove the max-subtraction is exact

Worked example

The stability trick must not change the answer. Compare the stable softmax (subtract the max) against the naive e^z / Σe^z on our logits - they should agree to machine precision. Self-contained:

import numpy as np
def softmax(z):
    e = np.exp(z - z.max()); return e / e.sum()
def naive(z):
    e = np.exp(z); return e / e.sum()
z = np.array([1., 2., 3.])
print(softmax(z).round(6))
print(naive(z).round(6))
print(np.allclose(softmax(z), naive(z)))

Both give [0.090031, 0.244728, 0.665241]; allclose = True

Why: Subtracting max(z) multiplies numerator and denominator by the same e^{-c}, which cancels - so the probabilities are identical, just computed without overflow risk. Always subtract the max in real code.

methodp₀p₁p₂
stable (z - max)0.0900310.2447280.665241
naive0.0900310.2447280.665241
np.allclose--True

46. Cross-entropy: penalize the true class's probability

Concept

Cross-entropy looks at the probability the model assigned to the true class and charges -log of it. Confident and right → near 0; confident and wrong → huge.

\[ \text{CE} = -\sum_{k} y_k \log p_k \;\overset{\text{one-hot } y}{=}\; -\log p_{\text{true}} \]

With a one-hot y, every term is zero except the true class, so the sum collapses to a single -log p_true. The binary special case is BCE: -y log p - (1-y) log(1-p).

47. Guess the shape of the answer: CE value on our logits

Estimation

Predict first

True class is 2, and p₂ = 0.6652. Cross-entropy is just -log(0.6652):

Commit before you compute: what does CE value on our logits come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Sanity: if p_true were 1.0, CE = -log 1 = 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. A perfect confident prediction has zero loss; as p_true → 0 the loss → +infinity.

48. CE value on our logits

Worked example

True class is 2, and p₂ = 0.6652. Cross-entropy is just -log(0.6652):

CE = -log(0.6652) = 0.4076

Why: The other two probabilities (0.09, 0.2447) never enter the loss VALUE - only p_true does. They will, however, enter the gradient.

\[ \text{CE} = -\log(0.6652) \approx 0.4076 \]

Sanity: if p_true were 1.0, CE = -log 1 = 0

Why: A perfect confident prediction has zero loss; as p_true → 0 the loss → +infinity. That steepness near 0 is what makes CE punish confident mistakes hard.

quantityvalue (verified)
p_true = p₂0.6652
-log(p₂)0.4076
torch cross_entropy0.4076

49. The elegant dL/dz = p - y

Section

Part 5 of 7 - the derivation, every step

50. Differentiate w.r.t. logits, not probabilities

Concept

The gradient dL/dp of cross-entropy alone is messy (it has a 1/p). The magic happens when we go one step further back and differentiate w.r.t. the logits z - the softmax Jacobian cancels that 1/p.

\[ \frac{\partial L}{\partial z_j} = \sum_k \frac{\partial L}{\partial p_k}\,\frac{\partial p_k}{\partial z_j} \quad\text{(chain rule through the softmax)} \]

We need two pieces: the loss-to-probability gradient dL/dp, and the softmax Jacobian dp/dz. Then we multiply and watch everything collapse.

51. Why go all the way back to the logits

Intuition

The network computes logits z → softmax → p → loss. Backprop hands each layer the gradient w.r.t. its output. The softmax layer's job is to convert dL/dp (messy, has a 1/p) into dL/dz (clean) before passing it on.

So stopping at dL/dp is the wrong altitude - it's an intermediate quantity. The number that actually flows into the weights is dL/dz. We chase it back one link, through the softmax, and the 1/p disappears.

52. Plan first: Piece 1: dL/dp for cross-entropy

Step zero

Discussion prompt

Piece 1: dL/dp for cross-entropy — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: L = -Σ y_k log p_k

Answer:

  1. L = -Σ y_k log p_k
  2. Differentiate w.r.t. a single p_k: d(-y_k log p_k)/dp_k = -y_k / p_k
  3. One-hot y: only the true class survives

53. Piece 1: dL/dp for cross-entropy

Worked example

L = -Σ y_k log p_k

Why: Cross-entropy written over all classes, before assuming one-hot.

\[ L = -\sum_k y_k \log p_k \]

Differentiate w.r.t. a single p_k: d(-y_k log p_k)/dp_k = -y_k / p_k

Why: Only the k-th term depends on p_k; d(log p)/dp = 1/p. So the loss-to-probability gradient is -y_k / p_k.

\[ \frac{\partial L}{\partial p_k} = -\frac{y_k}{p_k} \]

One-hot y: only the true class survives

Why: y_k = 0 except at the true class, so dL/dp is a single nonzero entry -1/p_true. For our data that's -1/0.6652 = -1.5032.

54. What has to happen first: Piece 2: the softmax Jacobian

Ranking

Put in order

Put the moves of Piece 2: the softmax Jacobian into the order they have to happen.

  1. Diagonal k = j: ∂pₖ/∂zₖ = pₖ(1 - pₖ)
  2. Off-diagonal k ≠ j: ∂pₖ/∂z_j = -pₖ p_j
  3. Both cases in one line

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The numerator e^{zₖ} depends on zₖ too, so you get pₖ minus pₖ² from the quotient rule.

55. Piece 2: the softmax Jacobian

Worked example

Differentiate pₖ = e^{zₖ}/Σⱼe^{zⱼ} w.r.t. z_j. Quotient rule splits into the diagonal case (k = j) and off-diagonal (k ≠ j):

Diagonal k = j: ∂pₖ/∂zₖ = pₖ(1 - pₖ)

Why: The numerator e^{zₖ} depends on zₖ too, so you get pₖ minus pₖ² from the quotient rule.

\[ \frac{\partial p_k}{\partial z_k} = p_k(1 - p_k) \]

Off-diagonal k ≠ j: ∂pₖ/∂z_j = -pₖ p_j

Why: Here only the denominator sees z_j, giving a negative coupling. Raising one logit steals probability from the others.

\[ \frac{\partial p_k}{\partial z_j} = -p_k p_j \]

Both cases in one line

Why: δ is the Kronecker delta (1 if k=j else 0). This single formula is the entire softmax Jacobian.

\[ \frac{\partial p_k}{\partial z_j} = p_k(\delta_{kj} - p_j) \]

56. Say it in words: Piece 2: the softmax Jacobian

Translation

\( \frac{\partial p_k}{\partial z_j} = p_k(\delta_{kj} - p_j) \)

Draw it

Translate both ways. First write the expression above as a sentence with no symbols in it at all. Then cover it, and write your sentence back as notation. If the two versions disagree, the disagreement is the thing to fix.

57. Plan first: Multiply the pieces - the cancellation

Step zero

Discussion prompt

Multiply the pieces - the cancellation — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Chain: ∂L/∂z_j = Σ_k (-y_k/p_k) · p_k(δ_kj - p_j)

Answer:

  1. Chain: ∂L/∂z_j = Σ_k (-y_k/p_k) · p_k(δ_kj - p_j)
  2. The pₖ cancels
  3. Use Σy_k = 1 and Σy_k δ_kj = y_j

58. Multiply the pieces - the cancellation

Worked example

Chain: ∂L/∂z_j = Σ_k (-y_k/p_k) · p_k(δ_kj - p_j)

Why: Substitute Piece 1 and Piece 2 into the chain-rule sum over k.

\[ \frac{\partial L}{\partial z_j} = \sum_k \Big(-\frac{y_k}{p_k}\Big)\,p_k(\delta_{kj} - p_j) \]

The pₖ cancels

Why: (-y_k/p_k)·p_k = -y_k. The messy 1/p is gone - this is the cancellation that makes the result clean.

\[ = -\sum_k y_k(\delta_{kj} - p_j) = -\sum_k y_k \delta_{kj} + p_j\sum_k y_k \]

Use Σy_k = 1 and Σy_k δ_kj = y_j

Why: Probabilities of the one-hot target sum to 1, and the delta picks out the j-th entry. So the first term is -y_j and the second is p_j·1.

\[ \frac{\partial L}{\partial z_j} = p_j - y_j \;\Longrightarrow\; \boxed{\;\frac{\partial L}{\partial z} = p - y\;} \]

59. Why p - y is so satisfying

Intuition

The gradient of the loss w.r.t. the logits is literally the prediction error: how far each predicted probability sits from its target. No 1/p, no Jacobian left over - just p - y.

The true class has p - y < 0, so descent raises its logit; every wrong class has p - y > 0, so descent lowers theirs. The size of each push is exactly how wrong that class was. That is why classifiers train cleanly.

60. Finish it with less help: Verify p - y numerically (Jacobian route)

Faded example

Fill in the blanks

Verify p - y numerically (Jacobian route), with the scaffolding fading: two lines are gone now — fill both.

import numpy as np
z = np.array([1., 2., 3.])
e = np.exp(z - z.max())
p = e / e.sum()
J = np.diag(p) - np.outer(p, p)
onehot = np.array([0., 0., 1.])
dLdp = -onehot / p
dLdz = J.T @ dLdp
print(dLdz.round(4))
print((p - onehot).round(4))

Why: Reproducing these unaided, rather than reading them, is what tells you the method has transferred. The full chain-rule product equals p - y exactly, confirming the algebra.

61. Verify p - y numerically (Jacobian route)

Worked example

Build the Jacobian J = diag(p) - p pᵀ, form dL/dp = -y/p, and confirm Jᵀ (dL/dp) = p - y. Self-contained:

import numpy as np
z = np.array([1., 2., 3.])
e = np.exp(z - z.max())
p = e / e.sum()
J = np.diag(p) - np.outer(p, p)
onehot = np.array([0., 0., 1.])
dLdp = -onehot / p
dLdz = J.T @ dLdp
print(dLdz.round(4))
print((p - onehot).round(4))

Jᵀ (dL/dp) = [0.0900, 0.2447, -0.3348] = p - onehot

Why: The full chain-rule product equals p - y exactly, confirming the algebra. The one-hot has its 1 at class 2, so only that entry is negative.

kvia Jacobianp - ymatch
00.09000.0900yes
10.24470.2447yes
2-0.3348-0.3348yes

62. What stays fixed: Verify p - y numerically (Jacobian route)

Invariant

Step through it

Step through Verify p - y numerically (Jacobian route) one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: k is 0
  2. Step 2: k is 1
  3. Step 3: k is 2

63. Verify against torch autograd

Worked example

The ultimate check: does real backprop agree? Build z with requires_grad, run cross_entropy, and read z.grad. Self-contained:

import torch, numpy as np
z = torch.tensor([1., 2., 3.], requires_grad=True)
loss = torch.nn.functional.cross_entropy(z.unsqueeze(0), torch.tensor([2]))
loss.backward()
p = torch.softmax(z, 0).detach().numpy()
print(p.round(4))
print(z.grad.numpy().round(4))
print((p - np.array([0., 0., 1.])).round(4))

z.grad = [0.0900, 0.2447, -0.3348] = p - onehot

Why: torch's autograd produces exactly p - y. The true class (2) gets the negative push; the impostors get positive pushes proportional to their probability.

classpy (one-hot)z.grad = p - y
00.09000+0.0900
10.24470+0.2447
2 (true)0.66521-0.3348

64. Fill in: y (one-hot) for Verify against torch autograd

Comparison

Comparison matrix

From Verify against torch autograd: refill the y (one-hot) column from what you know. The rest of the table is as it appeared.

classpy (one-hot)z.grad = p - y
00.09000+0.0900
10.24470+0.2447
2 (true)0.66521-0.3348

65. Restore the missing line: The binary case: sigmoid + BCE gives p - y too

Fill the middle

Fill in the blanks

From The binary case: sigmoid + BCE gives p - y too — one line has had its right-hand side removed. Put it back.

import torch
z = torch.tensor([0.5], requires_grad=True)
loss = torch.nn.functional.binary_cross_entropy_with_logits(z, torch.tensor([1.]))
loss.backward()
p = torch.sigmoid(z).item()
print(round(p, 4), round(z.grad.item(), 4), round(p - 1, 4))

Why: p is what everything below it consumes, so the wrong expression here fails later and somewhere else. Same clean form. The 0.3775 gap is exactly how far the prediction (0.6225) sits below the target (1.0).

66. The binary case: sigmoid + BCE gives p - y too

Worked example

For a single logit, softmax becomes the sigmoid σ(z) = 1/(1 + e^{-z}) and CE becomes BCE. The identity dL/dz = p - y still holds. Take z = 0.5, label y = 1:

import torch
z = torch.tensor([0.5], requires_grad=True)
loss = torch.nn.functional.binary_cross_entropy_with_logits(z, torch.tensor([1.]))
loss.backward()
p = torch.sigmoid(z).item()
print(round(p, 4), round(z.grad.item(), 4), round(p - 1, 4))

p = 0.6225, dL/dz = -0.3775 = p - 1

Why: Same clean form. The 0.3775 gap is exactly how far the prediction (0.6225) sits below the target (1.0).

quantityvalue (verified)
p = σ(0.5)0.6225
dL/dz (autograd)-0.3775
p - y-0.3775

67. Plan first: Prove sigmoid + BCE = p - y by hand

Step zero

Discussion prompt

Prove sigmoid + BCE = p - y by hand — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: L = -y log p - (1-y) log(1-p), p = σ(z)

Answer:

  1. L = -y log p - (1-y) log(1-p), p = σ(z)
  2. Sigmoid derivative: dp/dz = p(1 - p)
  3. Distribute p(1-p): the denominators cancel
  4. Check: p - y = 0.6225 - 1 = -0.3775

68. Prove sigmoid + BCE = p - y by hand

Worked example

L = -y log p - (1-y) log(1-p), p = σ(z)

Why: BCE in terms of the logit z through the sigmoid p.

\[ \frac{\partial L}{\partial p} = -\frac{y}{p} + \frac{1-y}{1-p} \]

Sigmoid derivative: dp/dz = p(1 - p)

Why: The 1-D version of the softmax diagonal. This is the factor that will cancel the denominators.

\[ \frac{\partial L}{\partial z} = \Big(-\frac{y}{p} + \frac{1-y}{1-p}\Big)\,p(1-p) \]

Distribute p(1-p): the denominators cancel

Why: -y/p · p(1-p) = -y(1-p); (1-y)/(1-p) · p(1-p) = (1-y)p. Add them.

\[ = -y(1-p) + (1-y)p = -y + yp + p - yp = p - y \]

Check: p - y = 0.6225 - 1 = -0.3775

Why: Matches the autograd value exactly. Both softmax+CE and sigmoid+BCE collapse to p - y - it is not a coincidence, it's the structure of the exponential-family loss paired with its natural link.

69. Draw the shape of it: Prove sigmoid + BCE = p - y by hand

Blank canvas

Draw it

Draw what Prove sigmoid + BCE = p - y by hand just did — the shape of it, not the line-by-line working. One picture, labels only where you need them. Then check it against the steps: anything you could not draw is a step you followed rather than understood.

70. Predict the next row: Sigmoid IS the two-class softmax

Pattern

Predict first

The table runs: -2.0 | 0.119203 | 0.119203 · 0.0 | 0.500000 | 0.500000 · 0.5 | 0.622459 | 0.622459

In Sigmoid IS the two-class softmax, given the rows so far: what is the next one — the row where z is 3.0?

Correct: 3.0 | 0.952574 | 0.952574

zsoftmax([0,z])[1]sigmoid(z)
-2.00.1192030.119203
0.00.5000000.500000
0.50.6224590.622459
3.00.9525740.952574

Why: The relationship between the columns, not the individual numbers, is what generates the next row. The two are algebraically identical, not merely close.

71. Sigmoid IS the two-class softmax

Worked example

The binary and multi-class cases are the same object. Softmax over logits [0, z] gives class-1 probability e^z/(1 + e^z) = σ(z). So sigmoid+BCE is just softmax+CE with two classes - which is why they share dL/dz = p - y. Check numerically:

import numpy as np
def softmax(z):
    e = np.exp(z - z.max()); return e / e.sum()
def sigmoid(z): return 1/(1 + np.exp(-z))
for z in [-2., 0., 0.5, 3.]:
    two_class = softmax(np.array([0., z]))[1]
    print(z, round(two_class, 6), round(sigmoid(z), 6))

softmax([0, z])[1] equals sigmoid(z) for every z

Why: The two are algebraically identical, not merely close. That is the deep reason the p - y result is one result, not two coincidences.

zsoftmax([0,z])[1]sigmoid(z)
-2.00.1192030.119203
0.00.5000000.500000
0.50.6224590.622459
3.00.9525740.952574

72. What happens as it grows: Sigmoid IS the two-class softmax

Scale up

Step through it

Step through Sigmoid IS the two-class softmax and watch the numbers move. Now imagine the input ten times bigger: which column is the one that stops this being practical?

  1. Step 1: z is -2.0
  2. Step 2: z is 0.0
  3. Step 3: z is 0.5
  4. Step 4: z is 3.0

73. Why MSE fails for classification

Section

Part 6 of 7 - the vanishing gradient

74. Pair sigmoid with MSE and a factor sneaks in

Concept

Suppose you (wrongly) train a classifier with sigmoid output and MSE loss L = ½(p - y)². The chain rule to the logit picks up the sigmoid derivative σ'(z) = p(1 - p):

\[ \frac{\partial L}{\partial z} = (p - y)\cdot\underbrace{p(1-p)}_{\sigma'(z)} \]

Compare with cross-entropy's clean p - y. MSE carries an extra p(1-p) multiplier - and that multiplier is exactly what breaks training.

75. The multiplier dies when you're confidently wrong

Intuition

p(1-p) is largest at p = 0.5 (value 0.25) and shrinks to 0 as p approaches 0 or 1. So when the model is confident - p ≈ 0 or p ≈ 1 - the multiplier is nearly zero.

Now picture a confident mistake: true label y = 1 but the model says p ≈ 0. Cross-entropy screams (p - y ≈ -1). MSE multiplies that by p(1-p) ≈ 0 and whispers - almost no gradient, right when you need the most.

76. Teach it back: The multiplier dies when you're confidently wrong

Explain it

Discussion prompt

Explain The multiplier dies when you're confidently wrong to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

p(1-p) is largest at p = 0.5 (value 0.25) and shrinks to 0 as p approaches 0 or 1. So when the model is confident - p ≈ 0 or p ≈ 1 - the multiplier is nearly zero.

77. What has to be given first: MSE vs CE on a confident mistake

Missing information

Discussion prompt

True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

CE keeps pushing hard on the confident mistake; MSE's gradient is 57x smaller because p(1-p) = 0.0177 crushed it. An MSE classifier gets STUCK on its worst errors.

78. MSE vs CE on a confident mistake

Worked example

True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:

import numpy as np
def sigmoid(z): return 1/(1 + np.exp(-z))
for z in [-4., 0.]:
    p = sigmoid(z)
    sg = p*(1 - p)
    ce = p - 1            # cross-entropy dL/dz
    mse = (p - 1)*sg      # MSE-through-sigmoid dL/dz
    print(z, round(p, 4), round(ce, 4), round(mse, 6))

At z = -4: CE grad -0.982 (strong) vs MSE grad -0.017 (≈ 0)

Why: CE keeps pushing hard on the confident mistake; MSE's gradient is 57x smaller because p(1-p) = 0.0177 crushed it. An MSE classifier gets STUCK on its worst errors.

logit zpp(1-p)CE dL/dzMSE dL/dz
-4.00.01800.0177-0.9820-0.017345
0.00.50000.2500-0.5000-0.125000

79. Work backwards from the answer: MSE vs CE on a confident mistake

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

At z = -4: CE grad -0.982 (strong) vs MSE grad -0.017 (≈ 0)

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

True label y = 1, but the logit is strongly negative (z = -4, so p ≈ 0.018). Compare the gradient each loss delivers. Self-contained:

80. Something is wrong here: using MSE to train a classifier

Anomaly

Predict first

A student writes this, and it looks reasonable:

A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix.

Match the loss to the output: cross-entropy for classification, MSE for regression.

Why: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix. Verified: gradient -0.017 vs CE's -0.982 at a confident error.

81. Trap: using MSE to train a classifier

Trap

The trap

A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.

Train with sigmoid + MSE, watch it stall

Why: The p(1-p) factor kills the gradient on confident mistakes, so learning stalls precisely on the hard examples the model most needs to fix. Verified: gradient -0.017 vs CE's -0.982 at a confident error.

The fix

Match the loss to the output: cross-entropy for classification, MSE for regression.

Use softmax/sigmoid + cross-entropy

Why: Then dL/dz = p - y - a full-strength gradient on every error, with no vanishing multiplier. The loss is deliberately chosen so its gradient stays healthy everywhere.

82. Which of these survive contact with Lesson 17: Loss Functions & Their Gradients?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A model outputs a prediction f; the truth is y. A loss L(y, f) collapses the mismatch into a single non-negative number - 0 means perfect, bigger means worse.; One regression example and one classification example thread the whole deck, so every slide compounds. Regression: three targets and three model predictions.; Mean squared error takes each residual yᵢ - fᵢ, squares it, and averages. Squaring makes every miss positive and punishes big misses hardest.
Breaks
MAE is a simple linear-looking function of the error, so its gradient is just ±1/n everywhere, including when the prediction is exactly right.; A loss is a loss - MSE minimizes squared error, so just apply sigmoid + MSE to the predicted probabilities for classification too.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 17: Loss Functions & Their Gradients puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

83. Hinge loss (SVM)

Section

Part 7 of 7 - margins & subgradients

84. Hinge: only punish inside the margin

Concept

The SVM loss uses labels y ∈ {-1, +1} and a raw score f(x). It penalizes a point only until it is correctly classified with margin (y·f ≥ 1); beyond that, zero.

\[ L = \max\!\big(0,\; 1 - y\,f(x)\big) \]

The quantity y·f is the signed margin: positive when the sign of f matches the label, and its size says how confidently. Hinge cares only when that margin is below 1.

85. By analogy: Hinge: only punish inside the margin

Analogy

Discussion prompt

Explain Hinge: only punish inside the margin by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The SVM loss uses labels y ∈ {-1, +1} and a raw score f(x). It penalizes a point only until it is correctly classified with margin (y·f ≥ 1); beyond that, zero.

86. What has to happen first: Hinge gradient - three regimes

Ranking

Put in order

Put the moves of Hinge gradient - three regimes into the order they have to happen.

  1. If y·f > 1: L = 0, gradient 0
  2. If y·f < 1: L = 1 - y·f, gradient -y
  3. At y·f = 1 exactly: subgradient in [-y, 0]

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The point is safely beyond the margin - already correct with room to spare - so it contributes no push at all.

87. Hinge gradient - three regimes

Worked example

Differentiate max(0, 1 - y·f) w.r.t. f. Two smooth pieces meet at the kink y·f = 1:

If y·f > 1: L = 0, gradient 0

Why: The point is safely beyond the margin - already correct with room to spare - so it contributes no push at all.

\[ \frac{\partial L}{\partial f} = \begin{cases} 0 & y f > 1 \\ -y & y f < 1 \end{cases} \]

If y·f < 1: L = 1 - y·f, gradient -y

Why: Inside the margin the loss is linear with slope -y in f. Descent (opposite -y) moves f in the +y direction, pushing the point toward the correct side.

At y·f = 1 exactly: subgradient in [-y, 0]

Why: Same corner story as MAE - the two slopes (0 and -y) disagree at the kink, so any value between them is a valid subgradient. Optimizers just pick one.

88. Decode the notation: Hinge gradient - three regimes

Notation

Annotate

From Hinge gradient - three regimes — read this one piece at a time. What is each part doing?

On: \( \frac{\partial L}{\partial f} = \begin{cases} 0 & y f > 1 \\ -y & y f < 1 \end{cases} \)

  • The point is safely beyond the margin - already correct with room to spare - so it contributes no push at all.
  • Inside the margin the loss is linear with slope -y in f. Descent (opposite -y) moves f in the +y direction, pushing the point toward the correct side.
  • Same corner story as MAE - the two slopes (0 and -y) disagree at the kink, so any value between them is a valid subgradient. Optimizers just pick one.

89. Guess the shape of the answer: Hinge on three points

Estimation

Predict first

Evaluate loss and gradient at three points. Self-contained:

Commit before you compute: what does Hinge on three points come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: (y=+1, f=1.5): y·f = 1.5 > 1 → loss 0, grad 0

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Correct and beyond the margin. The SVM ignores it entirely - only support vectors near the boundary matter.

90. Hinge on three points

Worked example

Evaluate loss and gradient at three points. Self-contained:

import numpy as np
def hinge(y, f): return max(0., 1 - y*f)
def hinge_grad(y, f): return 0. if y*f > 1 else -y
for y, f in [(1, 1.5), (1, 0.3), (-1, 0.3)]:
    print(y, f, round(y*f, 2), round(hinge(y, f), 2), hinge_grad(y, f))

(y=+1, f=1.5): y·f = 1.5 > 1 → loss 0, grad 0

Why: Correct and beyond the margin. The SVM ignores it entirely - only support vectors near the boundary matter.

yfy·flossgradient ∂L/∂f
+11.5+1.50.000
+10.3+0.30.70-1
-10.3-0.31.30+1

91. Loss → gradient cheat sheet

Pattern

The whole lesson in one table. Notice how the two classification rows share the same clean p - y, and how every kink comes with a subgradient.

lossvaluegradientuse / note
MSE‖y - f‖²/n-2(y - f)/nregression; grows with error, outlier-sensitive
MAE‖y - f‖₁/n-sign(y - f)/nregression; robust; subgradient at 0
softmax + CE-log p_truep - y (w.r.t. logits)multi-class; clean, full-strength gradient
sigmoid + BCE-log p_truep - y (w.r.t. logit)binary; same clean form
hingemax(0, 1 - yf)0 if yf>1 else -ySVM margin; subgradient at yf=1

92. Where does it stop working: Loss → gradient cheat sheet

Edge cases

Discussion prompt

Loss → gradient cheat sheet works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

The whole lesson in one table. Notice how the two classification rows share the same clean p - y, and how every kink comes with a subgradient.

93. Rule out three: Check yourself - the p - y result

Elimination

Eliminate the wrong options

For softmax + cross-entropy, the gradient of the loss with respect to the pre-softmax logits z is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. p - y (predicted probabilities minus the one-hot target)
  • B. y - p
  • C. p(1 - p)
  • D. -log p

Survives elimination: A

Why: In the chain rule dL/dz = Jᵀ(dL/dp), the softmax Jacobian's pₖ cancels cross-entropy's -yₖ/pₖ, leaving p - y. Verified: z = [1,2,3], true class 2 gives [0.0900, 0.2447, -0.3348].

94. Check yourself - the p - y result

Check

Differentiate with respect to the pre-softmax logits, not the probabilities.

Check your understanding

For softmax + cross-entropy, the gradient of the loss with respect to the pre-softmax logits z is:

  • A. p - y (predicted probabilities minus the one-hot target) (correct)
  • B. y - p
  • C. p(1 - p)
  • D. -log p

Answer: A

Why: In the chain rule dL/dz = Jᵀ(dL/dp), the softmax Jacobian's pₖ cancels cross-entropy's -yₖ/pₖ, leaving p - y. Verified: z = [1,2,3], true class 2 gives [0.0900, 0.2447, -0.3348].

Why B tempts people
Sign flipped. With L = -Σ y log p the gradient is p - y; y - p would point uphill and make gradient descent ASCEND the loss.
Why C tempts people
p(1 - p) is the sigmoid/softmax-diagonal DERIVATIVE. It appears (uncancelled) when you wrongly pair sigmoid with MSE - it is not in the clean CE result.
Why D tempts people
-log p is the loss VALUE for the true class, not its gradient with respect to the logits.

95. Answer it before you see the options: Check yourself - MSE for classification

Prediction

Predict first

Why is sigmoid + MSE a poor loss for classification?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Its gradient (p - y)·p(1 - p) vanishes when the model is confidently wrong

Why: The chain rule adds a p(1 - p) factor. For a confident mistake (p ≈ 0, y = 1) that factor ≈ 0, so the gradient ≈ 0 - learning stalls exactly where it matters most (-0.017 vs cross-entropy's -0.982).

96. Check yourself - MSE for classification

Check

Reason about the gradient magnitude on a confident mistake.

Check your understanding

Why is sigmoid + MSE a poor loss for classification?

  • A. Its gradient (p - y)·p(1 - p) vanishes when the model is confidently wrong (correct)
  • B. MSE cannot be computed on probabilities
  • C. MSE has no global minimum
  • D. MSE always overfits the training data

Answer: A

Why: The chain rule adds a p(1 - p) factor. For a confident mistake (p ≈ 0, y = 1) that factor ≈ 0, so the gradient ≈ 0 - learning stalls exactly where it matters most (-0.017 vs cross-entropy's -0.982).

Why B tempts people
MSE is perfectly computable on probabilities - the issue is the gradient SHAPE (a vanishing multiplier), not whether the value exists.
Why C tempts people
MSE is convex in the prediction and does have a minimum; the failure is a vanishing gradient en route, not a missing minimum.
Why D tempts people
Overfitting is unrelated - the failure is an optimization one (no learning signal on hard cases), not a generalization one.

97. Answer it before you see the options: Check yourself - the hinge margin

Prediction

Predict first

For hinge loss max(0, 1 - y·f), a correctly classified point with y·f = 1.5 contributes what gradient ∂L/∂f?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 0 - it is beyond the margin

Why: Since y·f = 1.5 > 1, the max selects 0 and the loss is flat there, so the gradient is 0. The SVM only pushes on points inside the margin (y·f < 1) - the support vectors.

98. Check yourself - the hinge margin

Check

When does the SVM stop caring about a point?

Check your understanding

For hinge loss max(0, 1 - y·f), a correctly classified point with y·f = 1.5 contributes what gradient ∂L/∂f?

  • A. 0 - it is beyond the margin (correct)
  • B. -y
  • C. +y
  • D. 1.5

Answer: A

Why: Since y·f = 1.5 > 1, the max selects 0 and the loss is flat there, so the gradient is 0. The SVM only pushes on points inside the margin (y·f < 1) - the support vectors.

Why B tempts people
-y is the gradient INSIDE the margin (y·f < 1). This point is safely outside, so it contributes nothing.
Why C tempts people
+y is never the hinge gradient in f; the active-region gradient is -y, and here it is 0 anyway.
Why D tempts people
1.5 is the signed margin value y·f, not a gradient. The gradient is 0 because the loss is flat beyond the margin.

99. Rule out three: Check yourself - MAE differentiability

Elimination

Eliminate the wrong options

Where is MAE = (1/n)Σ|yᵢ - fᵢ| non-differentiable, and what do we use there?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. At any point where fᵢ = yᵢ; use a subgradient in [-1/n, 1/n]
  • B. Nowhere - MAE is differentiable everywhere with gradient ±1/n
  • C. At its minimum only, where we use the second derivative
  • D. Everywhere - MAE has no usable gradient at all

Survives elimination: A

Why: |yᵢ - fᵢ| has a corner at fᵢ = yᵢ: the slope jumps from -1/n to +1/n, so no single derivative exists. Optimizers use any subgradient in [-1/n, 1/n] (numpy's sign(0) = 0 picks the middle).

100. Check yourself - MAE differentiability

Check

Think about the shape of |r| right at r = 0.

Check your understanding

Where is MAE = (1/n)Σ|yᵢ - fᵢ| non-differentiable, and what do we use there?

  • A. At any point where fᵢ = yᵢ; use a subgradient in [-1/n, 1/n] (correct)
  • B. Nowhere - MAE is differentiable everywhere with gradient ±1/n
  • C. At its minimum only, where we use the second derivative
  • D. Everywhere - MAE has no usable gradient at all

Answer: A

Why: |yᵢ - fᵢ| has a corner at fᵢ = yᵢ: the slope jumps from -1/n to +1/n, so no single derivative exists. Optimizers use any subgradient in [-1/n, 1/n] (numpy's sign(0) = 0 picks the middle).

Why B tempts people
This is the trap: it misses the kink. At fᵢ = yᵢ the left slope (-1/n) and right slope (+1/n) disagree, so the derivative is undefined there.
Why C tempts people
The non-differentiable points are the kinks (fᵢ = yᵢ), not 'the minimum', and the fix is a subgradient, not a second derivative.
Why D tempts people
MAE is differentiable everywhere EXCEPT the kinks; away from fᵢ = yᵢ its gradient is a perfectly usable ±1/n.

101. Your turn: build the losses

Section

The project

102. Project: losses and their gradients from scratch

Concept

Implement MSE (value + gradient) and softmax+CE (gradient p - y) in pure NumPy, then verify your cross-entropy gradient against torch's autograd. You derived every piece - now assemble it yourself.

#requirementtool
1MSE value + gradient -2(y - f)/nnumpy
2Stable softmax, then gradient p - ynp.exp(z - z.max())
3Verify dL/dz vs torch autogradtorch + np.allclose

Build rules: type every line yourself, use a numerically stable softmax (subtract the max), and compare to autograd with np.allclose(..., atol=1e-6). When something errors, read the shapes - don't delete the error.

103. Milestone 1 - MSE and its gradient

Worked example

Your turn: write the MSE value and gradient for y = [3,5,7], f = [2,6,4]. Predict the sign of grad[2] before you print (we under-predicted there).

Hint: value is np.mean((y - f)**2); gradient is -2*(y - f)/len(y). The big miss at i = 2 should get the largest-magnitude push.

import numpy as np
y = np.array([3., 5., 7.])
f = np.array([2., 6., 4.])
mse = np.mean((y - f)**2)
grad = -2*(y - f)/len(y)
print(round(mse, 4))
print(grad.round(4))
checkvalue
MSE3.6667
gradient[-0.6667, +0.6667, -2.0]
biggest pushgrad[2] = -2.0 (worst-fit point)

104. Milestone 2 - softmax + CE → p - y

Worked example

Your turn: write a stable softmax, then return p - onehot for logits [1,2,3], true class 2. Say aloud which entry will be negative before you print.

Hint: e = np.exp(z - z.max()); p = e/e.sum(); the gradient is p - onehot. Only the true class (2) should come out negative.

import numpy as np
def softmax(z):
    e = np.exp(z - z.max())
    return e / e.sum()
z = np.array([1., 2., 3.])
onehot = np.array([0., 0., 1.])
p = softmax(z)
print(p.round(4))
print((p - onehot).round(4))
classpp - y
00.0900+0.0900
10.2447+0.2447
2 (true)0.6652-0.3348

105. Watch it run: Milestone 2 - softmax + CE → p - y

Pattern

Step through it

Step through Milestone 2 - softmax + CE → p - y one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: class is 0
  2. Step 2: class is 1
  3. Step 3: class is 2 (true)

106. Milestone 3 - verify against torch

Worked example

Your turn: confirm your p - y matches torch's autograd gradient of cross-entropy. Predict: exact match, or only close?

Hint: build z with requires_grad=True, call cross_entropy(z.unsqueeze(0), tensor([2])), .backward(), then compare z.grad with np.allclose(..., atol=1e-6).

import numpy as np, torch
def softmax(z):
    e = np.exp(z - z.max()); return e/e.sum()
z = np.array([1., 2., 3.]); onehot = np.array([0., 0., 1.])
print((softmax(z) - onehot).round(4))
zt = torch.tensor([1., 2., 3.], requires_grad=True)
torch.nn.functional.cross_entropy(zt.unsqueeze(0), torch.tensor([2])).backward()
print(np.allclose(softmax(z) - onehot, zt.grad.numpy(), atol=1e-6))
sourcedL/dz class 2match
your p - y-0.3348-
torch autograd-0.3348-
np.allclose-True

107. What each one costs: Milestone 3 - verify against torch

Trade off

Comparison matrix

From Milestone 3 - verify against torch: every row here is a choice with a cost. Fill the dL/dz class 2 column, then say which row you would actually pick and what you give up for it.

sourcedL/dz class 2match
your p - y-0.3348-
torch autograd-0.3348-
np.allclose-True

108. The full program

Concept

import numpy as np, torch

def mse(y, f):       return np.mean((y - f)**2)
def mse_grad(y, f):  return -2*(y - f)/len(y)
def softmax(z):
    e = np.exp(z - z.max()); return e/e.sum()

y = np.array([3., 5., 7.]); f = np.array([2., 6., 4.])
print('MSE =', round(mse(y, f), 4), ' grad =', mse_grad(y, f).round(4))

z = np.array([1., 2., 3.]); onehot = np.array([0., 0., 1.])
ce_grad = softmax(z) - onehot            # dL/dz = p - y
zt = torch.tensor([1., 2., 3.], requires_grad=True)
torch.nn.functional.cross_entropy(zt.unsqueeze(0), torch.tensor([2])).backward()
print('CE grad =', ce_grad.round(4))
print('matches torch:', np.allclose(ce_grad, zt.grad.numpy(), atol=1e-6))
printed linevalue
MSE =3.6667 grad = [-0.6667 0.6667 -2.]
CE grad =[0.09 0.2447 -0.3348]
matches torch:True

If your MSE reads 3.6667, your CE gradient reads [0.09, 0.2447, -0.3348], and matches torch prints True - you built the training signal of every classifier from the math up.

109. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

printed linevalue
MSE =3.6667 grad = [-0.6667 0.6667 -2.]
CE grad =[0.09 0.2447 -0.3348]
matches torch:True

110. Show it off

Concept

Slides closed, out loud: explain (1) why softmax+CE collapses to p - y (which pₖ cancels what), (2) why MSE stalls on confident mistakes (name the vanishing factor), and (3) where MAE and hinge are non-differentiable and what you use there.

Stretch (homework): re-derive dL/dz = p - y for softmax+CE from the Jacobian by hand, then prove sigmoid+BCE gives the same. These exact gradients are the seed of Lesson 18's full backward pass.

111. Break it if you can: Show it off

Counterexample

Discussion prompt

Stretch (homework): re-derive dL/dz = p - y for softmax+CE from the Jacobian by hand, then prove sigmoid+BCE gives the same. These exact gradients are the seed of Lesson 18's full backward pass.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

112. Connect it up: Lesson 17: Loss Functions & Their Gradients

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — What a loss even is · MSE - mean squared error · MAE - mean absolute error · Cross-entropy & softmax · The elegant dL/dz = p - y · Why MSE fails for classification. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

113. What you can do now

Recap

movethe one thing to remember
softmax/sigmoid + CEdL/dz = p - y (the Jacobian cancels the 1/p)
MSE for classificationgradient vanishes when confidently wrong: p(1-p) → 0
MSE vs MAEMSE grows with error; MAE is flat ±1/n and robust
kinks (MAE, hinge)non-differentiable → pick a subgradient
choosing a lossmatch the loss to the output so the gradient stays healthy

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 17 (Week 6 - Loss Functions) — Barron - USAAIO Round 2 Preparation, 2026
  2. torch.nn.functional.cross_entropy / binary_cross_entropy_with_logits
  3. Goodfellow, Bengio, Courville - Deep Learning, Ch. 6.2 (Cost functions) & 3.13 (Cross-entropy) — MIT Press, 2016
  4. Every gradient, probability, and loss value produced by real execution — numpy 2.2 + torch 2.7.1, verification run July 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108