Lesson 61: Multiclass Classification — Softmax, OvR, OvO

USAAIO Lesson 61. It covers softmax regression, giving P(y=k|x) and the cross-entropy gradient P − Y_one_hot, then the One-vs-Rest and One-vs-One multiclass strategies, multinomial logistic regression against a set of separate binary logistic models, and weighted cross-entropy for class imbalance. You implement multinomial logistic regression from scratch and benchmark One-vs-Rest, One-vs-One, and multinomial on an imbalanced 10-class dataset. The lesson runs to 29 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Multiclass Classification: Softmax, OvR & OvO

Title

USAAIO · Lesson 61

From binary to K classes: the softmax distribution, its exact gradient, and two reduction strategies — One-vs-Rest and One-vs-One. Closes with class imbalance in the multiclass setting.

2. By the end of this lesson you can

Objectives

  1. Derive the softmax probability P(y=k|x) and its cross-entropy gradient dL/dz = P − Y_one_hot
  2. Implement multinomial logistic regression from scratch (forward pass + GD update)
  3. Distinguish OvR (K classifiers) from OvO (K(K−1)/2 classifiers) and choose each for the right model class
  4. Explain the subtle difference between multinomial and K independent binary logistic models
  5. Apply weighted cross-entropy and balanced sampling to handle multiclass imbalance

3. What survived from sklearn Pipeline, ColumnTransformer & GridSearchCV?

Warm-up

Discussion prompt

Before we open Lesson 61: Multiclass Classification — Softmax, OvR, OvO: without looking back, what was the main idea of sklearn Pipeline, ColumnTransformer & GridSearchCV, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

sklearn Pipeline chains estimators and prevents leakage; ColumnTransformer applies different transforms to different feature types; FunctionTransformer wraps any function; Pipeline+GridSearchCV tunes preprocessing and model hyperparameters jointly; joblib saves and loads versioned pipelines. Build a production-ready impute→encode→scale→LogReg pipeline, tune it end-to-end, and serialize it.

4. The softmax distribution

Section

Part 1 of 3

5. Softmax: logits → probabilities

Concept

A K-class model outputs K logits z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.

\[ P(y=k \mid x) = \frac{\exp(z_k)}{\sum_{j=1}^{K} \exp(z_j)} = \frac{\exp(w_k^\top x)}{\sum_{j=1}^{K} \exp(w_j^\top x)} \]

The denominator is the partition function. Larger logit → larger probability. All K probabilities sum to exactly 1 by construction.

6. Break it if you can: Softmax: logits → probabilities

Counterexample

Discussion prompt

A K-class model outputs K logits z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The denominator is the partition function. Larger logit → larger probability. All K probabilities sum to exactly 1 by construction.

7. Guess the shape of the answer: Softmax forward pass — traced

Estimation

Predict first

1 sample, 3 classes. Logits z = [2.0, 1.0, 0.1], true label y = 0.

Commit before you compute: what does Softmax forward pass — traced come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: CE loss = −log(0.6590) = 0.4170

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Only the true-class probability enters the cross-entropy loss; the rest influence through the softmax denominator.

8. Softmax forward pass — traced

Worked example

1 sample, 3 classes. Logits z = [2.0, 1.0, 0.1], true label y = 0.

import numpy as np
z = np.array([2.0, 1.0, 0.1])
exp_z = np.exp(z)                   # [7.3891, 2.7183, 1.1052]
p = exp_z / exp_z.sum()             # normalize
loss = -np.log(p[0])                # CE for true class 0
print('P:', p.round(4))
print('CE loss:', round(loss, 4))
stepclass 0class 1class 2
logit z2.00001.00000.1000
exp(z)7.38912.71831.1052
P(y=k|x)0.65900.24240.0986
P − Y (true=0)−0.34100.24240.0986

CE loss = −log(0.6590) = 0.4170

Why: Only the true-class probability enters the cross-entropy loss; the rest influence through the softmax denominator.

9. Which is which, by class 1

Discrimination

Sort into buckets

Sort these by class 1, from memory, without looking back at Softmax forward pass — traced. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1.0000
logit z
2.7183
exp(z)
0.2424
P(y=k|x); P − Y (true=0)
g1
class 1 is "1.0000" for logit z — that is what the table on "Softmax forward pass — traced" records, and it is the single property separating this group from the rest.
g2
class 1 is "2.7183" for exp(z) — that is what the table on "Softmax forward pass — traced" records, and it is the single property separating this group from the rest.
g3
class 1 is "0.2424" for P(y=k|x), P − Y (true=0) — that is what the table on "Softmax forward pass — traced" records, and it is the single property separating this group from the rest.

10. Cross-entropy gradient: dL/dz = P − Y

Concept

The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.

\[ \frac{\partial \mathcal{L}}{\partial z_k} = P_k - \mathbb{1}[y = k] = P_k - Y_k^{\text{one-hot}} \]

\[ \Rightarrow \frac{\partial \mathcal{L}}{\partial W} = X^\top (P - Y) \quad \text{(vectorized over the batch)} \]

Predicted probability minus one-hot truth — that's the whole gradient. The chain rule through the softmax and log cancel cleanly. (Numerical check: max error 3.16×10⁻¹¹.)

11. By analogy: Cross-entropy gradient: dL/dz = P − Y

Analogy

Discussion prompt

Explain Cross-entropy gradient: dL/dz = P − Y by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.

12. Something is wrong here: numerically unstable softmax

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compute exp(z) directly, then divide by the sum.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: With large logits exp(z_k) overflows IEEE float64.

Subtract max(z) before exponentiating — the softmax value is unchanged.

Why: With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.

13. Trap: numerically unstable softmax

Trap

The trap

Compute exp(z) directly, then divide by the sum.

z = [1000, 999, 0] → exp(1000) = overflow → NaN probabilities

Why: With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.

The fix

Subtract max(z) before exponentiating — the softmax value is unchanged.

z' = z − max(z) → exp(z') is bounded in (0,1]; divide by sum

Why: Shifting by max(z) keeps all exponents ≤ 0, so exp ≤ 1. The ratio is identical: exp(z_k)/Σexp(z_j) = exp(z_k−c)/Σexp(z_j−c). PyTorch's F.softmax does this automatically.

14. Break it on purpose: numerically unstable softmax

Break the constraint

Discussion prompt

The rule this trap just fixed:

Shifting by max(z) keeps all exponents ≤ 0, so exp ≤ 1. The ratio is identical: exp(z_k)/Σexp(z_j) = exp(z_k−c)/Σexp(z_j−c). PyTorch's F.softmax does this automatically.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.

15. What has to be given first: Multinomial logistic from scratch (Iris)

Missing information

Discussion prompt

Fit 3-class Iris with GD on softmax CE. Weight matrix W is (d+1)×K; gradient update: W -= lr * X^T @ (P − Y) / n.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Iris is nearly linearly separable; multinomial logistic converges quickly. sklearn's LBFGS reaches 97.33% — the gap is solver quality, not model capacity.

16. Multinomial logistic from scratch (Iris)

Worked example

Fit 3-class Iris with GD on softmax CE. Weight matrix W is (d+1)×K; gradient update: W -= lr * X^T @ (P − Y) / n.

import numpy as np
from sklearn.datasets import load_iris

def softmax(Z):
    E = np.exp(Z - Z.max(axis=1, keepdims=True))
    return E / E.sum(axis=1, keepdims=True)

iris = load_iris()
X = iris.data - iris.data.mean(0)          # center
Xb = np.c_[X, np.ones(len(X))]            # add bias
y = iris.target
Y = np.eye(3)[y]                           # one-hot

W = np.zeros((Xb.shape[1], 3))
for epoch in range(500):
    P = softmax(Xb @ W)
    L = -np.sum(Y * np.log(P + 1e-15)) / len(y)
    W -= 0.1 * Xb.T @ ((P - Y) / len(y))

acc = (np.argmax(Xb @ W, 1) == y).mean()
print(f'acc={acc:.4f}  loss={L:.4f}')
epochlossaccuracy
01.0986—
1000.3193—
4990.16470.9667

Loss 1.0986 → 0.1647, accuracy 96.67% on Iris in 500 GD steps

Why: Iris is nearly linearly separable; multinomial logistic converges quickly. sklearn's LBFGS reaches 97.33% — the gap is solver quality, not model capacity.

17. Fill in: accuracy for Multinomial logistic from scratch (Iris)

Comparison

Comparison matrix

From Multinomial logistic from scratch (Iris): refill the accuracy column from what you know. The rest of the table is as it appeared.

epochlossaccuracy
01.0986—
1000.3193—
4990.16470.9667

18. OvR and OvO reduction strategies

Section

Part 2 of 3

19. One-vs-Rest (OvR)

Concept

Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.

\[ f_k(x) = \sigma(w_k^\top x), \quad \hat{y} = \arg\max_k\, f_k(x) \]

For K=10: exactly 10 classifiers. Fast to train (K logistic models), easy to calibrate, and each classifier's weight vector is interpretable as 'what makes class k different from everything else'.

20. By analogy: One-vs-Rest (OvR)

Analogy

Discussion prompt

Explain One-vs-Rest (OvR) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.

21. One-vs-One (OvO)

Concept

Train K(K−1)/2 binary classifiers — one for every pair of classes. Predict by majority vote across all pairwise decisions.

\[ \hat{y} = \text{majority-vote over} \bigl\{f_{ij}(x)\bigr\}_{i<j} \]

K classesOvR classifiersOvO classifiers
333
5510
101045

OvO is preferred for SVM because each sub-problem is small (only 2-class data). OvR is preferred for logistic regression because calibrated probabilities let you just argmax.

22. What each one costs: One-vs-One (OvO)

Trade off

Comparison matrix

From One-vs-One (OvO): every row here is a choice with a cost. Fill the OvO classifiers column, then say which row you would actually pick and what you give up for it.

K classesOvR classifiersOvO classifiers
333
5510
101045

23. Something is wrong here: confusing OvR score calibration

Anomaly

Predict first

A student writes this, and it looks reasonable:

Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others.

For OvR, simply argmax the K scores (no normalization needed for classification).

Why: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others. The K scores don't need to sum to 1 and normalizing them is not mathematically justified.

24. Trap: confusing OvR score calibration

Trap

The trap

Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).

Treat OvR scores as a probability simplex by normalization

Why: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others. The K scores don't need to sum to 1 and normalizing them is not mathematically justified.

The fix

For OvR, simply argmax the K scores (no normalization needed for classification).

ŷ = argmax_k f_k(x) — pick the most confident positive detector

Why: If you need valid probabilities across K classes, use softmax (multinomial logistic) instead. OvR's scores are decision-function outputs, not a joint distribution.

25. Which of these survive contact with Lesson 61: Multiclass Classification —…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
A K-class model outputs K logits z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.; The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.; Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.
Breaks
Compute exp(z) directly, then divide by the sum.; Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 61: Multiclass Classification — Softmax, OvR, OvO puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

26. Multinomial vs K independent binary logistic

Concept

Multinomial logistic (softmax) trains W jointly — all K weight vectors are updated together through the shared partition function. The K probabilities always sum to 1.

OvR with K sigmoid models trains each binary classifier independently. No joint normalization — each classifier optimizes its own binary CE. Learning dynamics differ: OvR doesn't push down probabilities of wrong classes during a single update.

propertymultinomialOvR binary
joint normalizationyes (softmax)no
probabilities sum to 1yesonly by luck
update couplingall classes per stepone classifier per problem
gradient for true class kP_k − 1 (softmax CE)σ(z_k) − 1 (binary BCE)
gradient for wrong class jP_j (pushes down)no cross-class signal

27. Guess the shape of the answer: OvR vs OvO vs Multinomial benchmark

Estimation

Predict first

10-class imbalanced dataset (n=1000, class sizes 297→31). Compare OvR logistic, multinomial logistic, OvO SVM — 5-fold CV.

Commit before you compute: what does OvR vs OvO vs Multinomial benchmark come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: OvO SVM wins (+8.6 pp) because SVM's dual is inherently binary — each pairwise sub-problem uses the full kernel margin, not a linearization

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Multinomial and OvR logistic are linear in feature space; OvO SVM-RBF fits non-linear boundaries per pair.

28. OvR vs OvO vs Multinomial benchmark

Worked example

10-class imbalanced dataset (n=1000, class sizes 297→31). Compare OvR logistic, multinomial logistic, OvO SVM — 5-fold CV.

from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.multiclass import OneVsRestClassifier, OneVsOneClassifier
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
import numpy as np

X, y = make_classification(
    n_samples=1000, n_features=20, n_informative=15,
    n_classes=10, n_clusters_per_class=1,
    weights=[0.3,0.2,0.1,0.08,0.07,0.07,0.06,0.05,0.04,0.03],
    random_state=42)

ovr = OneVsRestClassifier(LogisticRegression(max_iter=500, random_state=0))
mn  = LogisticRegression(max_iter=500, random_state=0, solver='lbfgs')
ovo = OneVsOneClassifier(SVC(kernel='rbf', random_state=0))

for name, clf in [('OvR LR', ovr), ('MN LR', mn), ('OvO SVM', ovo)]:
    s = cross_val_score(clf, X, y, cv=5)
    print(f'{name}: {s.mean():.4f} ± {s.std():.4f}')
strategy5-fold accuracy# classifiers
OvR LogReg0.6590 ± 0.029410
Multinomial LogReg0.6590 ± 0.03541 (K-way)
OvO SVM (RBF)0.7450 ± 0.008445

OvO SVM wins (+8.6 pp) because SVM's dual is inherently binary — each pairwise sub-problem uses the full kernel margin, not a linearization

Why: Multinomial and OvR logistic are linear in feature space; OvO SVM-RBF fits non-linear boundaries per pair. The accuracy gap will close once you add kernel/polynomial features to logistic.

29. Fill in: 5-fold accuracy for OvR vs OvO vs Multinomial benchmark

Comparison

Comparison matrix

From OvR vs OvO vs Multinomial benchmark: refill the 5-fold accuracy column from what you know. The rest of the table is as it appeared.

strategy5-fold accuracy# classifiers
OvR LogReg0.6590 ± 0.029410
Multinomial LogReg0.6590 ± 0.03541 (K-way)
OvO SVM (RBF)0.7450 ± 0.008445

30. Class imbalance in multiclass

Section

Part 3 of 3

31. Why imbalance hurts multiclass more

Concept

In binary problems, imbalance biases the decision boundary toward the majority class. In K-class softmax, the rare-class weight vectors receive far fewer gradient updates — they stay near zero while majority-class vectors diverge.

The model learns to never predict rare classes: accuracy appears high (dominated by majority-class predictions) but recall on rare classes is near zero. Macro-F1 or per-class accuracy is the right metric here, not overall accuracy.

32. Two fixes: weighted CE and balanced sampling

Concept

Weighted cross-entropy: multiply each sample's CE contribution by its class weight. Balanced weights = 1/(K · frac_k), so a class with 3% frequency gets weight ≈3.26×.

\[ \mathcal{L}_{\text{weighted}} = -\frac{1}{n}\sum_{i=1}^n w_{y_i} \log P(y_i \mid x_i) \]

Class-balanced sampling: oversample rare classes (e.g. SMOTE) or undersample majority classes so each mini-batch has roughly equal representation. Complementary to weighting — use one or both.

33. Guess the shape of the answer: Balanced weights on 10-class imbalanced data

Estimation

Predict first

Compute sklearn balanced class weights and compare weighted vs unweighted multinomial logistic accuracy.

Commit before you compute: what does Balanced weights on 10-class imbalanced data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Weighted accuracy 0.5540 < unweighted 0.6590 — this is expected

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Weighting forces the model to attend to rare classes at the cost of overall accuracy.

34. Balanced weights on 10-class imbalanced data

Worked example

Compute sklearn balanced class weights and compare weighted vs unweighted multinomial logistic accuracy.

from sklearn.utils.class_weight import compute_class_weight
import numpy as np

cw = compute_class_weight('balanced', classes=np.arange(10), y=y)
print('class | weight')
for i in range(10):
    print(f'  {i:2d}  | {cw[i]:.4f}')

mn_w = LogisticRegression(max_iter=500, class_weight='balanced',
                           random_state=0, solver='lbfgs')
s_w = cross_val_score(mn_w, X, y, cv=5)
print(f'Weighted MN 5-fold: {s_w.mean():.4f} ± {s_w.std():.4f}')
classcountbalanced weight
0 (majority)2970.3367
4721.3889
7492.0408
9 (minority)313.2258

Weighted accuracy 0.5540 < unweighted 0.6590 — this is expected

Why: Weighting forces the model to attend to rare classes at the cost of overall accuracy. The right metric is per-class recall or macro-F1, not micro accuracy. Weighted cross-entropy is a mechanism, not a free accuracy boost.

35. Watch it run: Balanced weights on 10-class imbalanced data

Pattern

Step through it

Step through Balanced weights on 10-class imbalanced data one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: class is 0 (majority)
  2. Step 2: class is 4
  3. Step 3: class is 7
  4. Step 4: class is 9 (minority)

36. Rebuild the recipe: The multiclass logistic recipe

Ranking

Put in order

These are the steps of The multiclass logistic recipe, scrambled. Put them back in order before the next slide shows you.

  1. Softmax: logits Z = XW; P = softmax(Z); loss = CE
  2. Gradient: dL/dZ = (P − Y_one_hot)/n; dW = X^T dZ
  3. Choose strategy: logistic → multinomial (or OvR); SVM → OvO (pairwise binary)
  4. Key difference: multinomial couples all K classes per update; OvR trains K independent binary problems
  5. Imbalance: class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracy

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

37. The multiclass logistic recipe

Pattern

  1. Softmax: logits Z = XW; P = softmax(Z); loss = CE
  2. Gradient: dL/dZ = (P − Y_one_hot)/n; dW = X^T dZ
  3. Choose strategy: logistic → multinomial (or OvR); SVM → OvO (pairwise binary)
  4. Key difference: multinomial couples all K classes per update; OvR trains K independent binary problems
  5. Imbalance: class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracy

38. Where does it stop working: The multiclass logistic recipe

Edge cases

Discussion prompt

The multiclass logistic recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Softmax: logits Z = XW; P = softmax(Z); loss = CE
  2. Gradient: dL/dZ = (P − Y_one_hot)/n; dW = X^T dZ
  3. Choose strategy: logistic → multinomial (or OvR); SVM → OvO (pairwise binary)
  4. Key difference: multinomial couples all K classes per update; OvR trains K independent binary problems
  5. Imbalance: class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracy

39. Rule out three: Check yourself — softmax gradient

Elimination

Eliminate the wrong options

For a softmax model with logits z, the gradient of the cross-entropy loss w.r.t. z_k for the true class k=0 is:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. P_0 − 1
  • B. 1 − P_0
  • C. P_0
  • D. −log P_0

Survives elimination: A

Why: The gradient dL/dz_k = P_k − Y_k. For the true class Y_k = 1, so dL/dz_0 = P_0 − 1. This is negative (pushing the logit up) when P_0 < 1 — the correct direction.

40. Check yourself — softmax gradient

Check

Pen-and-paper first.

Check your understanding

For a softmax model with logits z, the gradient of the cross-entropy loss w.r.t. z_k for the true class k=0 is:

  • A. P_0 − 1 (correct)
  • B. 1 − P_0
  • C. P_0
  • D. −log P_0

Answer: A

Why: The gradient dL/dz_k = P_k − Y_k. For the true class Y_k = 1, so dL/dz_0 = P_0 − 1. This is negative (pushing the logit up) when P_0 < 1 — the correct direction.

Why B tempts people
1 − P_0 is positive, which would decrease the true-class logit — that's the wrong sign, it would hurt the model.
Why C tempts people
P_0 is the gradient for a wrong class (Y_k = 0), not the true class. For wrong classes dL/dz_j = P_j (push down).
Why D tempts people
−log P_0 is the CE loss value itself, not its gradient w.r.t. the logit.

41. Answer it before you see the options: Check yourself — OvR vs OvO

Prediction

Predict first

You want to extend a binary SVM (kernel trick, quadratic program in the dual) to 10 classes. Which reduction strategy is preferred and why?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: OvO — each of 45 pairwise sub-problems is a binary SVM, which the dual is designed for

Why: The SVM dual is inherently a binary QP — its kernel trick and margin maximization assume exactly two classes. OvO decomposes K-class SVM into K(K−1)/2 natural binary SVMs. OvR's K sub-problems use imbalanced negatives (all other classes), which degrades the margin.

42. Check yourself — OvR vs OvO

Check

Choose the right strategy.

Check your understanding

You want to extend a binary SVM (kernel trick, quadratic program in the dual) to 10 classes. Which reduction strategy is preferred and why?

  • A. OvO — each of 45 pairwise sub-problems is a binary SVM, which the dual is designed for (correct)
  • B. OvR — only 10 classifiers instead of 45, so it's always faster and equally accurate
  • C. Multinomial softmax — the gradient dL/dz = P − Y works for any classifier
  • D. OvR — because SVMs produce calibrated probabilities that can be argmaxed

Answer: A

Why: The SVM dual is inherently a binary QP — its kernel trick and margin maximization assume exactly two classes. OvO decomposes K-class SVM into K(K−1)/2 natural binary SVMs. OvR's K sub-problems use imbalanced negatives (all other classes), which degrades the margin.

Why B tempts people
Fewer classifiers doesn't mean equal or better accuracy. OvR's imbalanced negatives make each binary margin meaningless for SVM; OvO is empirically better here (we saw +8.6 pp in the benchmark).
Why C tempts people
Softmax is for logistic regression models, not SVMs. SVMs don't have a softmax output or cross-entropy loss.
Why D tempts people
SVMs don't naturally produce calibrated probabilities — Platt scaling is needed. Even then, OvR is not preferred for SVM.

43. Rule out three: Check yourself — class imbalance

Elimination

Eliminate the wrong options

You train a 10-class multinomial logistic model on severely imbalanced data (one class has 3% of samples). After adding class_weight='balanced', overall accuracy drops from 65.9% to 55.4%. What is the most likely correct interpretation?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The model is now attending to rare classes; macro-F1 likely improved even though micro accuracy fell
  • B. Balanced weighting is wrong here — it hurts the model; remove it
  • C. The model underfit; add more epochs to recover the accuracy
  • D. Weighting is equivalent to oversampling, so accuracy should not change

Survives elimination: A

Why: Balanced weighting intentionally sacrifices majority-class accuracy to improve minority-class recall. Overall accuracy is a micro metric dominated by the majority class — it falls. Macro-F1 (unweighted average over classes) typically improves, which is the real goal when class frequencies differ.

44. Check yourself — class imbalance

Check

Read the scenario carefully.

Check your understanding

You train a 10-class multinomial logistic model on severely imbalanced data (one class has 3% of samples). After adding class_weight='balanced', overall accuracy drops from 65.9% to 55.4%. What is the most likely correct interpretation?

  • A. The model is now attending to rare classes; macro-F1 likely improved even though micro accuracy fell (correct)
  • B. Balanced weighting is wrong here — it hurts the model; remove it
  • C. The model underfit; add more epochs to recover the accuracy
  • D. Weighting is equivalent to oversampling, so accuracy should not change

Answer: A

Why: Balanced weighting intentionally sacrifices majority-class accuracy to improve minority-class recall. Overall accuracy is a micro metric dominated by the majority class — it falls. Macro-F1 (unweighted average over classes) typically improves, which is the real goal when class frequencies differ.

Why B tempts people
The drop in accuracy is expected and desirable — it means the model is no longer ignoring rare classes. Removing weighting would restore accuracy but collapse minority-class recall to near zero.
Why C tempts people
This is not an underfitting phenomenon. More epochs on the same imbalanced loss would just re-learn the same biased solution.
Why D tempts people
Weighting and oversampling are mathematically related for small datasets but not equivalent in general. Either way, overall micro accuracy will change because the loss surface changes.

45. Your turn: implement it

Section

Project

46. Project: multiclass logistic classifier

Concept

Build a multinomial logistic classifier from scratch, then compare OvR / OvO / multinomial strategies on an imbalanced 10-class problem.

milestonetaskkey tool
1Implement softmax + CE forward pass, verify gradient numericallynumpy, finite differences
2Train multinomial logistic on Iris (GD loop)softmax, W update
3Benchmark OvR / MN / OvO on 10-class imbalanced datasklearn multiclass wrappers
4Add class_weight='balanced' and report macro-F1classification_report

Build rules: implement your own softmax function with max-shift for stability; verify the gradient before training. In the benchmark, measure macro-F1 alongside accuracy.

47. Break it if you can: Project: multiclass logistic classifier

Counterexample

Discussion prompt

Build a multinomial logistic classifier from scratch, then compare OvR / OvO / multinomial strategies on an imbalanced 10-class problem.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: implement your own softmax function with max-shift for stability; verify the gradient before training. In the benchmark, measure macro-F1 alongside accuracy.

48. Milestone 1 — softmax forward + gradient check

Worked example

Your turn: implement softmax(Z) and verify that your analytic gradient matches a numerical gradient (finite differences). Target max error < 1e-8.

Hint: stable softmax subtracts Z.max(axis=1, keepdims=True) before exp. Gradient check: perturb W by ±ε, compute (L_plus − L_minus)/(2ε) per parameter.

import numpy as np

def softmax(Z):
    E = np.exp(Z - Z.max(axis=1, keepdims=True))
    return E / E.sum(axis=1, keepdims=True)

# 5 samples, 3 classes, 4 features
rng = np.random.default_rng(42)
W = rng.normal(size=(4, 3))
X = rng.normal(size=(5, 4))
y = rng.integers(0, 3, size=5)
Y = np.eye(3)[y]

P = softmax(X @ W)
dW_analytic = X.T @ ((P - Y) / len(y))

# numerical gradient
eps = 1e-5; dW_num = np.zeros_like(W)
for i in range(4):
    for j in range(3):
        def loss(W_):
            return -np.sum(Y * np.log(softmax(X@W_)+1e-15))/len(y)
        Wp, Wm = W.copy(), W.copy()
        Wp[i,j] += eps; Wm[i,j] -= eps
        dW_num[i,j] = (loss(Wp)-loss(Wm))/(2*eps)

print('max err:', np.max(np.abs(dW_analytic - dW_num)))
checkresult
max gradient error3.16e-11
passes < 1e-8 thresholdyes

49. What each one costs: Milestone 1 — softmax forward + gradient check

Trade off

Comparison matrix

From Milestone 1 — softmax forward + gradient check: every row here is a choice with a cost. Fill the result column, then say which row you would actually pick and what you give up for it.

checkresult
max gradient error3.16e-11
passes < 1e-8 thresholdyes

50. Milestone 2 — multinomial GD on Iris

Worked example

Your turn: train multinomial logistic on Iris with 500 GD steps (lr=0.1). Predict the final accuracy and loss.

Hint: mean-center X; append a bias column; one-hot Y; update W -= lr * Xb.T @ ((P-Y)/n) each epoch.

from sklearn.datasets import load_iris
iris = load_iris()
X = iris.data - iris.data.mean(0)
Xb = np.c_[X, np.ones(len(X))]
y = iris.target
Y = np.eye(3)[y]

W = np.zeros((Xb.shape[1], 3))
for ep in range(500):
    P = softmax(Xb @ W)
    L = -np.sum(Y * np.log(P + 1e-15)) / len(y)
    W -= 0.1 * Xb.T @ ((P - Y) / len(y))

acc = (np.argmax(Xb @ W, 1) == y).mean()
print(f'acc={acc:.4f}  loss={L:.4f}')
epochlossaccuracy
01.0986—
1000.3193—
4990.16470.9667

51. Fill in: loss for Milestone 2 — multinomial GD on Iris

Comparison

Comparison matrix

From Milestone 2 — multinomial GD on Iris: refill the loss column from what you know. The rest of the table is as it appeared.

epochlossaccuracy
01.0986—
1000.3193—
4990.16470.9667

52. Milestone 3 — benchmark + macro-F1

Worked example

Your turn: on the 10-class imbalanced dataset, run OvR, multinomial, and OvO SVM. Then add balanced weighting and check whether macro-F1 improves.

Hint: from sklearn.metrics import classification_report; cross_val_score(..., scoring='f1_macro') for macro-F1 across CV folds.

from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.multiclass import OneVsRestClassifier, OneVsOneClassifier
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score

X, y = make_classification(
    n_samples=1000, n_features=20, n_informative=15,
    n_classes=10, n_clusters_per_class=1,
    weights=[0.3,0.2,0.1,0.08,0.07,0.07,0.06,0.05,0.04,0.03],
    random_state=42)

ovr = OneVsRestClassifier(LogisticRegression(max_iter=500, random_state=0))
mn  = LogisticRegression(max_iter=500, random_state=0, solver='lbfgs')
ovo = OneVsOneClassifier(SVC(kernel='rbf', random_state=0))
mn_w = LogisticRegression(max_iter=500, class_weight='balanced',
                           random_state=0, solver='lbfgs')

for name, clf in [('OvR LR',ovr),('MN LR',mn),('OvO SVM',ovo),('MN+bal',mn_w)]:
    acc = cross_val_score(clf, X, y, cv=5).mean()
    print(f'{name}: acc={acc:.4f}')
strategy5-fold accuracy
OvR LogReg0.6590
Multinomial LogReg0.6590
OvO SVM (RBF)0.7450
Multinomial + balanced0.5540

53. Which is which, by 5-fold accuracy

Discrimination

Sort into buckets

Sort these by 5-fold accuracy, from memory, without looking back at Milestone 3 — benchmark + macro-F1. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

0.6590
OvR LogReg; Multinomial LogReg
0.7450
OvO SVM (RBF)
0.5540
Multinomial + balanced
g1
5-fold accuracy is "0.6590" for OvR LogReg, Multinomial LogReg — that is what the table on "Milestone 3 — benchmark + macro-F1" records, and it is the single property separating this group from the rest.
g2
5-fold accuracy is "0.7450" for OvO SVM (RBF) — that is what the table on "Milestone 3 — benchmark + macro-F1" records, and it is the single property separating this group from the rest.
g3
5-fold accuracy is "0.5540" for Multinomial + balanced — that is what the table on "Milestone 3 — benchmark + macro-F1" records, and it is the single property separating this group from the rest.

54. Show it off

Concept

Out loud, slides closed: (1) Derive the softmax CE gradient — state why it equals P − Y. (2) Explain when OvO beats OvR. (3) Why does balanced weighting lower overall accuracy but improve model fairness?

Stretch (homework): add a mini-batch training loop for the multinomial scratch implementation; compare OvR/OvO/multinomial with macro-F1 (not just accuracy); derive the gradient from first principles as the math homework requires.

55. Connect it up: Lesson 61: Multiclass Classification — Softmax, OvR, OvO

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The softmax distribution · OvR and OvO reduction strategies · Class imbalance in multiclass · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. What you can do now

Recap

ideathe one thing to remember
softmax gradientdL/dz = P − Y (true class −(1−P), wrong class +P)
OvR vs OvOlogistic → OvR; SVM → OvO (binary dual fits naturally)
multinomial vs OvRmultinomial couples K-class updates; OvR trains K independent BCE problems
imbalance fixclass_weight='balanced'; evaluate macro-F1 not accuracy

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 61 — Softmax Regression, OvR, OvO, Multiclass Imbalance — Barron · USAAIO Round 2 Preparation, 2026
  2. Softmax gradient derivation (analytic vs numerical, max err 3.16e-11), Iris scratch GD, 10-class OvR/OvO/MN benchmark, all verified — numpy 2.2.6 + scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108