USAAIO Lesson 61. It covers softmax regression, giving P(y=k|x) and the cross-entropy gradient P − Y_one_hot, then the One-vs-Rest and One-vs-One multiclass strategies, multinomial logistic regression against a set of separate binary logistic models, and weighted cross-entropy for class imbalance. You implement multinomial logistic regression from scratch and benchmark One-vs-Rest, One-vs-One, and multinomial on an imbalanced 10-class dataset. The lesson runs to 29 slides.
Subject: Machine Learning · 56 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 61
From binary to K classes: the softmax distribution, its exact gradient, and two reduction strategies — One-vs-Rest and One-vs-One. Closes with class imbalance in the multiclass setting.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 61: Multiclass Classification — Softmax, OvR, OvO: without looking back, what was the main idea of sklearn Pipeline, ColumnTransformer & GridSearchCV, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
sklearn Pipeline chains estimators and prevents leakage; ColumnTransformer applies different transforms to different feature types; FunctionTransformer wraps any function; Pipeline+GridSearchCV tunes preprocessing and model hyperparameters jointly; joblib saves and loads versioned pipelines. Build a production-ready impute→encode→scale→LogReg pipeline, tune it end-to-end, and serialize it.
Section
Part 1 of 3
Concept
A K-class model outputs K logits z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.
\[ P(y=k \mid x) = \frac{\exp(z_k)}{\sum_{j=1}^{K} \exp(z_j)} = \frac{\exp(w_k^\top x)}{\sum_{j=1}^{K} \exp(w_j^\top x)} \]
The denominator is the partition function. Larger logit → larger probability. All K probabilities sum to exactly 1 by construction.
Counterexample
Discussion prompt
A K-class model outputs K logits z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The denominator is the partition function. Larger logit → larger probability. All K probabilities sum to exactly 1 by construction.
Estimation
Predict first
1 sample, 3 classes. Logits z = [2.0, 1.0, 0.1], true label y = 0.
Commit before you compute: what does Softmax forward pass — traced come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: CE loss = −log(0.6590) = 0.4170
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Only the true-class probability enters the cross-entropy loss; the rest influence through the softmax denominator.
Worked example
1 sample, 3 classes. Logits z = [2.0, 1.0, 0.1], true label y = 0.
import numpy as np
z = np.array([2.0, 1.0, 0.1])
exp_z = np.exp(z) # [7.3891, 2.7183, 1.1052]
p = exp_z / exp_z.sum() # normalize
loss = -np.log(p[0]) # CE for true class 0
print('P:', p.round(4))
print('CE loss:', round(loss, 4))| step | class 0 | class 1 | class 2 |
|---|---|---|---|
| logit z | 2.0000 | 1.0000 | 0.1000 |
| exp(z) | 7.3891 | 2.7183 | 1.1052 |
| P(y=k|x) | 0.6590 | 0.2424 | 0.0986 |
| P − Y (true=0) | −0.3410 | 0.2424 | 0.0986 |
CE loss = −log(0.6590) = 0.4170
Why: Only the true-class probability enters the cross-entropy loss; the rest influence through the softmax denominator.
Discrimination
Sort into buckets
Sort these by class 1, from memory, without looking back at Softmax forward pass — traced. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.
\[ \frac{\partial \mathcal{L}}{\partial z_k} = P_k - \mathbb{1}[y = k] = P_k - Y_k^{\text{one-hot}} \]
\[ \Rightarrow \frac{\partial \mathcal{L}}{\partial W} = X^\top (P - Y) \quad \text{(vectorized over the batch)} \]
Predicted probability minus one-hot truth — that's the whole gradient. The chain rule through the softmax and log cancel cleanly. (Numerical check: max error 3.16×10⁻¹¹.)
Analogy
Discussion prompt
Explain Cross-entropy gradient: dL/dz = P − Y by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Compute exp(z) directly, then divide by the sum.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: With large logits exp(z_k) overflows IEEE float64.
Subtract max(z) before exponentiating — the softmax value is unchanged.
Why: With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.
Trap
Compute exp(z) directly, then divide by the sum.
z = [1000, 999, 0] → exp(1000) = overflow → NaN probabilities
Why: With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.
Subtract max(z) before exponentiating — the softmax value is unchanged.
z' = z − max(z) → exp(z') is bounded in (0,1]; divide by sum
Why: Shifting by max(z) keeps all exponents ≤ 0, so exp ≤ 1. The ratio is identical: exp(z_k)/Σexp(z_j) = exp(z_k−c)/Σexp(z_j−c). PyTorch's F.softmax does this automatically.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Shifting by max(z) keeps all exponents ≤ 0, so exp ≤ 1. The ratio is identical: exp(z_k)/Σexp(z_j) = exp(z_k−c)/Σexp(z_j−c). PyTorch's F.softmax does this automatically.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
With large logits exp(z_k) overflows IEEE float64. All probabilities become NaN; training collapses silently.
Missing information
Discussion prompt
Fit 3-class Iris with GD on softmax CE. Weight matrix W is (d+1)×K; gradient update: W -= lr * X^T @ (P − Y) / n.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Iris is nearly linearly separable; multinomial logistic converges quickly. sklearn's LBFGS reaches 97.33% — the gap is solver quality, not model capacity.
Worked example
Fit 3-class Iris with GD on softmax CE. Weight matrix W is (d+1)×K; gradient update: W -= lr * X^T @ (P − Y) / n.
import numpy as np
from sklearn.datasets import load_iris
def softmax(Z):
E = np.exp(Z - Z.max(axis=1, keepdims=True))
return E / E.sum(axis=1, keepdims=True)
iris = load_iris()
X = iris.data - iris.data.mean(0) # center
Xb = np.c_[X, np.ones(len(X))] # add bias
y = iris.target
Y = np.eye(3)[y] # one-hot
W = np.zeros((Xb.shape[1], 3))
for epoch in range(500):
P = softmax(Xb @ W)
L = -np.sum(Y * np.log(P + 1e-15)) / len(y)
W -= 0.1 * Xb.T @ ((P - Y) / len(y))
acc = (np.argmax(Xb @ W, 1) == y).mean()
print(f'acc={acc:.4f} loss={L:.4f}')| epoch | loss | accuracy |
|---|---|---|
| 0 | 1.0986 | — |
| 100 | 0.3193 | — |
| 499 | 0.1647 | 0.9667 |
Loss 1.0986 → 0.1647, accuracy 96.67% on Iris in 500 GD steps
Why: Iris is nearly linearly separable; multinomial logistic converges quickly. sklearn's LBFGS reaches 97.33% — the gap is solver quality, not model capacity.
Comparison
Comparison matrix
From Multinomial logistic from scratch (Iris): refill the accuracy column from what you know. The rest of the table is as it appeared.
| epoch | loss | accuracy |
|---|---|---|
| 0 | 1.0986 | — |
| 100 | 0.3193 | — |
| 499 | 0.1647 | 0.9667 |
Section
Part 2 of 3
Concept
Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.
\[ f_k(x) = \sigma(w_k^\top x), \quad \hat{y} = \arg\max_k\, f_k(x) \]
For K=10: exactly 10 classifiers. Fast to train (K logistic models), easy to calibrate, and each classifier's weight vector is interpretable as 'what makes class k different from everything else'.
Analogy
Discussion prompt
Explain One-vs-Rest (OvR) by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.
Concept
Train K(K−1)/2 binary classifiers — one for every pair of classes. Predict by majority vote across all pairwise decisions.
\[ \hat{y} = \text{majority-vote over} \bigl\{f_{ij}(x)\bigr\}_{i<j} \]
| K classes | OvR classifiers | OvO classifiers |
|---|---|---|
| 3 | 3 | 3 |
| 5 | 5 | 10 |
| 10 | 10 | 45 |
OvO is preferred for SVM because each sub-problem is small (only 2-class data). OvR is preferred for logistic regression because calibrated probabilities let you just argmax.
Trade off
Comparison matrix
From One-vs-One (OvO): every row here is a choice with a cost. Fill the OvO classifiers column, then say which row you would actually pick and what you give up for it.
| K classes | OvR classifiers | OvO classifiers |
|---|---|---|
| 3 | 3 | 3 |
| 5 | 5 | 10 |
| 10 | 10 | 45 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others.
For OvR, simply argmax the K scores (no normalization needed for classification).
Why: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others. The K scores don't need to sum to 1 and normalizing them is not mathematically justified.
Trap
Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).
Treat OvR scores as a probability simplex by normalization
Why: Each OvR classifier was trained on a different binary problem — its P(positive|x) is not calibrated to compete against the others. The K scores don't need to sum to 1 and normalizing them is not mathematically justified.
For OvR, simply argmax the K scores (no normalization needed for classification).
ŷ = argmax_k f_k(x) — pick the most confident positive detector
Why: If you need valid probabilities across K classes, use softmax (multinomial logistic) instead. OvR's scores are decision-function outputs, not a joint distribution.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
z_k = w_k^T x. Softmax normalizes them into a valid probability distribution over K classes.; The gradient of softmax cross-entropy w.r.t. the logits has a beautifully simple closed form — one of the key facts for USAAIO.; Train K binary classifiers. Classifier k treats class k as positive and all others as negative. At test time, run all K and pick the class with the highest score.exp(z) directly, then divide by the sum.; Sum the K OvR probabilities and normalize to get a proper distribution: P(y=k|x) = f_k(x) / Σ_j f_j(x).Concept
Multinomial logistic (softmax) trains W jointly — all K weight vectors are updated together through the shared partition function. The K probabilities always sum to 1.
OvR with K sigmoid models trains each binary classifier independently. No joint normalization — each classifier optimizes its own binary CE. Learning dynamics differ: OvR doesn't push down probabilities of wrong classes during a single update.
| property | multinomial | OvR binary |
|---|---|---|
| joint normalization | yes (softmax) | no |
| probabilities sum to 1 | yes | only by luck |
| update coupling | all classes per step | one classifier per problem |
| gradient for true class k | P_k − 1 (softmax CE) | σ(z_k) − 1 (binary BCE) |
| gradient for wrong class j | P_j (pushes down) | no cross-class signal |
Estimation
Predict first
10-class imbalanced dataset (n=1000, class sizes 297→31). Compare OvR logistic, multinomial logistic, OvO SVM — 5-fold CV.
Commit before you compute: what does OvR vs OvO vs Multinomial benchmark come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: OvO SVM wins (+8.6 pp) because SVM's dual is inherently binary — each pairwise sub-problem uses the full kernel margin, not a linearization
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Multinomial and OvR logistic are linear in feature space; OvO SVM-RBF fits non-linear boundaries per pair.
Worked example
10-class imbalanced dataset (n=1000, class sizes 297→31). Compare OvR logistic, multinomial logistic, OvO SVM — 5-fold CV.
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.multiclass import OneVsRestClassifier, OneVsOneClassifier
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
import numpy as np
X, y = make_classification(
n_samples=1000, n_features=20, n_informative=15,
n_classes=10, n_clusters_per_class=1,
weights=[0.3,0.2,0.1,0.08,0.07,0.07,0.06,0.05,0.04,0.03],
random_state=42)
ovr = OneVsRestClassifier(LogisticRegression(max_iter=500, random_state=0))
mn = LogisticRegression(max_iter=500, random_state=0, solver='lbfgs')
ovo = OneVsOneClassifier(SVC(kernel='rbf', random_state=0))
for name, clf in [('OvR LR', ovr), ('MN LR', mn), ('OvO SVM', ovo)]:
s = cross_val_score(clf, X, y, cv=5)
print(f'{name}: {s.mean():.4f} ± {s.std():.4f}')| strategy | 5-fold accuracy | # classifiers |
|---|---|---|
| OvR LogReg | 0.6590 ± 0.0294 | 10 |
| Multinomial LogReg | 0.6590 ± 0.0354 | 1 (K-way) |
| OvO SVM (RBF) | 0.7450 ± 0.0084 | 45 |
OvO SVM wins (+8.6 pp) because SVM's dual is inherently binary — each pairwise sub-problem uses the full kernel margin, not a linearization
Why: Multinomial and OvR logistic are linear in feature space; OvO SVM-RBF fits non-linear boundaries per pair. The accuracy gap will close once you add kernel/polynomial features to logistic.
Comparison
Comparison matrix
From OvR vs OvO vs Multinomial benchmark: refill the 5-fold accuracy column from what you know. The rest of the table is as it appeared.
| strategy | 5-fold accuracy | # classifiers |
|---|---|---|
| OvR LogReg | 0.6590 ± 0.0294 | 10 |
| Multinomial LogReg | 0.6590 ± 0.0354 | 1 (K-way) |
| OvO SVM (RBF) | 0.7450 ± 0.0084 | 45 |
Section
Part 3 of 3
Concept
In binary problems, imbalance biases the decision boundary toward the majority class. In K-class softmax, the rare-class weight vectors receive far fewer gradient updates — they stay near zero while majority-class vectors diverge.
The model learns to never predict rare classes: accuracy appears high (dominated by majority-class predictions) but recall on rare classes is near zero. Macro-F1 or per-class accuracy is the right metric here, not overall accuracy.
Concept
Weighted cross-entropy: multiply each sample's CE contribution by its class weight. Balanced weights = 1/(K · frac_k), so a class with 3% frequency gets weight ≈3.26×.
\[ \mathcal{L}_{\text{weighted}} = -\frac{1}{n}\sum_{i=1}^n w_{y_i} \log P(y_i \mid x_i) \]
Class-balanced sampling: oversample rare classes (e.g. SMOTE) or undersample majority classes so each mini-batch has roughly equal representation. Complementary to weighting — use one or both.
Estimation
Predict first
Compute sklearn balanced class weights and compare weighted vs unweighted multinomial logistic accuracy.
Commit before you compute: what does Balanced weights on 10-class imbalanced data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Weighted accuracy 0.5540 < unweighted 0.6590 — this is expected
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Weighting forces the model to attend to rare classes at the cost of overall accuracy.
Worked example
Compute sklearn balanced class weights and compare weighted vs unweighted multinomial logistic accuracy.
from sklearn.utils.class_weight import compute_class_weight
import numpy as np
cw = compute_class_weight('balanced', classes=np.arange(10), y=y)
print('class | weight')
for i in range(10):
print(f' {i:2d} | {cw[i]:.4f}')
mn_w = LogisticRegression(max_iter=500, class_weight='balanced',
random_state=0, solver='lbfgs')
s_w = cross_val_score(mn_w, X, y, cv=5)
print(f'Weighted MN 5-fold: {s_w.mean():.4f} ± {s_w.std():.4f}')| class | count | balanced weight |
|---|---|---|
| 0 (majority) | 297 | 0.3367 |
| 4 | 72 | 1.3889 |
| 7 | 49 | 2.0408 |
| 9 (minority) | 31 | 3.2258 |
Weighted accuracy 0.5540 < unweighted 0.6590 — this is expected
Why: Weighting forces the model to attend to rare classes at the cost of overall accuracy. The right metric is per-class recall or macro-F1, not micro accuracy. Weighted cross-entropy is a mechanism, not a free accuracy boost.
Pattern
Step through it
Step through Balanced weights on 10-class imbalanced data one row at a time. What is driving the change, and what would the row after the last one be?
Ranking
Put in order
These are the steps of The multiclass logistic recipe, scrambled. Put them back in order before the next slide shows you.
class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracyWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracyEdge cases
Discussion prompt
The multiclass logistic recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
class_weight='balanced' or oversample; evaluate with macro-F1 or per-class recall, not overall accuracyElimination
Eliminate the wrong options
For a softmax model with logits z, the gradient of the cross-entropy loss w.r.t. z_k for the true class k=0 is:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The gradient dL/dz_k = P_k − Y_k. For the true class Y_k = 1, so dL/dz_0 = P_0 − 1. This is negative (pushing the logit up) when P_0 < 1 — the correct direction.
Check
Pen-and-paper first.
Check your understanding
For a softmax model with logits z, the gradient of the cross-entropy loss w.r.t. z_k for the true class k=0 is:
Answer: A
Why: The gradient dL/dz_k = P_k − Y_k. For the true class Y_k = 1, so dL/dz_0 = P_0 − 1. This is negative (pushing the logit up) when P_0 < 1 — the correct direction.
Prediction
Predict first
You want to extend a binary SVM (kernel trick, quadratic program in the dual) to 10 classes. Which reduction strategy is preferred and why?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: OvO — each of 45 pairwise sub-problems is a binary SVM, which the dual is designed for
Why: The SVM dual is inherently a binary QP — its kernel trick and margin maximization assume exactly two classes. OvO decomposes K-class SVM into K(K−1)/2 natural binary SVMs. OvR's K sub-problems use imbalanced negatives (all other classes), which degrades the margin.
Check
Choose the right strategy.
Check your understanding
You want to extend a binary SVM (kernel trick, quadratic program in the dual) to 10 classes. Which reduction strategy is preferred and why?
Answer: A
Why: The SVM dual is inherently a binary QP — its kernel trick and margin maximization assume exactly two classes. OvO decomposes K-class SVM into K(K−1)/2 natural binary SVMs. OvR's K sub-problems use imbalanced negatives (all other classes), which degrades the margin.
Elimination
Eliminate the wrong options
You train a 10-class multinomial logistic model on severely imbalanced data (one class has 3% of samples). After adding class_weight='balanced', overall accuracy drops from 65.9% to 55.4%. What is the most likely correct interpretation?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Balanced weighting intentionally sacrifices majority-class accuracy to improve minority-class recall. Overall accuracy is a micro metric dominated by the majority class — it falls. Macro-F1 (unweighted average over classes) typically improves, which is the real goal when class frequencies differ.
Check
Read the scenario carefully.
Check your understanding
You train a 10-class multinomial logistic model on severely imbalanced data (one class has 3% of samples). After adding class_weight='balanced', overall accuracy drops from 65.9% to 55.4%. What is the most likely correct interpretation?
Answer: A
Why: Balanced weighting intentionally sacrifices majority-class accuracy to improve minority-class recall. Overall accuracy is a micro metric dominated by the majority class — it falls. Macro-F1 (unweighted average over classes) typically improves, which is the real goal when class frequencies differ.
Section
Project
Concept
Build a multinomial logistic classifier from scratch, then compare OvR / OvO / multinomial strategies on an imbalanced 10-class problem.
| milestone | task | key tool |
|---|---|---|
| 1 | Implement softmax + CE forward pass, verify gradient numerically | numpy, finite differences |
| 2 | Train multinomial logistic on Iris (GD loop) | softmax, W update |
| 3 | Benchmark OvR / MN / OvO on 10-class imbalanced data | sklearn multiclass wrappers |
| 4 | Add class_weight='balanced' and report macro-F1 | classification_report |
Build rules: implement your own softmax function with max-shift for stability; verify the gradient before training. In the benchmark, measure macro-F1 alongside accuracy.
Counterexample
Discussion prompt
Build a multinomial logistic classifier from scratch, then compare OvR / OvO / multinomial strategies on an imbalanced 10-class problem.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: implement your own softmax function with max-shift for stability; verify the gradient before training. In the benchmark, measure macro-F1 alongside accuracy.
Worked example
Your turn: implement softmax(Z) and verify that your analytic gradient matches a numerical gradient (finite differences). Target max error < 1e-8.
Hint: stable softmax subtracts Z.max(axis=1, keepdims=True) before exp. Gradient check: perturb W by ±ε, compute (L_plus − L_minus)/(2ε) per parameter.
import numpy as np
def softmax(Z):
E = np.exp(Z - Z.max(axis=1, keepdims=True))
return E / E.sum(axis=1, keepdims=True)
# 5 samples, 3 classes, 4 features
rng = np.random.default_rng(42)
W = rng.normal(size=(4, 3))
X = rng.normal(size=(5, 4))
y = rng.integers(0, 3, size=5)
Y = np.eye(3)[y]
P = softmax(X @ W)
dW_analytic = X.T @ ((P - Y) / len(y))
# numerical gradient
eps = 1e-5; dW_num = np.zeros_like(W)
for i in range(4):
for j in range(3):
def loss(W_):
return -np.sum(Y * np.log(softmax(X@W_)+1e-15))/len(y)
Wp, Wm = W.copy(), W.copy()
Wp[i,j] += eps; Wm[i,j] -= eps
dW_num[i,j] = (loss(Wp)-loss(Wm))/(2*eps)
print('max err:', np.max(np.abs(dW_analytic - dW_num)))| check | result |
|---|---|
| max gradient error | 3.16e-11 |
| passes < 1e-8 threshold | yes |
Trade off
Comparison matrix
From Milestone 1 — softmax forward + gradient check: every row here is a choice with a cost. Fill the result column, then say which row you would actually pick and what you give up for it.
| check | result |
|---|---|
| max gradient error | 3.16e-11 |
| passes < 1e-8 threshold | yes |
Worked example
Your turn: train multinomial logistic on Iris with 500 GD steps (lr=0.1). Predict the final accuracy and loss.
Hint: mean-center X; append a bias column; one-hot Y; update W -= lr * Xb.T @ ((P-Y)/n) each epoch.
from sklearn.datasets import load_iris
iris = load_iris()
X = iris.data - iris.data.mean(0)
Xb = np.c_[X, np.ones(len(X))]
y = iris.target
Y = np.eye(3)[y]
W = np.zeros((Xb.shape[1], 3))
for ep in range(500):
P = softmax(Xb @ W)
L = -np.sum(Y * np.log(P + 1e-15)) / len(y)
W -= 0.1 * Xb.T @ ((P - Y) / len(y))
acc = (np.argmax(Xb @ W, 1) == y).mean()
print(f'acc={acc:.4f} loss={L:.4f}')| epoch | loss | accuracy |
|---|---|---|
| 0 | 1.0986 | — |
| 100 | 0.3193 | — |
| 499 | 0.1647 | 0.9667 |
Comparison
Comparison matrix
From Milestone 2 — multinomial GD on Iris: refill the loss column from what you know. The rest of the table is as it appeared.
| epoch | loss | accuracy |
|---|---|---|
| 0 | 1.0986 | — |
| 100 | 0.3193 | — |
| 499 | 0.1647 | 0.9667 |
Worked example
Your turn: on the 10-class imbalanced dataset, run OvR, multinomial, and OvO SVM. Then add balanced weighting and check whether macro-F1 improves.
Hint: from sklearn.metrics import classification_report; cross_val_score(..., scoring='f1_macro') for macro-F1 across CV folds.
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.multiclass import OneVsRestClassifier, OneVsOneClassifier
from sklearn.svm import SVC
from sklearn.model_selection import cross_val_score
X, y = make_classification(
n_samples=1000, n_features=20, n_informative=15,
n_classes=10, n_clusters_per_class=1,
weights=[0.3,0.2,0.1,0.08,0.07,0.07,0.06,0.05,0.04,0.03],
random_state=42)
ovr = OneVsRestClassifier(LogisticRegression(max_iter=500, random_state=0))
mn = LogisticRegression(max_iter=500, random_state=0, solver='lbfgs')
ovo = OneVsOneClassifier(SVC(kernel='rbf', random_state=0))
mn_w = LogisticRegression(max_iter=500, class_weight='balanced',
random_state=0, solver='lbfgs')
for name, clf in [('OvR LR',ovr),('MN LR',mn),('OvO SVM',ovo),('MN+bal',mn_w)]:
acc = cross_val_score(clf, X, y, cv=5).mean()
print(f'{name}: acc={acc:.4f}')| strategy | 5-fold accuracy |
|---|---|
| OvR LogReg | 0.6590 |
| Multinomial LogReg | 0.6590 |
| OvO SVM (RBF) | 0.7450 |
| Multinomial + balanced | 0.5540 |
Discrimination
Sort into buckets
Sort these by 5-fold accuracy, from memory, without looking back at Milestone 3 — benchmark + macro-F1. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Out loud, slides closed: (1) Derive the softmax CE gradient — state why it equals P − Y. (2) Explain when OvO beats OvR. (3) Why does balanced weighting lower overall accuracy but improve model fairness?
Stretch (homework): add a mini-batch training loop for the multinomial scratch implementation; compare OvR/OvO/multinomial with macro-F1 (not just accuracy); derive the gradient from first principles as the math homework requires.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The softmax distribution · OvR and OvO reduction strategies · Class imbalance in multiclass · Your turn: implement it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| idea | the one thing to remember |
|---|---|
| softmax gradient | dL/dz = P − Y (true class −(1−P), wrong class +P) |
| OvR vs OvO | logistic → OvR; SVM → OvO (binary dual fits naturally) |
| multinomial vs OvR | multinomial couples K-class updates; OvR trains K independent BCE problems |
| imbalance fix | class_weight='balanced'; evaluate macro-F1 not accuracy |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.