USAAIO Lesson 65, from Phase 3. It derives Multinomial, Bernoulli, and Complement Naive Bayes from first principles, covering the derivation of Laplace smoothing, log-space arithmetic to prevent underflow, Multinomial against Bernoulli Naive Bayes on count against binary features, and ComplementNB for imbalanced short texts. It also proves that Naive Bayes is a linear classifier in log space. You implement MultinomialNB from scratch and benchmark it against logistic regression. The lesson runs to 32 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 65 · Phase 3
Bayes' theorem meets counting: derive the MNB update, prevent underflow with log-space arithmetic, and prove that NB is secretly a linear classifier.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 65: Naive Bayes Classifiers: without looking back, what was the main idea of Bag of Words, TF-IDF, and Text Vectorization, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
Bag-of-Words (BoW) count vectors, TF-IDF weighting (formula + scratch implementation), n-grams, text preprocessing pipeline (lowercase / stopwords / stemming vs lemmatization), CountVectorizer and TfidfVectorizer, and a full TF-IDF + logistic regression text-classification pipeline.
Section
Part 1 of 4
Concept
Naive Bayes applies Bayes' theorem with one simplifying assumption: features are conditionally independent given the class.
\[ P(c \mid \mathbf{x}) \propto P(c) \prod_{i=1}^{d} P(x_i \mid c) \]
The 'naive' part is that joint P(x_1,...,x_d|c) is replaced by the product of marginals — rarely true in reality, but the classifier is surprisingly robust anyway.
Counterexample
Discussion prompt
Naive Bayes applies Bayes' theorem with one simplifying assumption: features are conditionally independent given the class.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
The 'naive' part is that joint P(x_1,...,x_d|c) is replaced by the product of marginals — rarely true in reality, but the classifier is surprisingly robust anyway.
Concept
For text, each feature x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.
\[ \hat{P}(w \mid c) = \frac{\text{count}(w, c) + \alpha}{\text{count}(*, c) + \alpha \cdot V} \]
alpha is the Laplace (additive) smoothing parameter; V is the vocabulary size. Setting alpha=0 gives maximum-likelihood estimation — and zero probabilities for unseen words.
Analogy
Discussion prompt
Explain Multinomial NB: the counting rule by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For text, each feature x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.
Intuition
Without smoothing, a word that never appeared in spam training docs gets P(word|spam) = 0. One zero kills the entire product — the class is assigned probability zero regardless of every other word.
Smoothing adds a 'pseudocount' alpha to every word in every class — as if you'd seen each word once before collecting any data. The model stays open to seeing anything.
| word | count in spam | V | P raw | P smoothed (alpha=1) |
|---|---|---|---|---|
| 'free' | 2 | 12 | 2/8 = 0.2500 | 3/20 = 0.1500 |
| 'meeting' | 0 | 12 | 0/8 = 0.0000 | 1/20 = 0.0500 |
| 'lottery' | 1 | 12 | 1/8 = 0.1250 | 2/20 = 0.1000 |
Comparison
Comparison matrix
From Why Laplace smoothing is non-negotiable: refill the count in spam column from what you know. The rest of the table is as it appeared.
| word | count in spam | V | P raw | P smoothed (alpha=1) |
|---|---|---|---|---|
| 'free' | 2 | 12 | 2/8 = 0.2500 | 3/20 = 0.1500 |
| 'meeting' | 0 | 12 | 0/8 = 0.0000 | 1/20 = 0.0500 |
| 'lottery' | 1 | 12 | 1/8 = 0.1250 | 2/20 = 0.1000 |
Estimation
Predict first
Four training docs: two spam (free money prize winner, free lottery win cash), two ham (meeting agenda project, project deadline review meeting). V=12, alpha=1.
Commit before you compute: what does Classify 'free money' — step by step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: class 0 (ham): -6.5820, class 1 (spam): -4.8929
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Argmax picks class 1 — spam. 'free' and 'money' both appear in spam training docs; neither appears in ham training docs.
Worked example
Four training docs: two spam (free money prize winner, free lottery win cash), two ham (meeting agenda project, project deadline review meeting). V=12, alpha=1.
import numpy as np
docs = ['free money prize winner', 'free lottery win cash',
'meeting agenda project', 'project deadline review meeting']
labels = [1, 1, 0, 0] # 1=spam, 0=ham
vocab = sorted(set(' '.join(docs).split()))
V = len(vocab); alpha = 1.0
vocab_idx = {w: i for i, w in enumerate(vocab)}
counts = {c: np.zeros(V) for c in [0, 1]}
for doc, label in zip(docs, labels):
for word in doc.split():
counts[label][vocab_idx[word]] += 1
log_probs = {}
for c in [0, 1]:
total = counts[c].sum()
log_probs[c] = np.log((counts[c] + alpha) / (total + alpha * V))
test = 'free money'
for c in [0, 1]:
score = np.log(0.5) # equal class priors
for w in test.split():
score += log_probs[c][vocab_idx[w]]
print(f'class {c}: {score:.4f}')class 0 (ham): -6.5820, class 1 (spam): -4.8929
Why: Argmax picks class 1 — spam. 'free' and 'money' both appear in spam training docs; neither appears in ham training docs.
| term | log P(w|spam) | log P(w|ham) |
|---|---|---|
| 'free' | log(3/20) = -1.8971 | log(1/19) = -2.9444 |
| 'money' | log(2/20) = -2.3026 | log(1/19) = -2.9444 |
| prior | log(0.5) = -0.6931 | log(0.5) = -0.6931 |
| total | -4.8929 (argmax) | -6.5820 |
Discrimination
Sort into buckets
Sort these by log P(w|ham), from memory, without looking back at Classify 'free money' — step by step. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 2 of 4
Concept
A 500-word email multiplies ~500 probabilities together. Each P(word|class) might be around 0.001. The product underflows to 0.0 before Python can compare the two classes.
\[ 0.001^{50} \approx 10^{-150} \quad\rightarrow\quad \texttt{float64 min} \approx 5 \times 10^{-324} \]
For a 50-word doc, 0.001^50 is already 1e-150 — well inside the danger zone. For 500 words it underflows to exactly 0.0 and comparison breaks.
Explain it
Discussion prompt
Explain Float underflow kills naive NB to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
A 500-word email multiplies ~500 probabilities together. Each P(word|class) might be around 0.001. The product underflows to 0.0 before Python can compare the two classes.
Concept
\[ \log \prod_i P(x_i \mid c) = \sum_i \log P(x_i \mid c) \]
Work entirely in log-space: precompute log P(word|class) once during training, then add them at prediction time. The argmax is identical — logs are monotone.
With 50 words each having log(0.001) = -6.908, the sum is 50 × -6.908 = -345.39 — a perfectly ordinary float64. No underflow, no information loss.
Explain it
Discussion prompt
Explain Log-space: sums instead of products to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Work entirely in log-space: precompute log P(word|class) once during training, then add them at prediction time. The argmax is identical — logs are monotone.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Compute the class score by multiplying raw probabilities:
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Seems like the direct application of Bayes.
Work entirely in log-space from the start:
Why: Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.
Trap
Compute the class score by multiplying raw probabilities:
score = P(c) * P(w1|c) * P(w2|c) * ... * P(w_n|c)
Why: Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.
Work entirely in log-space from the start:
log_score = log P(c) + sum(log P(w_i|c) for w_i in doc)
Why: Logs turn the product into a sum. The argmax is unchanged because log is monotone. No underflow possible — sums of finite floats are always finite.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Logs turn the product into a sum. The argmax is unchanged because log is monotone. No underflow possible — sums of finite floats are always finite.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.
Section
Part 3 of 4
Concept
NB is not one model — it depends on what P(x_i|c) models. Two variants dominate text classification:
| variant | feature x_i | parameter | best for |
|---|---|---|---|
| MultinomialNB | count of word i | P(word|c) via term freq | long docs, bag-of-words |
| BernoulliNB | 1 if word present | P(word present|c) | short docs, binary features |
| ComplementNB | count (complement) | trained on complement | imbalanced classes, short texts |
The difference matters: MNB rewards a word appearing 3 times vs 1 time in a doc. BNB sees only 'present or absent' — the count is discarded.
Trade off
Comparison matrix
From Two models, two feature encodings: every row here is a choice with a cost. Fill the feature x_i column, then say which row you would actually pick and what you give up for it.
| variant | feature x_i | parameter | best for |
|---|---|---|---|
| MultinomialNB | count of word i | P(word|c) via term freq | long docs, bag-of-words |
| BernoulliNB | 1 if word present | P(word present|c) | short docs, binary features |
| ComplementNB | count (complement) | trained on complement | imbalanced classes, short texts |
Estimation
Predict first
Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.
Commit before you compute: what does BNB vs MNB: when frequency matters come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MNB doc0 gets a 3x boost for the repeated 'spam' (feature value 3 vs 1).
Worked example
Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.
from sklearn.feature_extraction.text import CountVectorizer
docs2 = ['spam spam spam prize', 'spam prize', 'meeting project']
vc_count = CountVectorizer()
vc_bin = CountVectorizer(binary=True)
Xc = vc_count.fit_transform(docs2).toarray()
Xb = vc_bin.fit_transform(docs2).toarray()
print('features:', vc_count.get_feature_names_out())
print('MNB counts row0:', Xc[0])
print('BNB binary row0:', Xb[0])MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.
Why: MNB doc0 gets a 3x boost for the repeated 'spam' (feature value 3 vs 1). BNB binarizes — both docs show 'spam present=1'. The frequency signal is discarded.
| doc | MNB 'spam' feature | BNB 'spam' feature |
|---|---|---|
| 'spam spam spam prize' | 3 (contributes 3x log P) | 1 (present once) |
| 'spam prize' | 1 (contributes 1x log P) | 1 (present once) |
| verdict | MNB distinguishes them | BNB: identical |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.
Concept
ComplementNB (Rennie et al. 2003) estimates P(word | NOT c) instead of P(word | c), then predicts the class whose complement best explains the document.
\[ \hat{c} = \arg\min_c \sum_i x_i \cdot \log \hat{P}(w_i \mid \overline{c}) \]
Works better than MNB when classes are imbalanced or docs are short — the complement class typically has more training examples and produces a better-estimated distribution.
Sorting
Sort into buckets
These are the pieces of Lesson 65: Naive Bayes Classifiers, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4 of 4
Concept
For binary classification, the NB decision rule in log-space reduces to a dot product + bias — exactly a linear classifier.
\[ \log \frac{P(c=1 \mid \mathbf{x})}{P(c=0 \mid \mathbf{x})} = \underbrace{(\log \mathbf{p}_1 - \log \mathbf{p}_0)}_{\mathbf{w}} \cdot \mathbf{x} + \underbrace{(\log \pi_1 - \log \pi_0)}_{b} \]
Here p_1[i] = P(word_i | class=1) and p_0[i] = P(word_i | class=0). The weight for word i is the log-ratio of its class-conditional probabilities — a positive weight means the word is evidence for class 1.
Analogy
Discussion prompt
Explain The log-ratio weight vector by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For binary classification, the NB decision rule in log-space reduces to a dot product + bias — exactly a linear classifier.
Estimation
Predict first
Use load_digits (8×8 pixel counts, 10 classes). Extract 2-class slice (digits 0 vs 1). Build weight vector from log-prob difference; confirm it exactly replicates NB predictions.
Commit before you compute: what does Verify NB equals linear score on digits data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: match fraction = 1.0 — linear scores and NB argmax agree on 100% of test examples
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The derivation is exact, not approximate.
Worked example
Use load_digits (8×8 pixel counts, 10 classes). Extract 2-class slice (digits 0 vs 1). Build weight vector from log-prob difference; confirm it exactly replicates NB predictions.
import numpy as np
from sklearn.datasets import load_digits
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
digits = load_digits()
mask = (digits.target <= 1)
X, y = digits.data[mask].astype(int), digits.target[mask]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)
mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
w = mnb.feature_log_prob_[1] - mnb.feature_log_prob_[0]
b = mnb.class_log_prior_[1] - mnb.class_log_prior_[0]
preds_lin = (Xte @ w + b > 0).astype(int)
preds_nb = mnb.predict(Xte)
print('match fraction:', (preds_lin == preds_nb).mean())
print('accuracy:', accuracy_score(yte, preds_nb))match fraction = 1.0 — linear scores and NB argmax agree on 100% of test examples
Why: The derivation is exact, not approximate. NB is a linear classifier in log-space; the weight vector w encodes the log-likelihood ratio for each feature.
| result | value |
|---|---|
| match fraction (linear vs NB) | 1.0000 |
| 2-class NB accuracy (digit 0 vs 1) | 1.0000 |
| w[i] > 0 means | feature i supports digit 1 |
| w[i] < 0 means | feature i supports digit 0 |
Discrimination
Sort into buckets
Sort these by value, from memory, without looking back at Verify NB equals linear score on digits data. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Anomaly
Predict first
A student writes this, and it looks reasonable:
NB is simpler than logistic regression, so it must prefer fewer features to avoid overfitting.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This conflates model complexity with feature dimensionality.
NB estimates parameters independently per feature — it doesn't 'overfit' in the sense of gradient-based models.
Why: This conflates model complexity with feature dimensionality. NB has NO learned weights in the gradient-descent sense — it just counts. Rare words get small smoothed probabilities automatically. Aggressive feature pruning can hurt NB by removing genuine signal.
Trap
NB is simpler than logistic regression, so it must prefer fewer features to avoid overfitting.
Drop rare words before fitting NB to reduce feature dimensionality
Why: This conflates model complexity with feature dimensionality. NB has NO learned weights in the gradient-descent sense — it just counts. Rare words get small smoothed probabilities automatically. Aggressive feature pruning can hurt NB by removing genuine signal.
NB estimates parameters independently per feature — it doesn't 'overfit' in the sense of gradient-based models.
Use Laplace alpha to control smoothing, not aggressive feature pruning
Why: The only hyperparameter that matters is alpha. Higher alpha shrinks all estimates toward uniform — that's the NB regularizer. Feature selection is orthogonal and usually not needed.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.; Smoothing adds a 'pseudocount' alpha to every word in every class — as if you'd seen each word once before collecting any data. The model stays open to seeing anything.Concept
NB training is a single pass over the data: count occurrences per class, divide, take logs. No iterations, no matrix inversions, no gradient descent.
| model | train complexity | predict complexity | typical speedup vs LR |
|---|---|---|---|
| MultinomialNB | O(n * d) | O(d) | 10x–100x faster |
| Logistic Reg | O(n * d * iter) | O(d) | baseline |
| LinearSVC | O(n^2 * d) worst | O(d) | often slower |
For large vocabularies (d = 10 000+) and large corpora, NB trains in milliseconds where LR takes seconds. Prediction cost is the same (both are dot products).
Comparison
Comparison matrix
From Speed advantage: O(nd) train, O(d) predict: refill the typical speedup vs LR column from what you know. The rest of the table is as it appeared.
| model | train complexity | predict complexity | typical speedup vs LR |
|---|---|---|---|
| MultinomialNB | O(n * d) | O(d) | 10x–100x faster |
| Logistic Reg | O(n * d * iter) | O(d) | baseline |
| LinearSVC | O(n^2 * d) worst | O(d) | often slower |
Ranking
Put in order
These are the steps of The Naive Bayes recipe, scrambled. Put them back in order before the next slide shows you.
Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
Edge cases
Discussion prompt
The Naive Bayes recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
A vocabulary has V=1000 words. Class 'spam' has 500 total word tokens in training. A word 'gratis' appears 0 times in spam docs. With Laplace smoothing (alpha=1), what is P('gratis' | spam)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Laplace: P(w|c) = (count + alpha) / (total + alphaV) = (0+1)/(500+11000) = 1/1500 = 0.000667. The denominator adds alpha*V = 1000 pseudocounts, not just one.
Check
Think before clicking.
Check your understanding
A vocabulary has V=1000 words. Class 'spam' has 500 total word tokens in training. A word 'gratis' appears 0 times in spam docs. With Laplace smoothing (alpha=1), what is P('gratis' | spam)?
Answer: A
Why: Laplace: P(w|c) = (count + alpha) / (total + alphaV) = (0+1)/(500+11000) = 1/1500 = 0.000667. The denominator adds alpha*V = 1000 pseudocounts, not just one.
Prediction
Predict first
Why must Naive Bayes predictions use log-space (summing log probs) rather than multiplying raw probabilities?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Multiplying many small probabilities underflows to 0.0, making both classes score zero and breaking argmax
Why: 0.001^50 ≈ 1e-150 and 0.001^500 underflows to 0.0 in float64 (min ≈ 5e-324). When both classes underflow, argmax is undefined. Log-sums avoid this: 50 × log(0.001) = -345.39, always a valid float. The argmax is identical because log is strictly monotone.
Check
Don't estimate — reason from the math.
Check your understanding
Why must Naive Bayes predictions use log-space (summing log probs) rather than multiplying raw probabilities?
Answer: A
Why: 0.001^50 ≈ 1e-150 and 0.001^500 underflows to 0.0 in float64 (min ≈ 5e-324). When both classes underflow, argmax is undefined. Log-sums avoid this: 50 × log(0.001) = -345.39, always a valid float. The argmax is identical because log is strictly monotone.
Elimination
Eliminate the wrong options
Document A contains the word 'discount' 5 times; document B contains it once. MultinomialNB and BernoulliNB are both trained on a corpus. Which statement is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: MNB features are raw counts, so 'discount' contributes log P('discount'|c) × 5 times to A's score. BNB binarizes: 'discount' is simply present (1) in both A and B, contributing log P(present|c) once in both cases — the counts are discarded.
Check
Think about what each model sees.
Check your understanding
Document A contains the word 'discount' 5 times; document B contains it once. MultinomialNB and BernoulliNB are both trained on a corpus. Which statement is correct?
Answer: A
Why: MNB features are raw counts, so 'discount' contributes log P('discount'|c) × 5 times to A's score. BNB binarizes: 'discount' is simply present (1) in both A and B, contributing log P(present|c) once in both cases — the counts are discarded.
Section
Project
Concept
Implement a ScratchMNB class that matches sklearn's MultinomialNB on the digits dataset (treating pixel intensities as word counts).
| # | milestone | goal |
|---|---|---|
| 1 | Count table + Laplace | Build log_prob_ matrix of shape (n_classes, n_features) |
| 2 | Predict via log-space | log prior + X @ log_prob_.T → argmax |
| 3 | Verify vs sklearn | accuracy matches to 3 decimal places; predictions identical |
Build rules: one forward pass only (no loops over samples at predict time), alpha=1.0, equal class priors estimated from training labels.
Counterexample
Discussion prompt
Implement a ScratchMNB class that matches sklearn's MultinomialNB on the digits dataset (treating pixel intensities as word counts).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: one forward pass only (no loops over samples at predict time), alpha=1.0, equal class priors estimated from training labels.
Worked example
Your turn: load digits, split 70/30, build the per-class count array and apply Laplace smoothing. Predict what log_prob_[0, :].sum() will be (approximately).
Hint: counts[c] = X[y==c].sum(axis=0) + alpha; then log_prob_[c] = log(counts[c] / counts[c].sum()).
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
digits = load_digits()
Xd, yd = digits.data.astype(int), digits.target
Xtr, Xte, ytr, yte = train_test_split(Xd, yd, test_size=0.3, random_state=42)
alpha = 1.0
classes = np.unique(ytr)
n_feats = Xtr.shape[1]
log_prob = np.zeros((len(classes), n_feats))
for i, c in enumerate(classes):
cnts = Xtr[ytr==c].sum(axis=0) + alpha
log_prob[i] = np.log(cnts / cnts.sum())
print('log_prob shape:', log_prob.shape)
print('log_prob[0] sum:', round(log_prob[0].sum(), 4))| output | value |
|---|---|
| log_prob shape | (10, 64) |
| log_prob[0].sum() | a negative float (64 log-prob values summed; varies by class) |
| rows | one per digit class (0-9) |
| cols | one per pixel feature |
Worked example
Your turn: compute log priors from ytr, then predict via scores = X @ log_prob.T + log_prior. Predict the accuracy before running.
Hint: log_prior[c] = log(count(c in ytr) / len(ytr)); predicted = classes[argmax(scores, axis=1)].
from sklearn.metrics import accuracy_score
log_prior = np.log(np.array([(ytr==c).sum()/len(ytr) for c in classes]))
scores = Xte @ log_prob.T + log_prior
preds_scratch = classes[np.argmax(scores, axis=1)]
print('scratch accuracy:', accuracy_score(yte, preds_scratch))| metric | value |
|---|---|
| scratch NB accuracy on digits | 0.8944 |
| n_test samples | 540 |
| log_prior (balanced) | log(1/10) each = -2.3026 |
Worked example
Your turn: fit MultinomialNB(alpha=1.0) on the same split and compare accuracy and per-sample predictions. Predict whether they match exactly.
Hint: (preds_scratch == mnb.predict(Xte)).mean() should be 1.0.
from sklearn.naive_bayes import MultinomialNB
mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
acc_sk = accuracy_score(yte, mnb.predict(Xte))
print('sklearn accuracy:', acc_sk)
print('match fraction:',
(preds_scratch == mnb.predict(Xte)).mean())| model | accuracy | predictions match scratch |
|---|---|---|
| ScratchMNB | 0.8944 | — |
| sklearn MNB | 0.8944 | 1.0000 (exact match) |
| verdict | implementation correct |
Trade off
Comparison matrix
From Milestone 3 — verify against sklearn: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.
| model | accuracy | predictions match scratch |
|---|---|---|
| ScratchMNB | 0.8944 | — |
| sklearn MNB | 0.8944 | 1.0000 (exact match) |
| verdict | implementation correct |
Concept
import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import accuracy_score
digits = load_digits()
Xd, yd = digits.data.astype(int), digits.target
Xtr, Xte, ytr, yte = train_test_split(Xd, yd, test_size=0.3, random_state=42)
alpha = 1.0
classes = np.unique(ytr)
log_prob = np.zeros((len(classes), Xtr.shape[1]))
for i, c in enumerate(classes):
cnts = Xtr[ytr==c].sum(axis=0) + alpha
log_prob[i] = np.log(cnts / cnts.sum())
log_prior = np.log(np.array([(ytr==c).sum()/len(ytr) for c in classes]))
preds = classes[np.argmax(Xte @ log_prob.T + log_prior, axis=1)]
mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
print('scratch:', accuracy_score(yte, preds))
print('sklearn:', accuracy_score(yte, mnb.predict(Xte)))
print('match:', (preds == mnb.predict(Xte)).mean())| output | value |
|---|---|
| scratch accuracy | 0.8944 |
| sklearn accuracy | 0.8944 |
| match fraction | 1.0000 |
If your match fraction is 1.0 — you've implemented Naive Bayes from Bayes' theorem to working code. (Lesson 40 trained a linear model with gradient descent; NB gets there in a single pass.)
Comparison
Comparison matrix
From The full program: refill the value column from what you know. The rest of the table is as it appeared.
| output | value |
|---|---|
| scratch accuracy | 0.8944 |
| sklearn accuracy | 0.8944 |
| match fraction | 1.0000 |
Concept
Out loud, slides closed: (1) derive the Laplace smoothing formula from scratch and explain why alpha goes in both numerator and denominator, (2) explain why NB is immune to float underflow when implemented in log-space, (3) prove that binary MNB is a linear classifier by writing out the weight vector.
Stretch (homework): implement the spam classifier from the Lesson 65 worked example using the USAAIO docs corpus; compute F1 and inspect false positives. Next: logistic regression also becomes linear in log-space — but it learns weights via gradient descent (Lesson 40) rather than counting.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — The Bayes update for text · Log-space arithmetic · Bernoulli NB vs Multinomial NB · NB is a linear classifier · Your turn: build NB from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
(count + alpha) / (total + alpha * V)| idea | the one thing to remember |
|---|---|
| Laplace smoothing | add alpha to numerator, alpha*V to denominator |
| Log-space | sum log probs — never multiply raw probs |
| MNB vs BNB | MNB counts; BNB sees only presence |
| NB speed | O(nd) train, O(d) predict — one pass, no iterations |
| NB is linear | w = log p_1 - log p_0; same decision boundary as LR |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.