Lesson 65: Naive Bayes Classifiers

USAAIO Lesson 65, from Phase 3. It derives Multinomial, Bernoulli, and Complement Naive Bayes from first principles, covering the derivation of Laplace smoothing, log-space arithmetic to prevent underflow, Multinomial against Bernoulli Naive Bayes on count against binary features, and ComplementNB for imbalanced short texts. It also proves that Naive Bayes is a linear classifier in log space. You implement MultinomialNB from scratch and benchmark it against logistic regression. The lesson runs to 32 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Naive Bayes from First Principles

Title

USAAIO · Lesson 65 · Phase 3

Bayes' theorem meets counting: derive the MNB update, prevent underflow with log-space arithmetic, and prove that NB is secretly a linear classifier.

2. By the end of this lesson you can

Objectives

  1. Derive the Multinomial NB update from Bayes' theorem and apply Laplace smoothing
  2. Implement MultinomialNB from scratch — count table → log probabilities → argmax prediction
  3. Explain log-space computation and why it is non-negotiable for long documents
  4. Distinguish Bernoulli NB (binary presence) from Multinomial NB (term frequency)
  5. Prove NB is a linear classifier in log-space and interpret the weight vector

3. What survived from Bag of Words, TF-IDF, and Text Vectorization?

Warm-up

Discussion prompt

Before we open Lesson 65: Naive Bayes Classifiers: without looking back, what was the main idea of Bag of Words, TF-IDF, and Text Vectorization, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Bag-of-Words (BoW) count vectors, TF-IDF weighting (formula + scratch implementation), n-grams, text preprocessing pipeline (lowercase / stopwords / stemming vs lemmatization), CountVectorizer and TfidfVectorizer, and a full TF-IDF + logistic regression text-classification pipeline.

4. The Bayes update for text

Section

Part 1 of 4

5. Bayes' theorem for classification

Concept

Naive Bayes applies Bayes' theorem with one simplifying assumption: features are conditionally independent given the class.

\[ P(c \mid \mathbf{x}) \propto P(c) \prod_{i=1}^{d} P(x_i \mid c) \]

The 'naive' part is that joint P(x_1,...,x_d|c) is replaced by the product of marginals — rarely true in reality, but the classifier is surprisingly robust anyway.

6. Break it if you can: Bayes' theorem for classification

Counterexample

Discussion prompt

Naive Bayes applies Bayes' theorem with one simplifying assumption: features are conditionally independent given the class.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

The 'naive' part is that joint P(x_1,...,x_d|c) is replaced by the product of marginals — rarely true in reality, but the classifier is surprisingly robust anyway.

7. Multinomial NB: the counting rule

Concept

For text, each feature x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.

\[ \hat{P}(w \mid c) = \frac{\text{count}(w, c) + \alpha}{\text{count}(*, c) + \alpha \cdot V} \]

alpha is the Laplace (additive) smoothing parameter; V is the vocabulary size. Setting alpha=0 gives maximum-likelihood estimation — and zero probabilities for unseen words.

8. By analogy: Multinomial NB: the counting rule

Analogy

Discussion prompt

Explain Multinomial NB: the counting rule by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

For text, each feature x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.

9. Why Laplace smoothing is non-negotiable

Intuition

Without smoothing, a word that never appeared in spam training docs gets P(word|spam) = 0. One zero kills the entire product — the class is assigned probability zero regardless of every other word.

Smoothing adds a 'pseudocount' alpha to every word in every class — as if you'd seen each word once before collecting any data. The model stays open to seeing anything.

wordcount in spamVP rawP smoothed (alpha=1)
'free'2122/8 = 0.25003/20 = 0.1500
'meeting'0120/8 = 0.00001/20 = 0.0500
'lottery'1121/8 = 0.12502/20 = 0.1000

10. Fill in: count in spam for Why Laplace smoothing is non-negotiable

Comparison

Comparison matrix

From Why Laplace smoothing is non-negotiable: refill the count in spam column from what you know. The rest of the table is as it appeared.

wordcount in spamVP rawP smoothed (alpha=1)
'free'2122/8 = 0.25003/20 = 0.1500
'meeting'0120/8 = 0.00001/20 = 0.0500
'lottery'1121/8 = 0.12502/20 = 0.1000

11. Guess the shape of the answer: Classify 'free money' — step by step

Estimation

Predict first

Four training docs: two spam (free money prize winner, free lottery win cash), two ham (meeting agenda project, project deadline review meeting). V=12, alpha=1.

Commit before you compute: what does Classify 'free money' — step by step come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: class 0 (ham): -6.5820, class 1 (spam): -4.8929

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Argmax picks class 1 — spam. 'free' and 'money' both appear in spam training docs; neither appears in ham training docs.

12. Classify 'free money' — step by step

Worked example

Four training docs: two spam (free money prize winner, free lottery win cash), two ham (meeting agenda project, project deadline review meeting). V=12, alpha=1.

import numpy as np

docs   = ['free money prize winner', 'free lottery win cash',
          'meeting agenda project',  'project deadline review meeting']
labels = [1, 1, 0, 0]   # 1=spam, 0=ham
vocab  = sorted(set(' '.join(docs).split()))
V = len(vocab);  alpha = 1.0
vocab_idx = {w: i for i, w in enumerate(vocab)}

counts = {c: np.zeros(V) for c in [0, 1]}
for doc, label in zip(docs, labels):
    for word in doc.split():
        counts[label][vocab_idx[word]] += 1

log_probs = {}
for c in [0, 1]:
    total = counts[c].sum()
    log_probs[c] = np.log((counts[c] + alpha) / (total + alpha * V))

test = 'free money'
for c in [0, 1]:
    score = np.log(0.5)   # equal class priors
    for w in test.split():
        score += log_probs[c][vocab_idx[w]]
    print(f'class {c}: {score:.4f}')

class 0 (ham): -6.5820, class 1 (spam): -4.8929

Why: Argmax picks class 1 — spam. 'free' and 'money' both appear in spam training docs; neither appears in ham training docs.

termlog P(w|spam)log P(w|ham)
'free'log(3/20) = -1.8971log(1/19) = -2.9444
'money'log(2/20) = -2.3026log(1/19) = -2.9444
priorlog(0.5) = -0.6931log(0.5) = -0.6931
total-4.8929 (argmax)-6.5820

13. Which is which, by log P(w|ham)

Discrimination

Sort into buckets

Sort these by log P(w|ham), from memory, without looking back at Classify 'free money' — step by step. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

log(1/19) = -2.9444
'free'; 'money'
log(0.5) = -0.6931
prior
-6.5820
total
g1
log P(w|ham) is "log(1/19) = -2.9444" for 'free', 'money' — that is what the table on "Classify 'free money' — step by step" records, and it is the single property separating this group from the rest.
g2
log P(w|ham) is "log(0.5) = -0.6931" for prior — that is what the table on "Classify 'free money' — step by step" records, and it is the single property separating this group from the rest.
g3
log P(w|ham) is "-6.5820" for total — that is what the table on "Classify 'free money' — step by step" records, and it is the single property separating this group from the rest.

14. Log-space arithmetic

Section

Part 2 of 4

15. Float underflow kills naive NB

Concept

A 500-word email multiplies ~500 probabilities together. Each P(word|class) might be around 0.001. The product underflows to 0.0 before Python can compare the two classes.

\[ 0.001^{50} \approx 10^{-150} \quad\rightarrow\quad \texttt{float64 min} \approx 5 \times 10^{-324} \]

For a 50-word doc, 0.001^50 is already 1e-150 — well inside the danger zone. For 500 words it underflows to exactly 0.0 and comparison breaks.

16. Teach it back: Float underflow kills naive NB

Explain it

Discussion prompt

Explain Float underflow kills naive NB to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

A 500-word email multiplies ~500 probabilities together. Each P(word|class) might be around 0.001. The product underflows to 0.0 before Python can compare the two classes.

17. Log-space: sums instead of products

Concept

\[ \log \prod_i P(x_i \mid c) = \sum_i \log P(x_i \mid c) \]

Work entirely in log-space: precompute log P(word|class) once during training, then add them at prediction time. The argmax is identical — logs are monotone.

With 50 words each having log(0.001) = -6.908, the sum is 50 × -6.908 = -345.39 — a perfectly ordinary float64. No underflow, no information loss.

18. Teach it back: Log-space: sums instead of products

Explain it

Discussion prompt

Explain Log-space: sums instead of products to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

Work entirely in log-space: precompute log P(word|class) once during training, then add them at prediction time. The argmax is identical — logs are monotone.

19. Something is wrong here: multiplying raw probabilities in prediction

Anomaly

Predict first

A student writes this, and it looks reasonable:

Compute the class score by multiplying raw probabilities:

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Seems like the direct application of Bayes.

Work entirely in log-space from the start:

Why: Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.

20. Trap: multiplying raw probabilities in prediction

Trap

The trap

Compute the class score by multiplying raw probabilities:

score = P(c) * P(w1|c) * P(w2|c) * ... * P(w_n|c)

Why: Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.

The fix

Work entirely in log-space from the start:

log_score = log P(c) + sum(log P(w_i|c) for w_i in doc)

Why: Logs turn the product into a sum. The argmax is unchanged because log is monotone. No underflow possible — sums of finite floats are always finite.

21. Break it on purpose: multiplying raw probabilities in prediction

Break the constraint

Discussion prompt

The rule this trap just fixed:

Logs turn the product into a sum. The argmax is unchanged because log is monotone. No underflow possible — sums of finite floats are always finite.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Seems like the direct application of Bayes. But each factor is a float < 1, and 500 multiplications silently underflow to 0.0 — both classes score 0.0 and prediction is undefined or wrong.

22. Bernoulli NB vs Multinomial NB

Section

Part 3 of 4

23. Two models, two feature encodings

Concept

NB is not one model — it depends on what P(x_i|c) models. Two variants dominate text classification:

variantfeature x_iparameterbest for
MultinomialNBcount of word iP(word|c) via term freqlong docs, bag-of-words
BernoulliNB1 if word presentP(word present|c)short docs, binary features
ComplementNBcount (complement)trained on complementimbalanced classes, short texts

The difference matters: MNB rewards a word appearing 3 times vs 1 time in a doc. BNB sees only 'present or absent' — the count is discarded.

24. What each one costs: Two models, two feature encodings

Trade off

Comparison matrix

From Two models, two feature encodings: every row here is a choice with a cost. Fill the feature x_i column, then say which row you would actually pick and what you give up for it.

variantfeature x_iparameterbest for
MultinomialNBcount of word iP(word|c) via term freqlong docs, bag-of-words
BernoulliNB1 if word presentP(word present|c)short docs, binary features
ComplementNBcount (complement)trained on complementimbalanced classes, short texts

25. Guess the shape of the answer: BNB vs MNB: when frequency matters

Estimation

Predict first

Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.

Commit before you compute: what does BNB vs MNB: when frequency matters come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. MNB doc0 gets a 3x boost for the repeated 'spam' (feature value 3 vs 1).

26. BNB vs MNB: when frequency matters

Worked example

Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.

from sklearn.feature_extraction.text import CountVectorizer

docs2 = ['spam spam spam prize', 'spam prize', 'meeting project']
vc_count = CountVectorizer()
vc_bin   = CountVectorizer(binary=True)
Xc = vc_count.fit_transform(docs2).toarray()
Xb = vc_bin.fit_transform(docs2).toarray()
print('features:', vc_count.get_feature_names_out())
print('MNB counts row0:', Xc[0])
print('BNB binary row0:', Xb[0])

MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.

Why: MNB doc0 gets a 3x boost for the repeated 'spam' (feature value 3 vs 1). BNB binarizes — both docs show 'spam present=1'. The frequency signal is discarded.

docMNB 'spam' featureBNB 'spam' feature
'spam spam spam prize'3 (contributes 3x log P)1 (present once)
'spam prize'1 (contributes 1x log P)1 (present once)
verdictMNB distinguishes themBNB: identical

27. Work backwards from the answer: BNB vs MNB: when frequency matters

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

MNB row0 'spam' feature: count=3. BNB row0 'spam' feature: 1.

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Consider doc_A = 'spam spam spam prize'. MNB sees 'spam' with count=3 and weights it 3x. BNB sees presence=1 — same signal as doc_B = 'spam prize'.

28. Complement NB: train on what a class is NOT

Concept

ComplementNB (Rennie et al. 2003) estimates P(word | NOT c) instead of P(word | c), then predicts the class whose complement best explains the document.

\[ \hat{c} = \arg\min_c \sum_i x_i \cdot \log \hat{P}(w_i \mid \overline{c}) \]

Works better than MNB when classes are imbalanced or docs are short — the complement class typically has more training examples and produces a better-estimated distribution.

29. Where does each piece belong: Lesson 65: Naive Bayes Classifiers

Sorting

Sort into buckets

These are the pieces of Lesson 65: Naive Bayes Classifiers, out of order. Put each one back under the part of the lesson it belongs to.

The Bayes update for text
Bayes' theorem for classification; Multinomial NB: the counting rule; Why Laplace smoothing is non-negotiable
Log-space arithmetic
Float underflow kills naive NB; Log-space: sums instead of products
Bernoulli NB vs Multinomial NB
Two models, two feature encodings; BNB vs MNB: when frequency matters; Complement NB: train on what a class is NOT
s1
The Bayes update for text is where Lesson 65: Naive Bayes Classifiers puts Bayes' theorem for classification, Multinomial NB: the counting rule, Why Laplace smoothing is non-negotiable. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Log-space arithmetic is where Lesson 65: Naive Bayes Classifiers puts Float underflow kills naive NB, Log-space: sums instead of products. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Bernoulli NB vs Multinomial NB is where Lesson 65: Naive Bayes Classifiers puts Two models, two feature encodings, BNB vs MNB: when frequency matters, Complement NB: train on what a class is NOT. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

30. NB is a linear classifier

Section

Part 4 of 4

31. The log-ratio weight vector

Concept

For binary classification, the NB decision rule in log-space reduces to a dot product + bias — exactly a linear classifier.

\[ \log \frac{P(c=1 \mid \mathbf{x})}{P(c=0 \mid \mathbf{x})} = \underbrace{(\log \mathbf{p}_1 - \log \mathbf{p}_0)}_{\mathbf{w}} \cdot \mathbf{x} + \underbrace{(\log \pi_1 - \log \pi_0)}_{b} \]

Here p_1[i] = P(word_i | class=1) and p_0[i] = P(word_i | class=0). The weight for word i is the log-ratio of its class-conditional probabilities — a positive weight means the word is evidence for class 1.

32. By analogy: The log-ratio weight vector

Analogy

Discussion prompt

Explain The log-ratio weight vector by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

For binary classification, the NB decision rule in log-space reduces to a dot product + bias — exactly a linear classifier.

33. Guess the shape of the answer: Verify NB equals linear score on digits data

Estimation

Predict first

Use load_digits (8×8 pixel counts, 10 classes). Extract 2-class slice (digits 0 vs 1). Build weight vector from log-prob difference; confirm it exactly replicates NB predictions.

Commit before you compute: what does Verify NB equals linear score on digits data come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: match fraction = 1.0 — linear scores and NB argmax agree on 100% of test examples

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The derivation is exact, not approximate.

34. Verify NB equals linear score on digits data

Worked example

Use load_digits (8×8 pixel counts, 10 classes). Extract 2-class slice (digits 0 vs 1). Build weight vector from log-prob difference; confirm it exactly replicates NB predictions.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

digits = load_digits()
mask = (digits.target <= 1)
X, y = digits.data[mask].astype(int), digits.target[mask]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=42)

mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
w = mnb.feature_log_prob_[1] - mnb.feature_log_prob_[0]
b = mnb.class_log_prior_[1]  - mnb.class_log_prior_[0]
preds_lin = (Xte @ w + b > 0).astype(int)
preds_nb  = mnb.predict(Xte)
print('match fraction:', (preds_lin == preds_nb).mean())
print('accuracy:', accuracy_score(yte, preds_nb))

match fraction = 1.0 — linear scores and NB argmax agree on 100% of test examples

Why: The derivation is exact, not approximate. NB is a linear classifier in log-space; the weight vector w encodes the log-likelihood ratio for each feature.

resultvalue
match fraction (linear vs NB)1.0000
2-class NB accuracy (digit 0 vs 1)1.0000
w[i] > 0 meansfeature i supports digit 1
w[i] < 0 meansfeature i supports digit 0

35. Which is which, by value

Discrimination

Sort into buckets

Sort these by value, from memory, without looking back at Verify NB equals linear score on digits data. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1.0000
match fraction (linear vs NB); 2-class NB accuracy (digit 0 vs 1)
feature i supports digit 1
w[i] > 0 means
feature i supports digit 0
w[i] < 0 means
g1
value is "1.0000" for match fraction (linear vs NB), 2-class NB accuracy (digit 0 vs 1) — that is what the table on "Verify NB equals linear score on digits…" records, and it is the single property separating this group from the rest.
g2
value is "feature i supports digit 1" for w[i] > 0 means — that is what the table on "Verify NB equals linear score on digits…" records, and it is the single property separating this group from the rest.
g3
value is "feature i supports digit 0" for w[i] < 0 means — that is what the table on "Verify NB equals linear score on digits…" records, and it is the single property separating this group from the rest.

36. Something is wrong here: assuming NB needs fewer features than logistic…

Anomaly

Predict first

A student writes this, and it looks reasonable:

NB is simpler than logistic regression, so it must prefer fewer features to avoid overfitting.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This conflates model complexity with feature dimensionality.

NB estimates parameters independently per feature — it doesn't 'overfit' in the sense of gradient-based models.

Why: This conflates model complexity with feature dimensionality. NB has NO learned weights in the gradient-descent sense — it just counts. Rare words get small smoothed probabilities automatically. Aggressive feature pruning can hurt NB by removing genuine signal.

37. Trap: assuming NB needs fewer features than logistic regression

Trap

The trap

NB is simpler than logistic regression, so it must prefer fewer features to avoid overfitting.

Drop rare words before fitting NB to reduce feature dimensionality

Why: This conflates model complexity with feature dimensionality. NB has NO learned weights in the gradient-descent sense — it just counts. Rare words get small smoothed probabilities automatically. Aggressive feature pruning can hurt NB by removing genuine signal.

The fix

NB estimates parameters independently per feature — it doesn't 'overfit' in the sense of gradient-based models.

Use Laplace alpha to control smoothing, not aggressive feature pruning

Why: The only hyperparameter that matters is alpha. Higher alpha shrinks all estimates toward uniform — that's the NB regularizer. Feature selection is orthogonal and usually not needed.

38. Which of these survive contact with Lesson 65: Naive Bayes Classifiers?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Naive Bayes applies Bayes' theorem with one simplifying assumption: features are conditionally independent given the class.; For text, each feature x_i is the count of word i in the document. The parameter to estimate is P(word | class) — estimated by counting.; Smoothing adds a 'pseudocount' alpha to every word in every class — as if you'd seen each word once before collecting any data. The model stays open to seeing anything.
Breaks
Compute the class score by multiplying raw probabilities:; NB is simpler than logistic regression, so it must prefer fewer features to avoid overfitting.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 65: Naive Bayes Classifiers puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

39. Speed advantage: O(nd) train, O(d) predict

Concept

NB training is a single pass over the data: count occurrences per class, divide, take logs. No iterations, no matrix inversions, no gradient descent.

modeltrain complexitypredict complexitytypical speedup vs LR
MultinomialNBO(n * d)O(d)10x–100x faster
Logistic RegO(n * d * iter)O(d)baseline
LinearSVCO(n^2 * d) worstO(d)often slower

For large vocabularies (d = 10 000+) and large corpora, NB trains in milliseconds where LR takes seconds. Prediction cost is the same (both are dot products).

40. Fill in: typical speedup vs LR for Speed advantage: O(nd) train, O(d) predict

Comparison

Comparison matrix

From Speed advantage: O(nd) train, O(d) predict: refill the typical speedup vs LR column from what you know. The rest of the table is as it appeared.

modeltrain complexitypredict complexitytypical speedup vs LR
MultinomialNBO(n * d)O(d)10x–100x faster
Logistic RegO(n * d * iter)O(d)baseline
LinearSVCO(n^2 * d) worstO(d)often slower

41. Rebuild the recipe: The Naive Bayes recipe

Ranking

Put in order

These are the steps of The Naive Bayes recipe, scrambled. Put them back in order before the next slide shows you.

  1. Choose variant: MNB for term-frequency, BNB for binary presence, CNB for imbalanced/short texts
  2. Train: count(word, class), add alpha (Laplace), divide by (total + alpha*V), take log
  3. Predict: log prior + sum of log P(word|class) for words in doc → argmax
  4. Never multiply raw probs — always sum in log-space to prevent underflow
  5. Linear equivalence: for binary, NB weight_i = log P(w|c1) - log P(w|c0)

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

42. The Naive Bayes recipe

Pattern

  1. Choose variant: MNB for term-frequency, BNB for binary presence, CNB for imbalanced/short texts
  2. Train: count(word, class), add alpha (Laplace), divide by (total + alpha*V), take log
  3. Predict: log prior + sum of log P(word|class) for words in doc → argmax
  4. Never multiply raw probs — always sum in log-space to prevent underflow
  5. Linear equivalence: for binary, NB weight_i = log P(w|c1) - log P(w|c0)

43. Where does it stop working: The Naive Bayes recipe

Edge cases

Discussion prompt

The Naive Bayes recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Choose variant: MNB for term-frequency, BNB for binary presence, CNB for imbalanced/short texts
  2. Train: count(word, class), add alpha (Laplace), divide by (total + alpha*V), take log
  3. Predict: log prior + sum of log P(word|class) for words in doc → argmax
  4. Never multiply raw probs — always sum in log-space to prevent underflow
  5. Linear equivalence: for binary, NB weight_i = log P(w|c1) - log P(w|c0)

44. Rule out three: Check yourself — Laplace smoothing

Elimination

Eliminate the wrong options

A vocabulary has V=1000 words. Class 'spam' has 500 total word tokens in training. A word 'gratis' appears 0 times in spam docs. With Laplace smoothing (alpha=1), what is P('gratis' | spam)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 1 / 1500 = 0.000667
  • B. 0 / 500 = 0.000
  • C. 1 / 500 = 0.002
  • D. 1 / 1000 = 0.001

Survives elimination: A

Why: Laplace: P(w|c) = (count + alpha) / (total + alphaV) = (0+1)/(500+11000) = 1/1500 = 0.000667. The denominator adds alpha*V = 1000 pseudocounts, not just one.

45. Check yourself — Laplace smoothing

Check

Think before clicking.

Check your understanding

A vocabulary has V=1000 words. Class 'spam' has 500 total word tokens in training. A word 'gratis' appears 0 times in spam docs. With Laplace smoothing (alpha=1), what is P('gratis' | spam)?

  • A. 1 / 1500 = 0.000667 (correct)
  • B. 0 / 500 = 0.000
  • C. 1 / 500 = 0.002
  • D. 1 / 1000 = 0.001

Answer: A

Why: Laplace: P(w|c) = (count + alpha) / (total + alphaV) = (0+1)/(500+11000) = 1/1500 = 0.000667. The denominator adds alpha*V = 1000 pseudocounts, not just one.

Why B tempts people
This is the unsmoothed MLE estimate — it gives a zero probability that kills the entire class score for any document containing 'gratis'.
Why C tempts people
Adds 1 to the numerator but forgets to add alphaV to the denominator. The formula is (count+alpha)/(total+alphaV), not (count+alpha)/total.
Why D tempts people
Uses vocabulary size V in the denominator alone, ignoring the existing token count. The correct denominator is total + alpha*V = 500 + 1000 = 1500.

46. Answer it before you see the options: Check yourself — log-space

Prediction

Predict first

Why must Naive Bayes predictions use log-space (summing log probs) rather than multiplying raw probabilities?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Multiplying many small probabilities underflows to 0.0, making both classes score zero and breaking argmax

Why: 0.001^50 ≈ 1e-150 and 0.001^500 underflows to 0.0 in float64 (min ≈ 5e-324). When both classes underflow, argmax is undefined. Log-sums avoid this: 50 × log(0.001) = -345.39, always a valid float. The argmax is identical because log is strictly monotone.

47. Check yourself — log-space

Check

Don't estimate — reason from the math.

Check your understanding

Why must Naive Bayes predictions use log-space (summing log probs) rather than multiplying raw probabilities?

  • A. Multiplying many small probabilities underflows to 0.0, making both classes score zero and breaking argmax (correct)
  • B. Log probabilities are always larger than raw probabilities, so the argmax changes
  • C. Multiplying probabilities violates the conditional independence assumption
  • D. Float multiplication is slower than float addition, so log-space is a speed optimization

Answer: A

Why: 0.001^50 ≈ 1e-150 and 0.001^500 underflows to 0.0 in float64 (min ≈ 5e-324). When both classes underflow, argmax is undefined. Log-sums avoid this: 50 × log(0.001) = -345.39, always a valid float. The argmax is identical because log is strictly monotone.

Why B tempts people
Log probabilities are negative (since probs < 1), so they are SMALLER (more negative) than the raw probabilities. The argmax is preserved precisely because log is monotone, not because the values are larger.
Why C tempts people
The conditional independence assumption is a modeling choice about the joint distribution — it's unrelated to whether you work in probability space or log-space.
Why D tempts people
Speed is a minor side effect. The primary and critical reason is numerical stability — underflow silently destroys the computation, not just slows it.

48. Rule out three: Check yourself — MNB vs BNB

Elimination

Eliminate the wrong options

Document A contains the word 'discount' 5 times; document B contains it once. MultinomialNB and BernoulliNB are both trained on a corpus. Which statement is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. MNB scores 'discount' 5x as strongly in A vs B; BNB gives identical scores for A and B on that word
  • B. Both MNB and BNB score A higher than B because count matters to both
  • C. BNB scores A higher because repeated words get a higher binary feature value
  • D. MNB and BNB give identical predictions because they share the same log-probability table

Survives elimination: A

Why: MNB features are raw counts, so 'discount' contributes log P('discount'|c) × 5 times to A's score. BNB binarizes: 'discount' is simply present (1) in both A and B, contributing log P(present|c) once in both cases — the counts are discarded.

49. Check yourself — MNB vs BNB

Check

Think about what each model sees.

Check your understanding

Document A contains the word 'discount' 5 times; document B contains it once. MultinomialNB and BernoulliNB are both trained on a corpus. Which statement is correct?

  • A. MNB scores 'discount' 5x as strongly in A vs B; BNB gives identical scores for A and B on that word (correct)
  • B. Both MNB and BNB score A higher than B because count matters to both
  • C. BNB scores A higher because repeated words get a higher binary feature value
  • D. MNB and BNB give identical predictions because they share the same log-probability table

Answer: A

Why: MNB features are raw counts, so 'discount' contributes log P('discount'|c) × 5 times to A's score. BNB binarizes: 'discount' is simply present (1) in both A and B, contributing log P(present|c) once in both cases — the counts are discarded.

Why B tempts people
BNB binarizes its input (CountVectorizer with binary=True), so repeated occurrences collapse to 1. Count does NOT matter to BNB.
Why C tempts people
BNB features are binary (0 or 1) — there is no 'higher binary feature value' for 5 occurrences vs 1. Once the word is present, the feature is 1 regardless.
Why D tempts people
MNB and BNB have entirely different likelihood tables: MNB estimates P(count|c) via term frequency; BNB estimates P(present|c) via binary occurrence rate.

50. Your turn: build NB from scratch

Section

Project

51. Project: MultinomialNB from scratch

Concept

Implement a ScratchMNB class that matches sklearn's MultinomialNB on the digits dataset (treating pixel intensities as word counts).

#milestonegoal
1Count table + LaplaceBuild log_prob_ matrix of shape (n_classes, n_features)
2Predict via log-spacelog prior + X @ log_prob_.T → argmax
3Verify vs sklearnaccuracy matches to 3 decimal places; predictions identical

Build rules: one forward pass only (no loops over samples at predict time), alpha=1.0, equal class priors estimated from training labels.

52. Break it if you can: Project: MultinomialNB from scratch

Counterexample

Discussion prompt

Implement a ScratchMNB class that matches sklearn's MultinomialNB on the digits dataset (treating pixel intensities as word counts).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: one forward pass only (no loops over samples at predict time), alpha=1.0, equal class priors estimated from training labels.

53. Milestone 1 — build the count table and log probs

Worked example

Your turn: load digits, split 70/30, build the per-class count array and apply Laplace smoothing. Predict what log_prob_[0, :].sum() will be (approximately).

Hint: counts[c] = X[y==c].sum(axis=0) + alpha; then log_prob_[c] = log(counts[c] / counts[c].sum()).

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split

digits = load_digits()
Xd, yd = digits.data.astype(int), digits.target
Xtr, Xte, ytr, yte = train_test_split(Xd, yd, test_size=0.3, random_state=42)

alpha = 1.0
classes = np.unique(ytr)
n_feats = Xtr.shape[1]
log_prob = np.zeros((len(classes), n_feats))
for i, c in enumerate(classes):
    cnts = Xtr[ytr==c].sum(axis=0) + alpha
    log_prob[i] = np.log(cnts / cnts.sum())
print('log_prob shape:', log_prob.shape)
print('log_prob[0] sum:', round(log_prob[0].sum(), 4))
outputvalue
log_prob shape(10, 64)
log_prob[0].sum()a negative float (64 log-prob values summed; varies by class)
rowsone per digit class (0-9)
colsone per pixel feature

54. Milestone 2 — predict in log-space

Worked example

Your turn: compute log priors from ytr, then predict via scores = X @ log_prob.T + log_prior. Predict the accuracy before running.

Hint: log_prior[c] = log(count(c in ytr) / len(ytr)); predicted = classes[argmax(scores, axis=1)].

from sklearn.metrics import accuracy_score

log_prior = np.log(np.array([(ytr==c).sum()/len(ytr) for c in classes]))
scores = Xte @ log_prob.T + log_prior
preds_scratch = classes[np.argmax(scores, axis=1)]
print('scratch accuracy:', accuracy_score(yte, preds_scratch))
metricvalue
scratch NB accuracy on digits0.8944
n_test samples540
log_prior (balanced)log(1/10) each = -2.3026

55. Milestone 3 — verify against sklearn

Worked example

Your turn: fit MultinomialNB(alpha=1.0) on the same split and compare accuracy and per-sample predictions. Predict whether they match exactly.

Hint: (preds_scratch == mnb.predict(Xte)).mean() should be 1.0.

from sklearn.naive_bayes import MultinomialNB

mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
acc_sk = accuracy_score(yte, mnb.predict(Xte))
print('sklearn accuracy:', acc_sk)
print('match fraction:',
      (preds_scratch == mnb.predict(Xte)).mean())
modelaccuracypredictions match scratch
ScratchMNB0.8944—
sklearn MNB0.89441.0000 (exact match)
verdictimplementation correct

56. What each one costs: Milestone 3 — verify against sklearn

Trade off

Comparison matrix

From Milestone 3 — verify against sklearn: every row here is a choice with a cost. Fill the accuracy column, then say which row you would actually pick and what you give up for it.

modelaccuracypredictions match scratch
ScratchMNB0.8944—
sklearn MNB0.89441.0000 (exact match)
verdictimplementation correct

57. The full program

Concept

import numpy as np
from sklearn.datasets import load_digits
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import accuracy_score

digits = load_digits()
Xd, yd = digits.data.astype(int), digits.target
Xtr, Xte, ytr, yte = train_test_split(Xd, yd, test_size=0.3, random_state=42)

alpha = 1.0
classes = np.unique(ytr)
log_prob = np.zeros((len(classes), Xtr.shape[1]))
for i, c in enumerate(classes):
    cnts = Xtr[ytr==c].sum(axis=0) + alpha
    log_prob[i] = np.log(cnts / cnts.sum())
log_prior = np.log(np.array([(ytr==c).sum()/len(ytr) for c in classes]))
preds = classes[np.argmax(Xte @ log_prob.T + log_prior, axis=1)]

mnb = MultinomialNB(alpha=1.0).fit(Xtr, ytr)
print('scratch:', accuracy_score(yte, preds))
print('sklearn:', accuracy_score(yte, mnb.predict(Xte)))
print('match:',   (preds == mnb.predict(Xte)).mean())
outputvalue
scratch accuracy0.8944
sklearn accuracy0.8944
match fraction1.0000

If your match fraction is 1.0 — you've implemented Naive Bayes from Bayes' theorem to working code. (Lesson 40 trained a linear model with gradient descent; NB gets there in a single pass.)

58. Fill in: value for The full program

Comparison

Comparison matrix

From The full program: refill the value column from what you know. The rest of the table is as it appeared.

outputvalue
scratch accuracy0.8944
sklearn accuracy0.8944
match fraction1.0000

59. Show it off

Concept

Out loud, slides closed: (1) derive the Laplace smoothing formula from scratch and explain why alpha goes in both numerator and denominator, (2) explain why NB is immune to float underflow when implemented in log-space, (3) prove that binary MNB is a linear classifier by writing out the weight vector.

Stretch (homework): implement the spam classifier from the Lesson 65 worked example using the USAAIO docs corpus; compute F1 and inspect false positives. Next: logistic regression also becomes linear in log-space — but it learns weights via gradient descent (Lesson 40) rather than counting.

60. Connect it up: Lesson 65: Naive Bayes Classifiers

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The Bayes update for text · Log-space arithmetic · Bernoulli NB vs Multinomial NB · NB is a linear classifier · Your turn: build NB from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

ideathe one thing to remember
Laplace smoothingadd alpha to numerator, alpha*V to denominator
Log-spacesum log probs — never multiply raw probs
MNB vs BNBMNB counts; BNB sees only presence
NB speedO(nd) train, O(d) predict — one pass, no iterations
NB is linearw = log p_1 - log p_0; same decision boundary as LR

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 65 (Multinomial Naive Bayes) — Barron · USAAIO Round 2 Preparation, 2026
  2. Laplace smoothing, log-space NB, BNB vs MNB, ComplementNB, linear equivalence — all numbers verified with sklearn/numpy, scratch implementation matches sklearn — numpy 2.2.6 + scikit-learn, real execution, June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108