Lesson 64: Bag of Words, TF-IDF, and Text Vectorization

USAAIO Lesson 64. It covers bag-of-words count vectors and TF-IDF weighting, giving both the formula and a from-scratch implementation, then n-grams and the text-preprocessing pipeline - lowercasing, stopword removal, and stemming against lemmatization. It covers CountVectorizer and TfidfVectorizer, and ends with a full text-classification pipeline combining TF-IDF with logistic regression. The lesson runs to 30 slides.

Subject: Machine Learning · 59 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Bag of Words, TF-IDF, and Text Vectorization

Title

USAAIO · Lesson 64 · NLP Foundations

Turn raw text into numbers a classifier can learn from. Today: BoW, TF-IDF, n-grams, preprocessing, sklearn vectorizers, and a full classification pipeline.

2. By the end of this lesson you can

Objectives

  1. Build a Bag-of-Words count matrix from scratch and explain its sparsity
  2. Compute TF-IDF by hand (TF × IDF formula) and explain why it down-weights stop-words
  3. Explain the vocabulary trade-off of unigrams vs bigrams (n-grams)
  4. Apply the text preprocessing pipeline (lowercase → stopwords → stemming/lemma)
  5. Wire CountVectorizer / TfidfVectorizer into a logistic regression classifier

3. What survived from Transfer Learning, Fine-Tuning & Few-Shot Methods?

Warm-up

Discussion prompt

Before we open Lesson 64: Bag of Words, TF-IDF, and Text Vectorization: without looking back, what was the main idea of Transfer Learning, Fine-Tuning & Few-Shot Methods, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

fine-tuning strategies (linear probe → gradual unfreeze → full fine-tune), domain adaptation and distribution shift, few-shot learning (5-shot, prototypical networks), and zero-shot learning via textual descriptions (CLIP foundation).

4. Bag of Words — counting tokens

Section

Part 1 of 4

5. Representing text as numbers

Concept

Machine learning models need fixed-length numeric vectors. Text has variable length and contains symbols, not numbers — we need a bridge.

The simplest bridge: for each document, count how many times each vocabulary word appears. Ignore order, ignore grammar — just counts.

doccatchasesdogeatsfishrat
cat chases rat110001
cat eats fish100110
dog chases cat111000

6. What stays fixed: Representing text as numbers

Invariant

Step through it

Step through Representing text as numbers one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: doc is cat chases rat
  2. Step 2: doc is cat eats fish
  3. Step 3: doc is dog chases cat

7. Why BoW works at all

Intuition

If a document contains the word 'excellent' five times it is probably positive. If it contains 'terrible' three times it is probably negative. Word identity matters more than position for many classification tasks.

The count vector is almost entirely zeros — a vocabulary of 10 000 words with a 50-word document has 9 950 zeros. This sparsity is the main runtime concern (Lesson 32's dimensionality theme).

8. Break it if you can: Why BoW works at all

Counterexample

Discussion prompt

The count vector is almost entirely zeros — a vocabulary of 10 000 words with a 50-word document has 9 950 zeros. This sparsity is the main runtime concern (Lesson 32's dimensionality theme).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. Guess the shape of the answer: BoW with CountVectorizer

Estimation

Predict first

Fit a CountVectorizer on three sentences. Check the vocabulary index and the count matrix.

Commit before you compute: what does BoW with CountVectorizer come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: features: ['cat','chases','dog','eats','fish','rat']; vocab size 6

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every unique token becomes a column; unseen words in new documents get silently ignored (OOV problem).

10. BoW with CountVectorizer

Worked example

Fit a CountVectorizer on three sentences. Check the vocabulary index and the count matrix.

from sklearn.feature_extraction.text import CountVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
cv = CountVectorizer()
X = cv.fit_transform(corpus).toarray()
print(cv.get_feature_names_out())
print(X)

features: ['cat','chases','dog','eats','fish','rat']; vocab size 6

Why: Every unique token becomes a column; unseen words in new documents get silently ignored (OOV problem).

doccatchasesdogeatsfishrat
cat chases rat110001
cat eats fish100110
dog chases cat111000

11. What stays fixed: BoW with CountVectorizer

Invariant

Step through it

Step through BoW with CountVectorizer one row at a time. One of these columns never changes — find it, and say why it cannot.

  1. Step 1: doc is cat chases rat
  2. Step 2: doc is cat eats fish
  3. Step 3: doc is dog chases cat

12. TF-IDF — rewarding rare words

Section

Part 2 of 4

13. The TF-IDF formula

Concept

Raw counts treat every word equally. 'The' appearing 20 times conveys nothing. TF-IDF multiplies term frequency by inverse document frequency to down-weight ubiquitous words.

\[ \text{TF}(t,d) = \frac{\text{count}(t,d)}{|d|} \qquad \text{IDF}(t) = \log\!\left(\frac{N}{\text{df}(t)}\right) \]

\[ \text{TF-IDF}(t,d) = \text{TF}(t,d) \times \text{IDF}(t) \]

df(t) is the number of documents containing term t; N is the total document count. A word in every document has IDF = log(1) = 0 — it is zeroed out.

14. By analogy: The TF-IDF formula

Analogy

Discussion prompt

Explain The TF-IDF formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Raw counts treat every word equally. 'The' appearing 20 times conveys nothing. TF-IDF multiplies term frequency by inverse document frequency to down-weight ubiquitous words.

15. What has to be given first: TF-IDF from scratch

Missing information

Discussion prompt

Implement TF-IDF with Python dicts. Corpus: ["cat chases rat", "cat eats fish", "dog chases cat"]. Compute scores for doc0.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

'cat' appears in every document so its IDF is log(3/3)=0 — it is a stop-word here. 'rat' appears only once, giving it the highest score.

16. TF-IDF from scratch

Worked example

Implement TF-IDF with Python dicts. Corpus: ["cat chases rat", "cat eats fish", "dog chases cat"]. Compute scores for doc0.

import math
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
tok = [d.split() for d in corpus]
N = len(tok)
df = {w: sum(1 for d in tok if w in d) for w in set(w for d in tok for w in d)}

doc, L = tok[0], len(tok[0])  # 'cat chases rat'
for w in sorted(set(doc)):
    tf = doc.count(w) / L
    idf = math.log(N / df[w])
    print(f'{w:8s}: TF={tf:.4f}  IDF={idf:.4f}  TF-IDF={tf*idf:.4f}')

cat: TF-IDF=0.0000 (in all 3 docs, IDF=0); chases: 0.1352; rat: 0.3662

Why: 'cat' appears in every document so its IDF is log(3/3)=0 — it is a stop-word here. 'rat' appears only once, giving it the highest score.

wordTFIDFTF-IDF
cat0.33330.0000 (log 3/3)0.0000
chases0.33330.4055 (log 3/2)0.1352
rat0.33331.0986 (log 3/1)0.3662

17. Fill in: TF-IDF for TF-IDF from scratch

Comparison

Comparison matrix

From TF-IDF from scratch: refill the TF-IDF column from what you know. The rest of the table is as it appeared.

wordTFIDFTF-IDF
cat0.33330.0000 (log 3/3)0.0000
chases0.33330.4055 (log 3/2)0.1352
rat0.33331.0986 (log 3/1)0.3662

18. Something is wrong here: raw count vs TF-IDF for common words

Anomaly

Predict first

A student writes this, and it looks reasonable:

Use raw BoW counts for sentiment. 'The' appears 5 times → treat it as highly informative.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Stop-words dominate count vectors.

Apply TF-IDF weighting. 'the' in every document gets IDF=0; unique sentiment words score high.

Why: Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.

19. Trap: raw count vs TF-IDF for common words

Trap

The trap

Use raw BoW counts for sentiment. 'The' appears 5 times → treat it as highly informative.

Let the model learn that high count of 'the' predicts sentiment

Why: Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.

The fix

Apply TF-IDF weighting. 'the' in every document gets IDF=0; unique sentiment words score high.

IDF zeroes out universal stop-words automatically; unique discriminating words receive large scores

Why: TF-IDF is a principled reweighting — no stop-word list required, though a custom list can remove domain stop-words the IDF doesn't catch.

20. Break it on purpose: raw count vs TF-IDF for common words

Break the constraint

Discussion prompt

The rule this trap just fixed:

TF-IDF is a principled reweighting — no stop-word list required, though a custom list can remove domain stop-words the IDF doesn't catch.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.

21. N-grams and text preprocessing

Section

Part 3 of 4

22. N-grams: capturing local context

Concept

A unigram treats each word independently. A bigram treats each consecutive pair as a feature, preserving local context that unigrams miss.

n-gram typeexample tokens from 'not good'captures
unigram (n=1)'not', 'good'individual words
bigram (n=2)'not_good'negation context
trigram (n=3)'was_not_good'wider dependency

The cost: vocabulary grows exponentially. On the 6-word corpus, unigrams give 6 features, adding bigrams gives 12, adding trigrams gives 15 — with real corpora the explosion is dramatic.

23. What each one costs: N-grams: capturing local context

Trade off

Comparison matrix

From N-grams: capturing local context: every row here is a choice with a cost. Fill the example tokens from 'not good' column, then say which row you would actually pick and what you give up for it.

n-gram typeexample tokens from 'not good'captures
unigram (n=1)'not', 'good'individual words
bigram (n=2)'not_good'negation context
trigram (n=3)'was_not_good'wider dependency

24. Text preprocessing pipeline

Concept

  1. Lowercase — 'Cat' and 'cat' are one token
  2. Remove punctuation / special characters — strip noise
  3. Stopwords — drop 'the', 'a', 'is' (but beware: 'not' is a stop-word!)
  4. Stemming (fast, aggressive) — 'running' → 'run', 'ran' → 'ran' (rule-based)
  5. Lemmatization (slow, accurate) — 'running' → 'run', 'ran' → 'run' (dictionary-based)

Order matters: lowercase before stemming, stopwords before vectorizing. Stemming and lemmatization are interchangeable — pick lemmatization when morphological correctness matters (e.g., medical text).

25. Guess the shape of the answer: CountVectorizer with n-grams

Estimation

Predict first

Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.

Commit before you compute: what does CountVectorizer with n-grams come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: max n=1: 6; max n=2: 12; max n=3: 15

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each higher-order n-gram adds all new consecutive sequences.

26. CountVectorizer with n-grams

Worked example

Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.

from sklearn.feature_extraction.text import CountVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
for ng in [1, 2, 3]:
    cv = CountVectorizer(ngram_range=(1, ng))
    cv.fit(corpus)
    print(f'max n={ng}: vocab size = {len(cv.vocabulary_)}')

max n=1: 6; max n=2: 12; max n=3: 15

Why: Each higher-order n-gram adds all new consecutive sequences. On real corpora (vocab ~50k) bigrams easily produce millions of features — sparsity and memory become dominant concerns.

ngram_rangevocab size (6-word corpus)
(1,1) unigrams only6
(1,2) + bigrams12
(1,3) + trigrams15

27. Work backwards from the answer: CountVectorizer with n-grams

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

max n=1: 6; max n=2: 12; max n=3: 15

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.

28. Something is wrong here: stop-word removal deletes negation

Anomaly

Predict first

A student writes this, and it looks reasonable:

Always remove the default NLTK stop-word list before vectorizing — it cleans noise.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.

Audit the stop-word list for domain-critical words (especially negation: 'not', 'no', 'never').

Why: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.

29. Trap: stop-word removal deletes negation

Trap

The trap

Always remove the default NLTK stop-word list before vectorizing — it cleans noise.

Apply stop-words blindly; 'not' is in the default stop-word list and is removed

Why: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.

The fix

Audit the stop-word list for domain-critical words (especially negation: 'not', 'no', 'never').

Remove 'not' from the stop-word set, or use a bigram model to capture 'not_good' directly

Why: Blanket stop-word removal improves speed and sparsity, but can destroy discriminative signal. Always check which words are being dropped.

30. Which of these survive contact with Lesson 64: Bag of Words, TF-IDF, and Text…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Machine learning models need fixed-length numeric vectors. Text has variable length and contains symbols, not numbers — we need a bridge.; df(t) is the number of documents containing term t; N is the total document count. A word in every document has IDF = log(1) = 0 — it is zeroed out.; A unigram treats each word independently. A bigram treats each consecutive pair as a feature, preserving local context that unigrams miss.
Breaks
Use raw BoW counts for sentiment. 'The' appears 5 times → treat it as highly informative.; Always remove the default NLTK stop-word list before vectorizing — it cleans noise.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 64: Bag of Words, TF-IDF, and Text Vectorization puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. sklearn vectorizers and the full pipeline

Section

Part 4 of 4

32. TfidfVectorizer: fit vs transform

Concept

fit builds the vocabulary and computes IDF from the training corpus. transform applies those fixed IDF values to new documents — never refit on test data (data leakage, Lesson 32).

methoduses training vocab?updates IDF?call on
fit_transformbuilds ityestraining set only
transformyes (frozen)noval/test set

vocabulary_ stores the word → column mapping. Inspect it to debug unexpected feature columns or OOV drops.

33. Teach it back: TfidfVectorizer: fit vs transform

Explain it

Discussion prompt

Explain TfidfVectorizer: fit vs transform to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.

Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.

Answer:

fit builds the vocabulary and computes IDF from the training corpus. transform applies those fixed IDF values to new documents — never refit on test data (data leakage, Lesson 32).

34. Guess the shape of the answer: TF-IDF + Logistic Regression pipeline

Estimation

Predict first

Wire TfidfVectorizer into a Pipeline with LogisticRegression. Fit, predict, check accuracy.

Commit before you compute: what does TF-IDF + Logistic Regression pipeline come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: acc: 1.0; vocab_size: 18

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. TF-IDF provides a clean sparse representation.

35. TF-IDF + Logistic Regression pipeline

Worked example

Wire TfidfVectorizer into a Pipeline with LogisticRegression. Fit, predict, check accuracy.

from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
import warnings; warnings.filterwarnings('ignore')

train = ["I love this film excellent", "great movie wonderful story",
         "amazing performance brilliant",
         "terrible film awful", "bad movie boring", "horrible garbage waste"]
y_tr = [1, 1, 1, 0, 0, 0]
test  = ["excellent film loved it", "terrible awful waste", "great performance"]
y_te  = [1, 0, 1]

pipe = Pipeline([('tfidf', TfidfVectorizer()),
                 ('clf',   LogisticRegression(max_iter=1000, random_state=0))])
pipe.fit(train, y_tr)
print('acc:', accuracy_score(y_te, pipe.predict(test)))
print('vocab size:', len(pipe.named_steps['tfidf'].vocabulary_))

acc: 1.0; vocab_size: 18

Why: TF-IDF provides a clean sparse representation. Logistic regression (Lesson 40's pattern) over the TF-IDF features classifies sentiment correctly on this toy corpus.

test docpredictiontrue
excellent film loved it1 (positive)1
terrible awful waste0 (negative)0
great performance1 (positive)1

36. Fill in: prediction for TF-IDF + Logistic Regression pipeline

Comparison

Comparison matrix

From TF-IDF + Logistic Regression pipeline: refill the prediction column from what you know. The rest of the table is as it appeared.

test docpredictiontrue
excellent film loved it1 (positive)1
terrible awful waste0 (negative)0
great performance1 (positive)1

37. What TF-IDF cannot do

Concept

TF-IDF is a bag: it destroys word order. 'Dog bites man' and 'Man bites dog' produce identical vectors if the vocabulary counts are the same.

These failures motivate word embeddings (Lesson 65+), where 'happy' and 'joyful' map to nearby vectors and word order is preserved via position encodings.

38. By analogy: What TF-IDF cannot do

Analogy

Discussion prompt

Explain What TF-IDF cannot do by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

TF-IDF is a bag: it destroys word order. 'Dog bites man' and 'Man bites dog' produce identical vectors if the vocabulary counts are the same.

39. Rebuild the recipe: The text vectorization recipe

Ranking

Put in order

These are the steps of The text vectorization recipe, scrambled. Put them back in order before the next slide shows you.

  1. Preprocess: lowercase → remove punctuation → stopwords (audit for negation!) → stem/lemmatize
  2. Represent: CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoW
  3. N-grams: ngram_range=(1,2) for bigrams — watch vocab explosion
  4. Fit on train only: fit_transform(train), then transform(test) — no data leakage
  5. Downstream model: Pipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistent

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

40. The text vectorization recipe

Pattern

  1. Preprocess: lowercase → remove punctuation → stopwords (audit for negation!) → stem/lemmatize
  2. Represent: CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoW
  3. N-grams: ngram_range=(1,2) for bigrams — watch vocab explosion
  4. Fit on train only: fit_transform(train), then transform(test) — no data leakage
  5. Downstream model: Pipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistent

41. Where does it stop working: The text vectorization recipe

Edge cases

Discussion prompt

The text vectorization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Preprocess: lowercase → remove punctuation → stopwords (audit for negation!) → stem/lemmatize
  2. Represent: CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoW
  3. N-grams: ngram_range=(1,2) for bigrams — watch vocab explosion
  4. Fit on train only: fit_transform(train), then transform(test) — no data leakage
  5. Downstream model: Pipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistent

42. Rule out three: Check yourself — IDF of a universal word

Elimination

Eliminate the wrong options

A corpus has N=1000 documents. The word 'the' appears in all 1000. What is IDF('the') using the formula IDF = log(N/df)?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. log(1) = 0
  • B. log(1000) ≈ 6.91
  • C. 1/1000 = 0.001
  • D. 1 (by definition for universal words)

Survives elimination: A

Why: IDF = log(N/df) = log(1000/1000) = log(1) = 0. A word in every document contributes zero weight — exactly the stop-word suppression TF-IDF is designed for.

43. Check yourself — IDF of a universal word

Check

Work out the IDF before clicking.

Check your understanding

A corpus has N=1000 documents. The word 'the' appears in all 1000. What is IDF('the') using the formula IDF = log(N/df)?

  • A. log(1) = 0 (correct)
  • B. log(1000) ≈ 6.91
  • C. 1/1000 = 0.001
  • D. 1 (by definition for universal words)

Answer: A

Why: IDF = log(N/df) = log(1000/1000) = log(1) = 0. A word in every document contributes zero weight — exactly the stop-word suppression TF-IDF is designed for.

Why B tempts people
log(N) is the IDF of a word that appears in exactly ONE document (df=1), not all documents.
Why C tempts people
1/N is TF for a word appearing once in an N-token document; IDF is computed across documents, not within them.
Why D tempts people
There is no special case — the formula log(N/df) handles this; when df=N the result is log(1)=0, not 1.

44. Answer it before you see the options: Check yourself — fit vs transform

Prediction

Predict first

You fit a TfidfVectorizer on the full dataset (train + test), then split and train the classifier. What is wrong?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The IDF values were computed using test documents, leaking test distribution into training

Why: Fitting on train+test means the IDF for each token is computed using test documents. At deployment you won't have test data, so IDF values will differ — this is data leakage. Always fit_transform on train only, then transform test.

45. Check yourself — fit vs transform

Check

Data leakage or correct?

Check your understanding

You fit a TfidfVectorizer on the full dataset (train + test), then split and train the classifier. What is wrong?

  • A. The IDF values were computed using test documents, leaking test distribution into training (correct)
  • B. Nothing — TF-IDF has no learnable parameters so it cannot leak
  • C. The vocabulary will be too small because some test words are unseen
  • D. The pipeline breaks if test documents have words not in the training vocab

Answer: A

Why: Fitting on train+test means the IDF for each token is computed using test documents. At deployment you won't have test data, so IDF values will differ — this is data leakage. Always fit_transform on train only, then transform test.

Why B tempts people
TF-IDF does have learned parameters: the vocabulary mapping and IDF weights, both derived from the corpus it was fit on.
Why C tempts people
If you fit on train+test, the vocabulary is larger (includes test words), not smaller — this is the opposite of the real concern.
Why D tempts people
Unseen words are silently ignored by sklearn's transform — this is the OOV behavior, not the leakage bug described.

46. Rule out three: Check yourself — n-gram trade-off

Elimination

Eliminate the wrong options

Switching from unigrams to unigrams+bigrams (ngram_range=(1,2)) on a large corpus primarily causes:

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Vocabulary explosion and increased sparsity; captures local word order
  • B. Vocabulary shrinks because bigrams replace unigrams
  • C. IDF scores become undefined for bigrams
  • D. Each document vector becomes denser because more features fire

Survives elimination: A

Why: Adding bigrams multiplies the feature space. On the 6-word toy corpus unigrams→bigrams grows from 6 to 12 features; on a 50k-word corpus the explosion is enormous. Vectors remain sparse (most bigrams don't appear in a given doc) but memory and training time increase substantially.

47. Check yourself — n-gram trade-off

Check

What does adding bigrams buy you, and what does it cost?

Check your understanding

Switching from unigrams to unigrams+bigrams (ngram_range=(1,2)) on a large corpus primarily causes:

  • A. Vocabulary explosion and increased sparsity; captures local word order (correct)
  • B. Vocabulary shrinks because bigrams replace unigrams
  • C. IDF scores become undefined for bigrams
  • D. Each document vector becomes denser because more features fire

Answer: A

Why: Adding bigrams multiplies the feature space. On the 6-word toy corpus unigrams→bigrams grows from 6 to 12 features; on a 50k-word corpus the explosion is enormous. Vectors remain sparse (most bigrams don't appear in a given doc) but memory and training time increase substantially.

Why B tempts people
Bigrams are added on top of unigrams when ngram_range=(1,2) — the vocabulary grows, not shrinks.
Why C tempts people
IDF is computed the same way for bigrams: log(N/df(bigram)). Bigrams simply have their own df counts.
Why D tempts people
Vectors remain sparse: most bigrams are absent from any given document. The density does not increase substantially.

48. Your turn: build it

Section

Project

49. Project: TF-IDF from scratch + classification pipeline

Concept

Build TF-IDF from first principles with Python dicts, then wire it into an sklearn pipeline and classify sentiment.

#requirementtool
1TF-IDF from scratch (dicts)math.log + Counter
2sklearn vectorizers — BoW and TF-IDFCountVectorizer, TfidfVectorizer
3full pipeline: TF-IDF + LRPipeline, LogisticRegression

Build rules: never call fit_transform on test data; inspect vocabulary_; verify that a word in every document scores TF-IDF = 0.

50. Break it if you can: Project: TF-IDF from scratch + classification…

Counterexample

Discussion prompt

Build TF-IDF from first principles with Python dicts, then wire it into an sklearn pipeline and classify sentiment.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: never call fit_transform on test data; inspect vocabulary_; verify that a word in every document scores TF-IDF = 0.

51. Milestone 1 — TF-IDF from scratch

Worked example

Your turn: implement TF-IDF with only math and built-ins. Predict the TF-IDF of 'rat' in doc0.

Hint: TF = count(word, doc) / len(doc); IDF = log(N / df[word]); df = number of docs containing word.

import math
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
tok = [d.split() for d in corpus]
N = len(tok)
df = {w: sum(1 for d in tok if w in d) for w in set(w for d in tok for w in d)}
doc, L = tok[0], len(tok[0])
for w in sorted(set(doc)):
    tf  = doc.count(w) / L
    idf = math.log(N / df[w])
    print(f'{w:8s}: TF={tf:.4f}  IDF={idf:.4f}  TF-IDF={tf*idf:.4f}')
wordTFIDFTF-IDF
cat0.33330.00000.0000
chases0.33330.40550.1352
rat0.33331.09860.3662

52. What each one costs: Milestone 1 — TF-IDF from scratch

Trade off

Comparison matrix

From Milestone 1 — TF-IDF from scratch: every row here is a choice with a cost. Fill the TF column, then say which row you would actually pick and what you give up for it.

wordTFIDFTF-IDF
cat0.33330.00000.0000
chases0.33330.40550.1352
rat0.33331.09860.3662

53. Milestone 2 — sklearn vectorizers

Worked example

Your turn: fit CountVectorizer and TfidfVectorizer on the same corpus. Inspect vocabulary and compare doc0 representations.

Hint: .fit_transform(corpus).toarray() then .get_feature_names_out(). Note sklearn adds +1 smoothing to IDF by default (smooth_idf=True) — values will differ slightly from your scratch version.

from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
cv = CountVectorizer()
X_bow = cv.fit_transform(corpus).toarray()
print('BoW vocab:', cv.get_feature_names_out())
print('BoW doc0: ', X_bow[0])

tv = TfidfVectorizer(norm=None, smooth_idf=False)
X_tf = tv.fit_transform(corpus).toarray()
print('TF-IDF doc0 (no smooth):', X_tf[0].round(4))
featureBoW doc0TF-IDF doc0 (no smooth)
cat11.0000
chases11.4055
dog00.0000
eats00.0000
fish00.0000
rat12.0986

54. Which is which, by BoW doc0

Discrimination

Sort into buckets

Sort these by BoW doc0, from memory, without looking back at Milestone 2 — sklearn vectorizers. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

1
cat; chases; rat
0
dog; eats; fish
g1
BoW doc0 is "1" for cat, chases, rat — that is what the table on "Milestone 2 — sklearn vectorizers" records, and it is the single property separating this group from the rest.
g2
BoW doc0 is "0" for dog, eats, fish — that is what the table on "Milestone 2 — sklearn vectorizers" records, and it is the single property separating this group from the rest.

55. Milestone 3 — full classification pipeline

Worked example

Your turn: wire TfidfVectorizer + LogisticRegression into a Pipeline. Predict whether the pipeline achieves perfect accuracy on the toy set.

Hint: Pipeline([('tfidf', TfidfVectorizer()), ('clf', LogisticRegression(...))]); fit on train, predict on test, compute accuracy_score.

from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
import warnings; warnings.filterwarnings('ignore')

train=["I love this film excellent","great movie wonderful story",
       "amazing performance brilliant",
       "terrible film awful","bad movie boring","horrible garbage waste"]
y_tr=[1,1,1,0,0,0]
test=["excellent film loved it","terrible awful waste","great performance"]
y_te=[1,0,1]

pipe=Pipeline([('tfidf',TfidfVectorizer()),
               ('clf', LogisticRegression(max_iter=1000, random_state=0))])
pipe.fit(train,y_tr)
print('acc:', accuracy_score(y_te, pipe.predict(test)))
print('vocab size:', len(pipe.named_steps['tfidf'].vocabulary_))
test docpredtrue
excellent film loved it11
terrible awful waste00
great performance11

56. Fill in: pred for Milestone 3 — full classification pipeline

Comparison

Comparison matrix

From Milestone 3 — full classification pipeline: refill the pred column from what you know. The rest of the table is as it appeared.

test docpredtrue
excellent film loved it11
terrible awful waste00
great performance11

57. Show it off

Concept

Out loud, slides closed: (1) derive the TF-IDF formula from scratch and explain why IDF zeroes out stop-words; (2) explain what happens to 'rabbit' in a test document if 'rabbit' never appeared in training; (3) name two things TF-IDF cannot represent.

Stretch (homework): compare BoW vs TF-IDF vs bigrams on a sentiment dataset; implement TF-IDF that handles repeated terms across docs (smooth_idf=True); analyze which features logistic regression weighted highest with coef_. Next up: word embeddings — dense vectors that capture semantics.

58. Connect it up: Lesson 64: Bag of Words, TF-IDF, and Text Vectorization

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Bag of Words — counting tokens · TF-IDF — rewarding rare words · N-grams and text preprocessing · sklearn vectorizers and the full pipeline · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. What you can do now

Recap

ideathe one thing to remember
BoWcounts per word; order discarded; sparse
TF-IDFTF × IDF; IDF = log(N/df); stop-words → 0
n-gramsbigrams capture context; vocab explodes exponentially
fit vs transformfit on train only; transform test (no leakage)
limitationsno order, no synonyms, no semantics → embeddings next

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 64 (Text Representations: BoW, TF-IDF, N-grams) — Barron · USAAIO Round 2 Preparation, 2026
  2. TF-IDF from scratch and sklearn TfidfVectorizer / CountVectorizer pipeline verified — numpy 2.2.6, scikit-learn, Python 3.12, real execution June 2026

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108