USAAIO Lesson 64. It covers bag-of-words count vectors and TF-IDF weighting, giving both the formula and a from-scratch implementation, then n-grams and the text-preprocessing pipeline - lowercasing, stopword removal, and stemming against lemmatization. It covers CountVectorizer and TfidfVectorizer, and ends with a full text-classification pipeline combining TF-IDF with logistic regression. The lesson runs to 30 slides.
Subject: Machine Learning · 59 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 64 · NLP Foundations
Turn raw text into numbers a classifier can learn from. Today: BoW, TF-IDF, n-grams, preprocessing, sklearn vectorizers, and a full classification pipeline.
Objectives
CountVectorizer / TfidfVectorizer into a logistic regression classifierWarm-up
Discussion prompt
Before we open Lesson 64: Bag of Words, TF-IDF, and Text Vectorization: without looking back, what was the main idea of Transfer Learning, Fine-Tuning & Few-Shot Methods, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
fine-tuning strategies (linear probe → gradual unfreeze → full fine-tune), domain adaptation and distribution shift, few-shot learning (5-shot, prototypical networks), and zero-shot learning via textual descriptions (CLIP foundation).
Section
Part 1 of 4
Concept
Machine learning models need fixed-length numeric vectors. Text has variable length and contains symbols, not numbers — we need a bridge.
The simplest bridge: for each document, count how many times each vocabulary word appears. Ignore order, ignore grammar — just counts.
| doc | cat | chases | dog | eats | fish | rat |
|---|---|---|---|---|---|---|
| cat chases rat | 1 | 1 | 0 | 0 | 0 | 1 |
| cat eats fish | 1 | 0 | 0 | 1 | 1 | 0 |
| dog chases cat | 1 | 1 | 1 | 0 | 0 | 0 |
Invariant
Step through it
Step through Representing text as numbers one row at a time. One of these columns never changes — find it, and say why it cannot.
Intuition
If a document contains the word 'excellent' five times it is probably positive. If it contains 'terrible' three times it is probably negative. Word identity matters more than position for many classification tasks.
The count vector is almost entirely zeros — a vocabulary of 10 000 words with a 50-word document has 9 950 zeros. This sparsity is the main runtime concern (Lesson 32's dimensionality theme).
Counterexample
Discussion prompt
The count vector is almost entirely zeros — a vocabulary of 10 000 words with a 50-word document has 9 950 zeros. This sparsity is the main runtime concern (Lesson 32's dimensionality theme).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Estimation
Predict first
Fit a CountVectorizer on three sentences. Check the vocabulary index and the count matrix.
Commit before you compute: what does BoW with CountVectorizer come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: features: ['cat','chases','dog','eats','fish','rat']; vocab size 6
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Every unique token becomes a column; unseen words in new documents get silently ignored (OOV problem).
Worked example
Fit a CountVectorizer on three sentences. Check the vocabulary index and the count matrix.
from sklearn.feature_extraction.text import CountVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
cv = CountVectorizer()
X = cv.fit_transform(corpus).toarray()
print(cv.get_feature_names_out())
print(X)features: ['cat','chases','dog','eats','fish','rat']; vocab size 6
Why: Every unique token becomes a column; unseen words in new documents get silently ignored (OOV problem).
| doc | cat | chases | dog | eats | fish | rat |
|---|---|---|---|---|---|---|
| cat chases rat | 1 | 1 | 0 | 0 | 0 | 1 |
| cat eats fish | 1 | 0 | 0 | 1 | 1 | 0 |
| dog chases cat | 1 | 1 | 1 | 0 | 0 | 0 |
Invariant
Step through it
Step through BoW with CountVectorizer one row at a time. One of these columns never changes — find it, and say why it cannot.
Section
Part 2 of 4
Concept
Raw counts treat every word equally. 'The' appearing 20 times conveys nothing. TF-IDF multiplies term frequency by inverse document frequency to down-weight ubiquitous words.
\[ \text{TF}(t,d) = \frac{\text{count}(t,d)}{|d|} \qquad \text{IDF}(t) = \log\!\left(\frac{N}{\text{df}(t)}\right) \]
\[ \text{TF-IDF}(t,d) = \text{TF}(t,d) \times \text{IDF}(t) \]
df(t) is the number of documents containing term t; N is the total document count. A word in every document has IDF = log(1) = 0 — it is zeroed out.
Analogy
Discussion prompt
Explain The TF-IDF formula by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Raw counts treat every word equally. 'The' appearing 20 times conveys nothing. TF-IDF multiplies term frequency by inverse document frequency to down-weight ubiquitous words.
Missing information
Discussion prompt
Implement TF-IDF with Python dicts. Corpus: ["cat chases rat", "cat eats fish", "dog chases cat"]. Compute scores for doc0.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
'cat' appears in every document so its IDF is log(3/3)=0 — it is a stop-word here. 'rat' appears only once, giving it the highest score.
Worked example
Implement TF-IDF with Python dicts. Corpus: ["cat chases rat", "cat eats fish", "dog chases cat"]. Compute scores for doc0.
import math
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
tok = [d.split() for d in corpus]
N = len(tok)
df = {w: sum(1 for d in tok if w in d) for w in set(w for d in tok for w in d)}
doc, L = tok[0], len(tok[0]) # 'cat chases rat'
for w in sorted(set(doc)):
tf = doc.count(w) / L
idf = math.log(N / df[w])
print(f'{w:8s}: TF={tf:.4f} IDF={idf:.4f} TF-IDF={tf*idf:.4f}')cat: TF-IDF=0.0000 (in all 3 docs, IDF=0); chases: 0.1352; rat: 0.3662
Why: 'cat' appears in every document so its IDF is log(3/3)=0 — it is a stop-word here. 'rat' appears only once, giving it the highest score.
| word | TF | IDF | TF-IDF |
|---|---|---|---|
| cat | 0.3333 | 0.0000 (log 3/3) | 0.0000 |
| chases | 0.3333 | 0.4055 (log 3/2) | 0.1352 |
| rat | 0.3333 | 1.0986 (log 3/1) | 0.3662 |
Comparison
Comparison matrix
From TF-IDF from scratch: refill the TF-IDF column from what you know. The rest of the table is as it appeared.
| word | TF | IDF | TF-IDF |
|---|---|---|---|
| cat | 0.3333 | 0.0000 (log 3/3) | 0.0000 |
| chases | 0.3333 | 0.4055 (log 3/2) | 0.1352 |
| rat | 0.3333 | 1.0986 (log 3/1) | 0.3662 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Use raw BoW counts for sentiment. 'The' appears 5 times → treat it as highly informative.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Stop-words dominate count vectors.
Apply TF-IDF weighting. 'the' in every document gets IDF=0; unique sentiment words score high.
Why: Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.
Trap
Use raw BoW counts for sentiment. 'The' appears 5 times → treat it as highly informative.
Let the model learn that high count of 'the' predicts sentiment
Why: Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.
Apply TF-IDF weighting. 'the' in every document gets IDF=0; unique sentiment words score high.
IDF zeroes out universal stop-words automatically; unique discriminating words receive large scores
Why: TF-IDF is a principled reweighting — no stop-word list required, though a custom list can remove domain stop-words the IDF doesn't catch.
Break the constraint
Discussion prompt
The rule this trap just fixed:
TF-IDF is a principled reweighting — no stop-word list required, though a custom list can remove domain stop-words the IDF doesn't catch.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Stop-words dominate count vectors. A 1 000-review corpus where 'the' appears in every review assigns it 0 IDF — no discriminative power — yet count-based models waste capacity fitting it.
Section
Part 3 of 4
Concept
A unigram treats each word independently. A bigram treats each consecutive pair as a feature, preserving local context that unigrams miss.
| n-gram type | example tokens from 'not good' | captures |
|---|---|---|
| unigram (n=1) | 'not', 'good' | individual words |
| bigram (n=2) | 'not_good' | negation context |
| trigram (n=3) | 'was_not_good' | wider dependency |
The cost: vocabulary grows exponentially. On the 6-word corpus, unigrams give 6 features, adding bigrams gives 12, adding trigrams gives 15 — with real corpora the explosion is dramatic.
Trade off
Comparison matrix
From N-grams: capturing local context: every row here is a choice with a cost. Fill the example tokens from 'not good' column, then say which row you would actually pick and what you give up for it.
| n-gram type | example tokens from 'not good' | captures |
|---|---|---|
| unigram (n=1) | 'not', 'good' | individual words |
| bigram (n=2) | 'not_good' | negation context |
| trigram (n=3) | 'was_not_good' | wider dependency |
Concept
Order matters: lowercase before stemming, stopwords before vectorizing. Stemming and lemmatization are interchangeable — pick lemmatization when morphological correctness matters (e.g., medical text).
Estimation
Predict first
Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.
Commit before you compute: what does CountVectorizer with n-grams come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: max n=1: 6; max n=2: 12; max n=3: 15
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each higher-order n-gram adds all new consecutive sequences.
Worked example
Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.
from sklearn.feature_extraction.text import CountVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
for ng in [1, 2, 3]:
cv = CountVectorizer(ngram_range=(1, ng))
cv.fit(corpus)
print(f'max n={ng}: vocab size = {len(cv.vocabulary_)}')max n=1: 6; max n=2: 12; max n=3: 15
Why: Each higher-order n-gram adds all new consecutive sequences. On real corpora (vocab ~50k) bigrams easily produce millions of features — sparsity and memory become dominant concerns.
| ngram_range | vocab size (6-word corpus) |
|---|---|
| (1,1) unigrams only | 6 |
| (1,2) + bigrams | 12 |
| (1,3) + trigrams | 15 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
max n=1: 6; max n=2: 12; max n=3: 15
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Pass ngram_range=(1,2) to capture both unigrams and bigrams. Observe vocabulary explosion.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Always remove the default NLTK stop-word list before vectorizing — it cleans noise.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.
Audit the stop-word list for domain-critical words (especially negation: 'not', 'no', 'never').
Why: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.
Trap
Always remove the default NLTK stop-word list before vectorizing — it cleans noise.
Apply stop-words blindly; 'not' is in the default stop-word list and is removed
Why: 'not good', 'not bad' become just 'good', 'bad' — two sentences with opposite sentiment now have identical vectors.
Audit the stop-word list for domain-critical words (especially negation: 'not', 'no', 'never').
Remove 'not' from the stop-word set, or use a bigram model to capture 'not_good' directly
Why: Blanket stop-word removal improves speed and sparsity, but can destroy discriminative signal. Always check which words are being dropped.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
df(t) is the number of documents containing term t; N is the total document count. A word in every document has IDF = log(1) = 0 — it is zeroed out.; A unigram treats each word independently. A bigram treats each consecutive pair as a feature, preserving local context that unigrams miss.Section
Part 4 of 4
Concept
fit builds the vocabulary and computes IDF from the training corpus. transform applies those fixed IDF values to new documents — never refit on test data (data leakage, Lesson 32).
| method | uses training vocab? | updates IDF? | call on |
|---|---|---|---|
| fit_transform | builds it | yes | training set only |
| transform | yes (frozen) | no | val/test set |
vocabulary_ stores the word → column mapping. Inspect it to debug unexpected feature columns or OOV drops.
Explain it
Discussion prompt
Explain TfidfVectorizer: fit vs transform to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
fit builds the vocabulary and computes IDF from the training corpus. transform applies those fixed IDF values to new documents — never refit on test data (data leakage, Lesson 32).
Estimation
Predict first
Wire TfidfVectorizer into a Pipeline with LogisticRegression. Fit, predict, check accuracy.
Commit before you compute: what does TF-IDF + Logistic Regression pipeline come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: acc: 1.0; vocab_size: 18
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. TF-IDF provides a clean sparse representation.
Worked example
Wire TfidfVectorizer into a Pipeline with LogisticRegression. Fit, predict, check accuracy.
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
import warnings; warnings.filterwarnings('ignore')
train = ["I love this film excellent", "great movie wonderful story",
"amazing performance brilliant",
"terrible film awful", "bad movie boring", "horrible garbage waste"]
y_tr = [1, 1, 1, 0, 0, 0]
test = ["excellent film loved it", "terrible awful waste", "great performance"]
y_te = [1, 0, 1]
pipe = Pipeline([('tfidf', TfidfVectorizer()),
('clf', LogisticRegression(max_iter=1000, random_state=0))])
pipe.fit(train, y_tr)
print('acc:', accuracy_score(y_te, pipe.predict(test)))
print('vocab size:', len(pipe.named_steps['tfidf'].vocabulary_))acc: 1.0; vocab_size: 18
Why: TF-IDF provides a clean sparse representation. Logistic regression (Lesson 40's pattern) over the TF-IDF features classifies sentiment correctly on this toy corpus.
| test doc | prediction | true |
|---|---|---|
| excellent film loved it | 1 (positive) | 1 |
| terrible awful waste | 0 (negative) | 0 |
| great performance | 1 (positive) | 1 |
Comparison
Comparison matrix
From TF-IDF + Logistic Regression pipeline: refill the prediction column from what you know. The rest of the table is as it appeared.
| test doc | prediction | true |
|---|---|---|
| excellent film loved it | 1 (positive) | 1 |
| terrible awful waste | 0 (negative) | 0 |
| great performance | 1 (positive) | 1 |
Concept
TF-IDF is a bag: it destroys word order. 'Dog bites man' and 'Man bites dog' produce identical vectors if the vocabulary counts are the same.
These failures motivate word embeddings (Lesson 65+), where 'happy' and 'joyful' map to nearby vectors and word order is preserved via position encodings.
Analogy
Discussion prompt
Explain What TF-IDF cannot do by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
TF-IDF is a bag: it destroys word order. 'Dog bites man' and 'Man bites dog' produce identical vectors if the vocabulary counts are the same.
Ranking
Put in order
These are the steps of The text vectorization recipe, scrambled. Put them back in order before the next slide shows you.
CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoWngram_range=(1,2) for bigrams — watch vocab explosionfit_transform(train), then transform(test) — no data leakagePipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistentWhy: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.
Pattern
CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoWngram_range=(1,2) for bigrams — watch vocab explosionfit_transform(train), then transform(test) — no data leakagePipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistentEdge cases
Discussion prompt
The text vectorization recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
CountVectorizer for raw BoW; TfidfVectorizer for IDF-reweighted BoWngram_range=(1,2) for bigrams — watch vocab explosionfit_transform(train), then transform(test) — no data leakagePipeline([('tfidf', ...), ('clf', ...)]) keeps fit/transform consistentElimination
Eliminate the wrong options
A corpus has N=1000 documents. The word 'the' appears in all 1000. What is IDF('the') using the formula IDF = log(N/df)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: IDF = log(N/df) = log(1000/1000) = log(1) = 0. A word in every document contributes zero weight — exactly the stop-word suppression TF-IDF is designed for.
Check
Work out the IDF before clicking.
Check your understanding
A corpus has N=1000 documents. The word 'the' appears in all 1000. What is IDF('the') using the formula IDF = log(N/df)?
Answer: A
Why: IDF = log(N/df) = log(1000/1000) = log(1) = 0. A word in every document contributes zero weight — exactly the stop-word suppression TF-IDF is designed for.
Prediction
Predict first
You fit a TfidfVectorizer on the full dataset (train + test), then split and train the classifier. What is wrong?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The IDF values were computed using test documents, leaking test distribution into training
Why: Fitting on train+test means the IDF for each token is computed using test documents. At deployment you won't have test data, so IDF values will differ — this is data leakage. Always fit_transform on train only, then transform test.
Check
Data leakage or correct?
Check your understanding
You fit a TfidfVectorizer on the full dataset (train + test), then split and train the classifier. What is wrong?
Answer: A
Why: Fitting on train+test means the IDF for each token is computed using test documents. At deployment you won't have test data, so IDF values will differ — this is data leakage. Always fit_transform on train only, then transform test.
Elimination
Eliminate the wrong options
Switching from unigrams to unigrams+bigrams (ngram_range=(1,2)) on a large corpus primarily causes:
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Adding bigrams multiplies the feature space. On the 6-word toy corpus unigrams→bigrams grows from 6 to 12 features; on a 50k-word corpus the explosion is enormous. Vectors remain sparse (most bigrams don't appear in a given doc) but memory and training time increase substantially.
Check
What does adding bigrams buy you, and what does it cost?
Check your understanding
Switching from unigrams to unigrams+bigrams (ngram_range=(1,2)) on a large corpus primarily causes:
Answer: A
Why: Adding bigrams multiplies the feature space. On the 6-word toy corpus unigrams→bigrams grows from 6 to 12 features; on a 50k-word corpus the explosion is enormous. Vectors remain sparse (most bigrams don't appear in a given doc) but memory and training time increase substantially.
Section
Project
Concept
Build TF-IDF from first principles with Python dicts, then wire it into an sklearn pipeline and classify sentiment.
| # | requirement | tool |
|---|---|---|
| 1 | TF-IDF from scratch (dicts) | math.log + Counter |
| 2 | sklearn vectorizers — BoW and TF-IDF | CountVectorizer, TfidfVectorizer |
| 3 | full pipeline: TF-IDF + LR | Pipeline, LogisticRegression |
Build rules: never call fit_transform on test data; inspect vocabulary_; verify that a word in every document scores TF-IDF = 0.
Counterexample
Discussion prompt
Build TF-IDF from first principles with Python dicts, then wire it into an sklearn pipeline and classify sentiment.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: never call fit_transform on test data; inspect vocabulary_; verify that a word in every document scores TF-IDF = 0.
Worked example
Your turn: implement TF-IDF with only math and built-ins. Predict the TF-IDF of 'rat' in doc0.
Hint: TF = count(word, doc) / len(doc); IDF = log(N / df[word]); df = number of docs containing word.
import math
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
tok = [d.split() for d in corpus]
N = len(tok)
df = {w: sum(1 for d in tok if w in d) for w in set(w for d in tok for w in d)}
doc, L = tok[0], len(tok[0])
for w in sorted(set(doc)):
tf = doc.count(w) / L
idf = math.log(N / df[w])
print(f'{w:8s}: TF={tf:.4f} IDF={idf:.4f} TF-IDF={tf*idf:.4f}')| word | TF | IDF | TF-IDF |
|---|---|---|---|
| cat | 0.3333 | 0.0000 | 0.0000 |
| chases | 0.3333 | 0.4055 | 0.1352 |
| rat | 0.3333 | 1.0986 | 0.3662 |
Trade off
Comparison matrix
From Milestone 1 — TF-IDF from scratch: every row here is a choice with a cost. Fill the TF column, then say which row you would actually pick and what you give up for it.
| word | TF | IDF | TF-IDF |
|---|---|---|---|
| cat | 0.3333 | 0.0000 | 0.0000 |
| chases | 0.3333 | 0.4055 | 0.1352 |
| rat | 0.3333 | 1.0986 | 0.3662 |
Worked example
Your turn: fit CountVectorizer and TfidfVectorizer on the same corpus. Inspect vocabulary and compare doc0 representations.
Hint: .fit_transform(corpus).toarray() then .get_feature_names_out(). Note sklearn adds +1 smoothing to IDF by default (smooth_idf=True) — values will differ slightly from your scratch version.
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
corpus = ["cat chases rat", "cat eats fish", "dog chases cat"]
cv = CountVectorizer()
X_bow = cv.fit_transform(corpus).toarray()
print('BoW vocab:', cv.get_feature_names_out())
print('BoW doc0: ', X_bow[0])
tv = TfidfVectorizer(norm=None, smooth_idf=False)
X_tf = tv.fit_transform(corpus).toarray()
print('TF-IDF doc0 (no smooth):', X_tf[0].round(4))| feature | BoW doc0 | TF-IDF doc0 (no smooth) |
|---|---|---|
| cat | 1 | 1.0000 |
| chases | 1 | 1.4055 |
| dog | 0 | 0.0000 |
| eats | 0 | 0.0000 |
| fish | 0 | 0.0000 |
| rat | 1 | 2.0986 |
Discrimination
Sort into buckets
Sort these by BoW doc0, from memory, without looking back at Milestone 2 — sklearn vectorizers. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Your turn: wire TfidfVectorizer + LogisticRegression into a Pipeline. Predict whether the pipeline achieves perfect accuracy on the toy set.
Hint: Pipeline([('tfidf', TfidfVectorizer()), ('clf', LogisticRegression(...))]); fit on train, predict on test, compute accuracy_score.
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
import warnings; warnings.filterwarnings('ignore')
train=["I love this film excellent","great movie wonderful story",
"amazing performance brilliant",
"terrible film awful","bad movie boring","horrible garbage waste"]
y_tr=[1,1,1,0,0,0]
test=["excellent film loved it","terrible awful waste","great performance"]
y_te=[1,0,1]
pipe=Pipeline([('tfidf',TfidfVectorizer()),
('clf', LogisticRegression(max_iter=1000, random_state=0))])
pipe.fit(train,y_tr)
print('acc:', accuracy_score(y_te, pipe.predict(test)))
print('vocab size:', len(pipe.named_steps['tfidf'].vocabulary_))| test doc | pred | true |
|---|---|---|
| excellent film loved it | 1 | 1 |
| terrible awful waste | 0 | 0 |
| great performance | 1 | 1 |
Comparison
Comparison matrix
From Milestone 3 — full classification pipeline: refill the pred column from what you know. The rest of the table is as it appeared.
| test doc | pred | true |
|---|---|---|
| excellent film loved it | 1 | 1 |
| terrible awful waste | 0 | 0 |
| great performance | 1 | 1 |
Concept
Out loud, slides closed: (1) derive the TF-IDF formula from scratch and explain why IDF zeroes out stop-words; (2) explain what happens to 'rabbit' in a test document if 'rabbit' never appeared in training; (3) name two things TF-IDF cannot represent.
Stretch (homework): compare BoW vs TF-IDF vs bigrams on a sentiment dataset; implement TF-IDF that handles repeated terms across docs (smooth_idf=True); analyze which features logistic regression weighted highest with coef_. Next up: word embeddings — dense vectors that capture semantics.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Bag of Words — counting tokens · TF-IDF — rewarding rare words · N-grams and text preprocessing · sklearn vectorizers and the full pipeline · Your turn: build it. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
ngram_range=(1,2) for bigrams and predict vocabulary growthTfidfVectorizer + LogisticRegression in an sklearn Pipeline without data leakage| idea | the one thing to remember |
|---|---|
| BoW | counts per word; order discarded; sparse |
| TF-IDF | TF × IDF; IDF = log(N/df); stop-words → 0 |
| n-grams | bigrams capture context; vocab explodes exponentially |
| fit vs transform | fit on train only; transform test (no leakage) |
| limitations | no order, no synonyms, no semantics → embeddings next |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.