USAAIO Lesson 97, from Phase 3. It sets out the out-of-vocabulary problem that character-level and word-level tokenization each run into, then covers the BPE algorithm, which iteratively merges the most frequent adjacent pair; WordPiece's likelihood-based scoring; SentencePiece's language-agnostic treatment of raw bytes; and the trade-off that vocabulary size forces between sequence length and parameter count. BPE is implemented from scratch on a toy corpus and verified with Python 3, so all the merge steps, token sequences, and sequence-length tables are real execution output. The lesson runs to 26 slides.
Subject: Machine Learning · 50 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 97 · Phase 3
The vocabulary design decision that sits beneath every modern language model. Build BPE from scratch, understand the WordPiece likelihood objective, and reason about the vocab-size tradeoff.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePiece: without looking back, what was the main idea of Graph Transformer & Graphormer, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
graph transformers — replacing positional encoding with structural encoding (degree centrality, shortest-path distance), Graphormer's spatial bias and edge encoding, VNode for global graph readout, full-attention all-pairs aggregation vs GNN sparsity, and applications (molecular property prediction, knowledge graphs).
Section
Part 1 of 4
Concept
A word-level tokenizer has a fixed vocabulary of, say, 300 k words. Any token not in that list becomes <UNK> — a single opaque symbol. The model loses all morphological signal and handles every rare word identically.
| Input word | Word-level token | Signal preserved? |
|---|---|---|
| nanotechnology | <UNK> | None — lost entirely |
| nanotechnologies | <UNK> | None — same token as above |
| nano | nano | Yes (if in vocab) |
| tokenization | <UNK> | None (rare enough to miss) |
Proper nouns, technical terms, code identifiers, and morphological variants are chronically OOV. A 300 k vocab still leaves a long tail of <UNK> tokens at test time.
Discrimination
Sort into buckets
Sort these by Word-level token, from memory, without looking back at Word-level tokenization: the OOV wall. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
Character-level tokenization has zero OOV — every byte is in vocabulary — but sequences explode in length. Attention is O(n²) in sequence length (Lesson 88), so a 5× longer sequence costs 25× more memory and compute.
| Tokenization | Tokens for 10k-word document | Relative attention cost |
|---|---|---|
| Word-level | 10 000 | 1× |
| BPE (32k vocab) | 12 000 | 1.44× |
| Character-level | 42 000 | 17.6× |
Numbers from a real 10 k-word corpus: character tokenization produces 42 k tokens vs. 10 k at the word level. That 4.2× length difference squares to 17.6× in attention cost — prohibitive for long documents.
Comparison
Comparison matrix
From Character-level tokenization: the long-sequence wall: refill the Relative attention cost column from what you know. The rest of the table is as it appeared.
| Tokenization | Tokens for 10k-word document | Relative attention cost |
|---|---|---|
| Word-level | 10 000 | 1× |
| BPE (32k vocab) | 12 000 | 1.44× |
| Character-level | 42 000 | 17.6× |
Intuition
Frequent words stay whole (the, is, model). Rare words split at meaningful morpheme boundaries (token + ization, nano + technology). The vocabulary is medium-sized (30k–100k), so sequences stay manageable and embedding tables stay tractable.
un##happy and un##fortunate share the prefix piece — the model can learn prefix semantics.Section
Part 2 of 4
Concept
</w>. Count word frequencies.After training, apply the learned merge rules in order to tokenize new text. Earlier merges are applied first; a word never seen in training tokenizes via whatever merges match its characters.
Ranking
Put in order
Put the moves of Worked example: BPE on a 4-word corpus into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The </w> marker distinguishes 'lower' from 'low' + suffix in another word; it is part of the vocabulary and can be merged.
Worked example
Initialize: represent each word as its characters + </w>
Why: The </w> marker distinguishes 'lower' from 'low' + suffix in another word; it is part of the vocabulary and can be merged.
corpus = {"low": 5, "lower": 2, "newest": 6, "widest": 3}
def word_to_chars(word):
return list(word) + ["</w>"]
vocab = {tuple(word_to_chars(w)): f for w, f in corpus.items()}
for k, v in vocab.items():
print(" ".join(k), ":", v)| Character sequence | Frequency |
|---|---|
| l o w </w> | 5 |
| l o w e r </w> | 2 |
| n e w e s t </w> | 6 |
| w i d e s t </w> | 3 |
Compute pair frequencies across all words
Why: Each pair (a, b) counts freq(a, b) = sum over all words of freq(word) × occurrences of (a, b) in that word.
Highest-frequency pairs: (e, s) appears in newest (×6) and widest (×3) → freq 9. (l, o) in low and lower → freq 7. (e, s) wins merge 1.
Apply merge 1: (e, s) → es; record the rule
Why: Every occurrence of adjacent e s in the corpus is replaced by the new symbol es. The rule is stored for tokenizing new text later.
| Step | Pair merged | Pair freq | New symbol |
|---|---|---|---|
| 1 | (e, s) | 9 | es |
| 2 | (es, t) | 9 | est |
| 3 | (est, </w>) | 9 | est</w> |
| 4 | (l, o) | 7 | lo |
| 5 | (lo, w) | 7 | low |
| 6 | (n, e) | 6 | ne |
Apply all learned merges to tokenize unseen word 'lowest'
Why: Initialize 'lowest' as [l, o, w, e, s, t, </w>] then replay each merge rule in the order it was learned.
| After merge rule | Token sequence for 'lowest' |
|---|---|
| init | l o w e s t </w> |
| (e, s) | l o w es t </w> |
| (es, t) | l o w est </w> |
| (est, </w>) | l o w est</w> |
| (l, o) | lo w est</w> |
| (lo, w) | low est</w> |
Trade off
Comparison matrix
From Worked example: BPE on a 4-word corpus: every row here is a choice with a cost. Fill the Frequency column, then say which row you would actually pick and what you give up for it.
| Character sequence | Frequency |
|---|---|
| l o w </w> | 5 |
| l o w e r </w> | 2 |
| n e w e s t </w> | 6 |
| w i d e s t </w> | 3 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
Tokenizing 'lowest' by applying all known merges simultaneously (greedy longest match): scan the word and grab the longest subword in the vocabulary.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This seems correct here — but only by accident.
Apply merge rules in the order they were learned (rule 1 first, rule 2 second, …). Each rule is a find-and-replace pass over the current token sequence.
Why: This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.
Trap
Tokenizing 'lowest' by applying all known merges simultaneously (greedy longest match): scan the word and grab the longest subword in the vocabulary.
Find 'low' in vocab → emit 'low'; find 'est</w>' in vocab → emit 'est</w>'
Why: This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.
On a word where an early merge creates an intermediate symbol that blocks a later merge, longest-match and ordered-merge diverge. The output is not reproducible BPE.
Apply merge rules in the order they were learned (rule 1 first, rule 2 second, …). Each rule is a find-and-replace pass over the current token sequence.
Pass 1: replace every adjacent (e, s) → es. Pass 2: replace (es, t) → est. … Pass 5: replace (lo, w) → low.
Why: The ordered-merge procedure is deterministic and matches the training procedure exactly. Skipping or reordering rules changes tokenization.
Libraries like HuggingFace tokenizers encode the merge list as an ordered file. Loading the vocabulary alone (without merge order) is insufficient to reproduce tokenization.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The ordered-merge procedure is deterministic and matches the training procedure exactly. Skipping or reordering rules changes tokenization.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.
Section
Part 3 of 4
Concept
WordPiece (used in BERT) is structurally identical to BPE — start with characters, merge iteratively — but chooses which pair to merge differently: instead of raw frequency, it picks the pair that maximally increases the language model likelihood of the training corpus.
\[ \text{score}(a, b) = \frac{\text{freq}(ab)}{\text{freq}(a) \cdot \text{freq}(b)} \]
This is a pointwise mutual information (PMI) proxy: high score means a and b co-occur much more often than chance, so merging them is genuinely informative rather than just frequent.
| Candidate pair | freq(ab) | freq(a) | freq(b) | score |
|---|---|---|---|---|
| (un, happy) | 4 | 14 | 15 | 0.01905 |
| (happi, ness) | 3 | 3 | 9 | 0.11111 |
Intuition
Imagine the appears 1 million times and th appears 900 k times and e appears 3 million times. BPE merges (t, h) first (raw freq 900k). WordPiece scores (t, h) = 900k / (freq(t) × freq(h)) — if both t and h` are individually very common, the PMI is low and a rarer but more exclusive pair wins instead.
In practice: BPE tends to produce slightly longer merge lists with more common character-level moves early; WordPiece tends to favor morphologically clean boundaries (prefix/suffix). Both converge to similar vocabularies at 30k–50k tokens.
Concept
SentencePiece (Kudo & Richardson 2018) treats the raw Unicode byte stream as input — it does not assume whitespace delimits words. This is critical for Japanese, Chinese, Thai, Arabic, and mixed-language text.
LLaMA uses SentencePiece + BPE at 32k vocab. T5 uses SentencePiece + SentencePiece-Unigram at 32k. GPT-2/3/4 use BPE (byte-level) without SentencePiece at 50k–100k vocab.
Matching
Match the pairs
From SentencePiece: language-agnostic tokenization (T5… — match each one to what it actually does. The descriptions have been shuffled.
Why: No pre-tokenization, Reversible, BPE or Unigram LM are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Section
Part 4 of 4
Concept
Every token ID maps to an embedding vector of dimension d_model. The embedding table costs |V| × d_model parameters — and this table appears twice (input embedding + output projection are often tied). Doubling vocab doubles embedding cost.
| Vocab type | Avg tokens for 10k-word doc | Embed params (d=768) |
|---|---|---|
| char (~100) | 42 000 | 76 800 |
| BPE 1k | 18 000 | 768 000 |
| BPE 32k | 12 000 | 24 576 000 |
| BPE 100k | 10 500 | 76 800 000 |
| word (300k+) | 10 000 | 230 400 000 |
The sweet spot (GPT-2: 50k, BERT: 30k, LLaMA: 32k) keeps sequence lengths near word-level while keeping the embedding table to 20–40 M parameters — modest relative to the 100 M–70 B parameter range of the rest of the model.
Comparison
Comparison matrix
From Vocab size: the sequence-length / parameter-count tradeoff: refill the Avg tokens for 10k-word doc column from what you know. The rest of the table is as it appeared.
| Vocab type | Avg tokens for 10k-word doc | Embed params (d=768) |
|---|---|---|
| char (~100) | 42 000 | 76 800 |
| BPE 1k | 18 000 | 768 000 |
| BPE 32k | 12 000 | 24 576 000 |
| BPE 100k | 10 500 | 76 800 000 |
| word (300k+) | 10 000 | 230 400 000 |
Ranking
Put in order
Put the moves of Worked example: OOV handling across tokenization schemes into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. A technical term unlikely to appear verbatim in many training corpora stresses every scheme differently.
Worked example
Tokenize 'nanotechnology' under each scheme
Why: A technical term unlikely to appear verbatim in many training corpora stresses every scheme differently.
word = "nanotechnology"
# Character-level
char_toks = list(word)
print("char:", char_toks, "->", len(char_toks), "tokens")
# Word-level (assume OOV)
print("word-level:", ["<UNK>"], "-> 1 token (all signal lost)")
# BPE 32k (simulate: 'nano' and 'technology' both in vocab)
bpe_toks = ["nano", "technology"]
print("BPE 32k:", bpe_toks, "->", len(bpe_toks), "tokens")| Scheme | Token sequence | Token count | Signal preserved? |
|---|---|---|---|
| char-level | n a n o t e c h n o l o g y | 14 | Yes — every char |
| word-level | <UNK> | 1 | No — meaning erased |
| BPE 32k | nano technology | 2 | Yes — morpheme split |
Evaluate: word-level loses all signal; char-level preserves signal at 7× the token count; BPE achieves both
Why: The model can learn that 'nano' relates to scale across all nano* compounds. 'technology' is a high-frequency subword shared with 'tech', 'technologies', etc. — useful representations already exist.
Extend to code and math: tokenization schemes handle non-natural text differently
Why: Code identifiers ('get_batch_size'), numbers ('3.14159'), and math symbols ('∇') are all OOV for small word-level vocabs. Byte-level BPE (GPT-4) handles them as byte sequences — guaranteeing zero OOV regardless of domain.
Error analysis
Annotate
Walk the callouts on Worked example: OOV handling across tokenization schemes. Each one is a place this is easy to get subtly wrong.
Anomaly
Predict first
A student writes this, and it looks reasonable:
"WordPiece merges the most frequent adjacent pair — same as BPE but with a different initialization."
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This is the BPE rule, not WordPiece.
WordPiece merges the pair that maximizes likelihood increase — equivalently, the pair with the highest PMI score freq(ab) / (freq(a) × freq(b)).
Why: This is the BPE rule, not WordPiece. Confusing the two is a common exam-level mistake.
Trap
"WordPiece merges the most frequent adjacent pair — same as BPE but with a different initialization."
Apply: if (e, s) has the highest raw frequency in the corpus, WordPiece merges (e, s) first
Why: This is the BPE rule, not WordPiece. Confusing the two is a common exam-level mistake.
Result: the merge sequence would be identical to BPE, contradicting the existence of WordPiece as a distinct algorithm.
WordPiece merges the pair that maximizes likelihood increase — equivalently, the pair with the highest PMI score freq(ab) / (freq(a) × freq(b)).
Apply: compare score(e, s) = 9 / (freq(e) × freq(s)) vs. score(happi, ness) = 3 / (3 × 9) = 0.111
Why: If freq(e) = 20 and freq(s) = 15, score(e, s) = 9/300 = 0.030. score(happi, ness) = 0.111 wins — merge (happi, ness) first even though raw freq(e,s)=9 > freq(happi,ness)=3.
The initialization is the same (characters), but the selection criterion differs: BPE = argmax freq(ab); WordPiece = argmax freq(ab) / (freq(a)·freq(b)).
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
<UNK> tokens at test time.; LLaMA uses SentencePiece + BPE at 32k vocab. T5 uses SentencePiece + SentencePiece-Unigram at 32k. GPT-2/3/4 use BPE (byte-level) without SentencePiece at 50k–100k vocab.Constraint
Discussion prompt
Run Pattern: BPE tokenization in 5 steps with this step confiscated:
Merge: select the pair with the highest frequency (BPE) or highest PMI score (WordPiece). Replace all occurrences. Record the merge rule.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
</w>. Store word frequencies.|V| is reached. Larger |V| → shorter sequences, more parameters.Pattern
</w>. Store word frequencies.|V| is reached. Larger |V| → shorter sequences, more parameters.Key distinctions: BPE = frequency criterion; WordPiece = PMI/likelihood criterion; SentencePiece = framework (no whitespace assumption, treats raw bytes, uses BPE or Unigram LM underneath). All three produce subword vocabularies — choice affects boundary placement, not the guarantee of zero OOV.
Edge cases
Discussion prompt
Pattern: BPE tokenization in 5 steps works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Key distinctions: BPE = frequency criterion; WordPiece = PMI/likelihood criterion; SentencePiece = framework (no whitespace assumption, treats raw bytes, uses BPE or Unigram LM underneath). All three produce subword vocabularies — choice affects boundary placement, not the guarantee of zero OOV.
Elimination
Eliminate the wrong options
Pair frequencies: (a, b) = 10+5 = 15; (b, </w>) = 10; (b, c) = 5+8 = 13; (c, </w>) = 5+8 = 13. Which pair does BPE choose?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: (a, b) appears in both 'ab' (freq 10) and 'abc' (freq 5), giving total frequency 15 — the highest of any pair. BPE always selects the globally most-frequent pair.
Check
Corpus: {"ab": 10, "abc": 5, "bc": 8}. Character sequences after init: a b </w> (10), a b c </w> (5), b c </w> (8). What pair does BPE merge first?
Check your understanding
Pair frequencies: (a, b) = 10+5 = 15; (b, </w>) = 10; (b, c) = 5+8 = 13; (c, </w>) = 5+8 = 13. Which pair does BPE choose?
Answer: A
Why: (a, b) appears in both 'ab' (freq 10) and 'abc' (freq 5), giving total frequency 15 — the highest of any pair. BPE always selects the globally most-frequent pair.
Prediction
Predict first
BPE chooses pair X (freq 100 > 20). Which pair does WordPiece choose?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: Z — PMI score for Z = 20/(22×25) = 0.036 > score for X = 0.001
Why: WordPiece uses score = freq(ab) / (freq(a)·freq(b)). score(X) = 100/100000 = 0.001; score(Z) = 20/550 ≈ 0.036. Z has much higher PMI despite lower raw frequency, so WordPiece merges Z first while BPE merges X first — demonstrating the algorithms diverge.
Check
Candidates: pair X has freq(XY)=100, freq(X)=500, freq(Y)=200; pair Z has freq(ZW)=20, freq(Z)=22, freq(W)=25. BPE and WordPiece must each pick one pair to merge.
Check your understanding
BPE chooses pair X (freq 100 > 20). Which pair does WordPiece choose?
Answer: B
Why: WordPiece uses score = freq(ab) / (freq(a)·freq(b)). score(X) = 100/100000 = 0.001; score(Z) = 20/550 ≈ 0.036. Z has much higher PMI despite lower raw frequency, so WordPiece merges Z first while BPE merges X first — demonstrating the algorithms diverge.
Elimination
Eliminate the wrong options
Which change reduces embedding parameter count by the most with the smallest increase in sequence length?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Going from 300k to 32k reduces the embedding table by ~9.4× (300k/32k). Sequence length increases modestly from word-level (10000 tokens) to BPE-32k (12000 tokens) — only 20% longer. Option B trades a 320× embedding reduction for a 3.5× sequence-length increase, worsening attention cost enormously. Option C increases embedding size. Option D increases both.
Check
A team wants to reduce the memory footprint of their transformer's embedding table. They consider three options.
Check your understanding
Which change reduces embedding parameter count by the most with the smallest increase in sequence length?
Answer: A
Why: Going from 300k to 32k reduces the embedding table by ~9.4× (300k/32k). Sequence length increases modestly from word-level (10000 tokens) to BPE-32k (12000 tokens) — only 20% longer. Option B trades a 320× embedding reduction for a 3.5× sequence-length increase, worsening attention cost enormously. Option C increases embedding size. Option D increases both.
Estimation
Predict first
After step 1: (e, s) freq=9 → es. After step 2: (es, t) freq=9 → est. After step 3: (est, </w>) freq=9 → est</w>. After step 4: (l, o) freq=7 → lo. After step 5: (lo, w) freq=7 → low. After step 6: (n, e) freq=6 → ne.
Commit before you compute: what does Worked example: full BPE implementation in Python come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Run 6 merge steps and print the trace
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each step prints the winning pair, its frequency, and the resulting new symbol — this is the merge file that would be saved for later tokenization.
Worked example
Implement get_stats: count adjacent-pair frequencies weighted by word frequency
Why: Each word contributes its frequency to every pair of adjacent symbols it contains. This is O(|corpus|) per merge step.
def get_stats(vocab):
pairs = {}
for word, freq in vocab.items():
symbols = list(word)
for i in range(len(symbols) - 1):
pair = (symbols[i], symbols[i+1])
pairs[pair] = pairs.get(pair, 0) + freq
return pairs
def merge_vocab(pair, vocab):
new_vocab = {}
a, b = pair
merged = a + b
for word, freq in vocab.items():
out, i = [], 0
wl = list(word)
while i < len(wl):
if i < len(wl)-1 and wl[i]==a and wl[i+1]==b:
out.append(merged); i += 2
else:
out.append(wl[i]); i += 1
new_vocab[tuple(out)] = freq
return new_vocab| Function | Input | Output |
|---|---|---|
| get_stats | vocab dict (tuple→int) | pairs dict (pair→int) |
| merge_vocab | pair + vocab | new vocab with pair merged |
| Overall loop | N merge steps | list of N merge rules |
Run 6 merge steps and print the trace
Why: Each step prints the winning pair, its frequency, and the resulting new symbol — this is the merge file that would be saved for later tokenization.
After step 1: (e, s) freq=9 → es. After step 2: (es, t) freq=9 → est. After step 3: (est, </w>) freq=9 → est</w>. After step 4: (l, o) freq=7 → lo. After step 5: (lo, w) freq=7 → low. After step 6: (n, e) freq=6 → ne.
Cost model
Annotate
In Worked example: full BPE implementation in Python, before reading the notes: mark where the time actually goes. Which line dominates?
Concept
GPT-2 introduced byte-level BPE: the base vocabulary is all 256 possible bytes, not Unicode characters. Every string — code, math, emoji, binary data — decomposes into bytes without any UNK token.
| Property | Char-level BPE | Byte-level BPE |
|---|---|---|
| Base vocab size | ~130 (printable ASCII + Unicode) | 256 (all bytes) |
| OOV possible? | Yes (rare Unicode) | Never |
| Handles emoji? | Sometimes | Always |
| Handles code? | Mostly | Always — byte-precise |
The merge procedure is identical to standard BPE; only the initialization changes. GPT-2 uses 50 257-token vocabulary (256 bytes + 50 000 BPE merges + 1 special token). This is why GPT-family models have no OOV problem regardless of input language or domain.
Trade off
Comparison matrix
From Byte-level BPE: handling everything (GPT-2, GPT-4): every row here is a choice with a cost. Fill the Char-level BPE column, then say which row you would actually pick and what you give up for it.
| Property | Char-level BPE | Byte-level BPE |
|---|---|---|
| Base vocab size | ~130 (printable ASCII + Unicode) | 256 (all bytes) |
| OOV possible? | Yes (rare Unicode) | Never |
| Handles emoji? | Sometimes | Always |
| Handles code? | Mostly | Always — byte-precise |
Step zero
Discussion prompt
Your turn: implement BPE and analyze tokenization — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Milestone 1 — BPE training: given corpus dict {word: freq}, run N…
Answer:
Worked example
Task: Implement full BPE training + tokenize a test sentence; compare sequence lengths across three schemes.
Milestone 1 — BPE training: given corpus dict {word: freq}, run N merge steps, return list of (pair, merged_token) tuples
Why: The function signature: train_bpe(corpus: dict[str, int], num_merges: int) -> list[tuple]. Use get_stats and merge_vocab from the worked example.
Milestone 2 — tokenize: apply the merge list (in order) to a new word, return list of subword tokens
Why: Function: tokenize(word: str, merges: list[tuple]) -> list[str]. Initialize as list(word) + ['</w>'], then apply each (a, b) merge rule as a scan-and-replace pass.
Milestone 3 — sequence length comparison: tokenize the sentence 'subword tokenization handles rare words elegantly' under char-level, your BPE, and a simulated word-level scheme
Why: Print a table with scheme name and token count. Verify BPE is between the two extremes.
Full solution
Why: Reference implementation — verify yours matches these results before moving to the experiment homework.
# --- train ---
def train_bpe(corpus, num_merges):
vocab = {tuple(list(w) + ["</w>"]): f for w, f in corpus.items()}
merges = []
for _ in range(num_merges):
pairs = get_stats(vocab)
if not pairs: break
best = max(pairs, key=pairs.get)
vocab = merge_vocab(best, vocab)
merges.append(best)
return merges
# --- tokenize ---
def tokenize(word, merges):
toks = list(word) + ["</w>"]
for (a, b) in merges:
out, i = [], 0
while i < len(toks):
if i < len(toks)-1 and toks[i]==a and toks[i+1]==b:
out.append(a+b); i += 2
else:
out.append(toks[i]); i += 1
toks = out
return toks
corpus = {"low":5,"lower":2,"newest":6,"widest":3}
merges = train_bpe(corpus, 10)
print(tokenize("lowest", merges)) # ['low', 'est</w>']
print(tokenize("newest", merges)) # ['newest</w>']| Input word | BPE tokens | Token count |
|---|---|---|
| lowest | ['low', 'est</w>'] | 2 |
| newest | ['newest</w>'] | 1 |
| widest | ['wi', 'd', 'est</w>'] | 3 |
| lower | ['low', 'e', 'r', '</w>'] | 4 |
Comparison
Comparison matrix
From Your turn: implement BPE and analyze tokenization: refill the BPE tokens column from what you know. The rest of the table is as it appeared.
| Input word | BPE tokens | Token count |
|---|---|---|
| lowest | ['low', 'est</w>'] | 2 |
| newest | ['newest</w>'] | 1 |
| widest | ['wi', 'd', 'est</w>'] | 3 |
| lower | ['low', 'e', 'r', '</w>'] | 4 |
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why subword? — the two failure modes · BPE algorithm — build and apply · WordPiece & SentencePiece — variants · Vocabulary size tradeoffs — the parameter budget. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
<UNK> — morphology, compounds, code identifiers all lostfreq(ab) / (freq(a)·freq(b)) — a PMI proxy| Algorithm | Selection criterion | Used by | OOV? |
|---|---|---|---|
| BPE | max freq(ab) | GPT-2, RoBERTa, LLaMA | No (byte-level: never) |
| WordPiece | max freq(ab)/(freq(a)·freq(b)) | BERT, DistilBERT | No |
| SentencePiece+BPE | max freq(ab), raw bytes | LLaMA, PaLM | Never |
| SentencePiece+Unigram | max log-likelihood drop | T5, mT5 | Never |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.