Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePiece

USAAIO Lesson 97, from Phase 3. It sets out the out-of-vocabulary problem that character-level and word-level tokenization each run into, then covers the BPE algorithm, which iteratively merges the most frequent adjacent pair; WordPiece's likelihood-based scoring; SentencePiece's language-agnostic treatment of raw bytes; and the trade-off that vocabulary size forces between sequence length and parameter count. BPE is implemented from scratch on a toy corpus and verified with Python 3, so all the merge steps, token sequences, and sequence-length tables are real execution output. The lesson runs to 26 slides.

Subject: Machine Learning · 50 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Subword Tokenization BPE · WordPiece · SentencePiece

Title

USAAIO · Lesson 97 · Phase 3

The vocabulary design decision that sits beneath every modern language model. Build BPE from scratch, understand the WordPiece likelihood objective, and reason about the vocab-size tradeoff.

2. By the end of this lesson you can

Objectives

  1. State the OOV problem for word-level tokenization and the long-sequence problem for character-level tokenization
  2. Implement BPE from scratch: initialize with characters, collect pair frequencies, merge greedily, apply learned merges to new text
  3. Explain how WordPiece differs from BPE: it merges the pair that maximizes likelihood increase, not raw frequency
  4. Describe SentencePiece (T5, LLaMA): language-agnostic, treats raw bytes, no whitespace assumption
  5. Reason about the vocab-size tradeoff — larger vocab → shorter sequences but bigger embedding table — and recall the 30k–100k typical range

3. What survived from Graph Transformer & Graphormer?

Warm-up

Discussion prompt

Before we open Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePiece: without looking back, what was the main idea of Graph Transformer & Graphormer, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

graph transformers — replacing positional encoding with structural encoding (degree centrality, shortest-path distance), Graphormer's spatial bias and edge encoding, VNode for global graph readout, full-attention all-pairs aggregation vs GNN sparsity, and applications (molecular property prediction, knowledge graphs).

4. Why subword? — the two failure modes

Section

Part 1 of 4

5. Word-level tokenization: the OOV wall

Concept

A word-level tokenizer has a fixed vocabulary of, say, 300 k words. Any token not in that list becomes <UNK> — a single opaque symbol. The model loses all morphological signal and handles every rare word identically.

Input wordWord-level tokenSignal preserved?
nanotechnology<UNK>None — lost entirely
nanotechnologies<UNK>None — same token as above
nanonanoYes (if in vocab)
tokenization<UNK>None (rare enough to miss)

Proper nouns, technical terms, code identifiers, and morphological variants are chronically OOV. A 300 k vocab still leaves a long tail of <UNK> tokens at test time.

6. Which is which, by Word-level token

Discrimination

Sort into buckets

Sort these by Word-level token, from memory, without looking back at Word-level tokenization: the OOV wall. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

<UNK>
nanotechnology; nanotechnologies; tokenization
nano
nano
g1
Word-level token is "<UNK>" for nanotechnology, nanotechnologies, tokenization — that is what the table on "Word-level tokenization: the OOV wall" records, and it is the single property separating this group from the rest.
g2
Word-level token is "nano" for nano — that is what the table on "Word-level tokenization: the OOV wall" records, and it is the single property separating this group from the rest.

7. Character-level tokenization: the long-sequence wall

Concept

Character-level tokenization has zero OOV — every byte is in vocabulary — but sequences explode in length. Attention is O(n²) in sequence length (Lesson 88), so a 5× longer sequence costs 25× more memory and compute.

TokenizationTokens for 10k-word documentRelative attention cost
Word-level10 0001×
BPE (32k vocab)12 0001.44×
Character-level42 00017.6×

Numbers from a real 10 k-word corpus: character tokenization produces 42 k tokens vs. 10 k at the word level. That 4.2× length difference squares to 17.6× in attention cost — prohibitive for long documents.

8. Fill in: Relative attention cost for Character-level tokenization: the…

Comparison

Comparison matrix

From Character-level tokenization: the long-sequence wall: refill the Relative attention cost column from what you know. The rest of the table is as it appeared.

TokenizationTokens for 10k-word documentRelative attention cost
Word-level10 0001×
BPE (32k vocab)12 0001.44×
Character-level42 00017.6×

9. Subword: best of both worlds

Intuition

Frequent words stay whole (the, is, model). Rare words split at meaningful morpheme boundaries (token + ization, nano + technology). The vocabulary is medium-sized (30k–100k), so sequences stay manageable and embedding tables stay tractable.

Zero OOV
Any unseen word decomposes into known subword pieces (worst case: individual characters).
Short sequences
Common words are single tokens. Sequence length is close to word-level, not character-level.
Morphology signal
un##happy and un##fortunate share the prefix piece — the model can learn prefix semantics.

10. BPE algorithm — build and apply

Section

Part 2 of 4

11. BPE: the three-phase loop

Concept

  1. Initialize: split every word in the training corpus into characters + end-of-word marker </w>. Count word frequencies.
  2. Collect pair stats: scan all adjacent symbol pairs across all word occurrences; multiply by word frequency to get pair frequency.
  3. Merge: combine the most-frequent pair into a single new symbol. Record the merge rule. Repeat from step 2 until the target vocabulary size is reached.

After training, apply the learned merge rules in order to tokenize new text. Earlier merges are applied first; a word never seen in training tokenizes via whatever merges match its characters.

12. What has to happen first: Worked example: BPE on a 4-word corpus

Ranking

Put in order

Put the moves of Worked example: BPE on a 4-word corpus into the order they have to happen.

  1. Initialize: represent each word as its characters + </w>
  2. Compute pair frequencies across all words
  3. Apply merge 1: (e, s) → es; record the rule
  4. Apply all learned merges to tokenize unseen word 'lowest'

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The </w> marker distinguishes 'lower' from 'low' + suffix in another word; it is part of the vocabulary and can be merged.

13. Worked example: BPE on a 4-word corpus

Worked example

Initialize: represent each word as its characters + </w>

Why: The </w> marker distinguishes 'lower' from 'low' + suffix in another word; it is part of the vocabulary and can be merged.

corpus = {"low": 5, "lower": 2, "newest": 6, "widest": 3}

def word_to_chars(word):
    return list(word) + ["</w>"]

vocab = {tuple(word_to_chars(w)): f for w, f in corpus.items()}
for k, v in vocab.items():
    print(" ".join(k), ":", v)
Character sequenceFrequency
l o w </w>5
l o w e r </w>2
n e w e s t </w>6
w i d e s t </w>3

Compute pair frequencies across all words

Why: Each pair (a, b) counts freq(a, b) = sum over all words of freq(word) × occurrences of (a, b) in that word.

Highest-frequency pairs: (e, s) appears in newest (×6) and widest (×3) → freq 9. (l, o) in low and lower → freq 7. (e, s) wins merge 1.

Apply merge 1: (e, s) → es; record the rule

Why: Every occurrence of adjacent e s in the corpus is replaced by the new symbol es. The rule is stored for tokenizing new text later.

StepPair mergedPair freqNew symbol
1(e, s)9es
2(es, t)9est
3(est, </w>)9est</w>
4(l, o)7lo
5(lo, w)7low
6(n, e)6ne

Apply all learned merges to tokenize unseen word 'lowest'

Why: Initialize 'lowest' as [l, o, w, e, s, t, </w>] then replay each merge rule in the order it was learned.

After merge ruleToken sequence for 'lowest'
initl o w e s t </w>
(e, s)l o w es t </w>
(es, t)l o w est </w>
(est, </w>)l o w est</w>
(l, o)lo w est</w>
(lo, w)low est</w>

14. What each one costs: Worked example: BPE on a 4-word corpus

Trade off

Comparison matrix

From Worked example: BPE on a 4-word corpus: every row here is a choice with a cost. Fill the Frequency column, then say which row you would actually pick and what you give up for it.

Character sequenceFrequency
l o w </w>5
l o w e r </w>2
n e w e s t </w>6
w i d e s t </w>3

15. Something is wrong here: applying BPE merges out of order

Anomaly

Predict first

A student writes this, and it looks reasonable:

Tokenizing 'lowest' by applying all known merges simultaneously (greedy longest match): scan the word and grab the longest subword in the vocabulary.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This seems correct here — but only by accident.

Apply merge rules in the order they were learned (rule 1 first, rule 2 second, …). Each rule is a find-and-replace pass over the current token sequence.

Why: This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.

16. Trap: applying BPE merges out of order

Trap

The trap

Tokenizing 'lowest' by applying all known merges simultaneously (greedy longest match): scan the word and grab the longest subword in the vocabulary.

Find 'low' in vocab → emit 'low'; find 'est</w>' in vocab → emit 'est</w>'

Why: This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.

On a word where an early merge creates an intermediate symbol that blocks a later merge, longest-match and ordered-merge diverge. The output is not reproducible BPE.

The fix

Apply merge rules in the order they were learned (rule 1 first, rule 2 second, …). Each rule is a find-and-replace pass over the current token sequence.

Pass 1: replace every adjacent (e, s) → es. Pass 2: replace (es, t) → est. … Pass 5: replace (lo, w) → low.

Why: The ordered-merge procedure is deterministic and matches the training procedure exactly. Skipping or reordering rules changes tokenization.

Libraries like HuggingFace tokenizers encode the merge list as an ordered file. Loading the vocabulary alone (without merge order) is insufficient to reproduce tokenization.

17. Break it on purpose: applying BPE merges out of order

Break the constraint

Discussion prompt

The rule this trap just fixed:

The ordered-merge procedure is deterministic and matches the training procedure exactly. Skipping or reordering rules changes tokenization.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

This seems correct here — but only by accident. Longest-match does not reproduce BPE tokenization in general.

18. WordPiece & SentencePiece — variants

Section

Part 3 of 4

19. WordPiece: likelihood-based merging (BERT)

Concept

WordPiece (used in BERT) is structurally identical to BPE — start with characters, merge iteratively — but chooses which pair to merge differently: instead of raw frequency, it picks the pair that maximally increases the language model likelihood of the training corpus.

\[ \text{score}(a, b) = \frac{\text{freq}(ab)}{\text{freq}(a) \cdot \text{freq}(b)} \]

This is a pointwise mutual information (PMI) proxy: high score means a and b co-occur much more often than chance, so merging them is genuinely informative rather than just frequent.

Candidate pairfreq(ab)freq(a)freq(b)score
(un, happy)414150.01905
(happi, ness)3390.11111

20. WordPiece vs BPE: when the winner changes

Intuition

Imagine the appears 1 million times and th appears 900 k times and e appears 3 million times. BPE merges (t, h) first (raw freq 900k). WordPiece scores (t, h) = 900k / (freq(t) × freq(h)) — if both t and h` are individually very common, the PMI is low and a rarer but more exclusive pair wins instead.

In practice: BPE tends to produce slightly longer merge lists with more common character-level moves early; WordPiece tends to favor morphologically clean boundaries (prefix/suffix). Both converge to similar vocabularies at 30k–50k tokens.

21. SentencePiece: language-agnostic tokenization (T5, LLaMA)

Concept

SentencePiece (Kudo & Richardson 2018) treats the raw Unicode byte stream as input — it does not assume whitespace delimits words. This is critical for Japanese, Chinese, Thai, Arabic, and mixed-language text.

No pre-tokenization
Whitespace is just another character (encoded as ▁). No language-specific pre-processing step.
Reversible
Original text reconstructed exactly from token IDs: ▁hello ▁world → 'hello world'.
BPE or Unigram LM
SentencePiece is a framework; the segmentation algorithm can be BPE or Unigram LM (alternative to WordPiece).

LLaMA uses SentencePiece + BPE at 32k vocab. T5 uses SentencePiece + SentencePiece-Unigram at 32k. GPT-2/3/4 use BPE (byte-level) without SentencePiece at 50k–100k vocab.

22. Which is which: SentencePiece: language-agnostic tokenization…

Matching

Match the pairs

From SentencePiece: language-agnostic tokenization (T5… — match each one to what it actually does. The descriptions have been shuffled.

  • c1. No pre-tokenization
  • c2. Reversible
  • c3. BPE or Unigram LM
  • b1. Whitespace is just another character (encoded as ▁). No language-specific pre-processing step.
  • b2. Original text reconstructed exactly from token IDs: ▁hello ▁world → 'hello world'.
  • b3. SentencePiece is a framework; the segmentation algorithm can be BPE or Unigram LM (alternative to WordPiece).

Why: No pre-tokenization, Reversible, BPE or Unigram LM are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

23. Vocabulary size tradeoffs — the parameter budget

Section

Part 4 of 4

24. Vocab size: the sequence-length / parameter-count tradeoff

Concept

Every token ID maps to an embedding vector of dimension d_model. The embedding table costs |V| × d_model parameters — and this table appears twice (input embedding + output projection are often tied). Doubling vocab doubles embedding cost.

Vocab typeAvg tokens for 10k-word docEmbed params (d=768)
char (~100)42 00076 800
BPE 1k18 000768 000
BPE 32k12 00024 576 000
BPE 100k10 50076 800 000
word (300k+)10 000230 400 000

The sweet spot (GPT-2: 50k, BERT: 30k, LLaMA: 32k) keeps sequence lengths near word-level while keeping the embedding table to 20–40 M parameters — modest relative to the 100 M–70 B parameter range of the rest of the model.

25. Fill in: Avg tokens for 10k-word doc for Vocab size: the sequence-length /…

Comparison

Comparison matrix

From Vocab size: the sequence-length / parameter-count tradeoff: refill the Avg tokens for 10k-word doc column from what you know. The rest of the table is as it appeared.

Vocab typeAvg tokens for 10k-word docEmbed params (d=768)
char (~100)42 00076 800
BPE 1k18 000768 000
BPE 32k12 00024 576 000
BPE 100k10 50076 800 000
word (300k+)10 000230 400 000

26. What has to happen first: Worked example: OOV handling across tokenization…

Ranking

Put in order

Put the moves of Worked example: OOV handling across tokenization schemes into the order they have to happen.

  1. Tokenize 'nanotechnology' under each scheme
  2. Evaluate: word-level loses all signal; char-level preserves signal at 7× the token count; BPE achieves both
  3. Extend to code and math: tokenization schemes handle non-natural text differently

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. A technical term unlikely to appear verbatim in many training corpora stresses every scheme differently.

27. Worked example: OOV handling across tokenization schemes

Worked example

Tokenize 'nanotechnology' under each scheme

Why: A technical term unlikely to appear verbatim in many training corpora stresses every scheme differently.

word = "nanotechnology"

# Character-level
char_toks = list(word)
print("char:", char_toks, "->", len(char_toks), "tokens")

# Word-level (assume OOV)
print("word-level:", ["<UNK>"], "-> 1 token (all signal lost)")

# BPE 32k (simulate: 'nano' and 'technology' both in vocab)
bpe_toks = ["nano", "technology"]
print("BPE 32k:", bpe_toks, "->", len(bpe_toks), "tokens")
SchemeToken sequenceToken countSignal preserved?
char-leveln a n o t e c h n o l o g y14Yes — every char
word-level<UNK>1No — meaning erased
BPE 32knano technology2Yes — morpheme split

Evaluate: word-level loses all signal; char-level preserves signal at 7× the token count; BPE achieves both

Why: The model can learn that 'nano' relates to scale across all nano* compounds. 'technology' is a high-frequency subword shared with 'tech', 'technologies', etc. — useful representations already exist.

Extend to code and math: tokenization schemes handle non-natural text differently

Why: Code identifiers ('get_batch_size'), numbers ('3.14159'), and math symbols ('∇') are all OOV for small word-level vocabs. Byte-level BPE (GPT-4) handles them as byte sequences — guaranteeing zero OOV regardless of domain.

28. Inspect it line by line: Worked example: OOV handling across…

Error analysis

Annotate

Walk the callouts on Worked example: OOV handling across tokenization schemes. Each one is a place this is easy to get subtly wrong.

  • A technical term unlikely to appear verbatim in many training corpora stresses every scheme differently.
  • The model can learn that 'nano' relates to scale across all nano* compounds. 'technology' is a high-frequency subword shared with 'tech', 'technologies', etc. — useful representations already exist.
  • Code identifiers ('get_batch_size'), numbers ('3.14159'), and math symbols ('∇') are all OOV for small word-level vocabs. Byte-level BPE (GPT-4) handles them as byte sequences — guaranteeing zero OOV regardless of domain.

29. Something is wrong here: confusing the merge objective in WordPiece vs. BPE

Anomaly

Predict first

A student writes this, and it looks reasonable:

"WordPiece merges the most frequent adjacent pair — same as BPE but with a different initialization."

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This is the BPE rule, not WordPiece.

WordPiece merges the pair that maximizes likelihood increase — equivalently, the pair with the highest PMI score freq(ab) / (freq(a) × freq(b)).

Why: This is the BPE rule, not WordPiece. Confusing the two is a common exam-level mistake.

30. Trap: confusing the merge objective in WordPiece vs. BPE

Trap

The trap

"WordPiece merges the most frequent adjacent pair — same as BPE but with a different initialization."

Apply: if (e, s) has the highest raw frequency in the corpus, WordPiece merges (e, s) first

Why: This is the BPE rule, not WordPiece. Confusing the two is a common exam-level mistake.

Result: the merge sequence would be identical to BPE, contradicting the existence of WordPiece as a distinct algorithm.

The fix

WordPiece merges the pair that maximizes likelihood increase — equivalently, the pair with the highest PMI score freq(ab) / (freq(a) × freq(b)).

Apply: compare score(e, s) = 9 / (freq(e) × freq(s)) vs. score(happi, ness) = 3 / (3 × 9) = 0.111

Why: If freq(e) = 20 and freq(s) = 15, score(e, s) = 9/300 = 0.030. score(happi, ness) = 0.111 wins — merge (happi, ness) first even though raw freq(e,s)=9 > freq(happi,ness)=3.

The initialization is the same (characters), but the selection criterion differs: BPE = argmax freq(ab); WordPiece = argmax freq(ab) / (freq(a)·freq(b)).

31. Which of these survive contact with Lesson 97: Subword Tokenization — BPE…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Proper nouns, technical terms, code identifiers, and morphological variants are chronically OOV. A 300 k vocab still leaves a long tail of <UNK> tokens at test time.; LLaMA uses SentencePiece + BPE at 32k vocab. T5 uses SentencePiece + SentencePiece-Unigram at 32k. GPT-2/3/4 use BPE (byte-level) without SentencePiece at 50k–100k vocab.
Breaks
Tokenizing 'lowest' by applying all known merges simultaneously (greedy longest match): scan the word and grab the longest subword in the vocabulary.; "WordPiece merges the most frequent adjacent pair — same as BPE but with a different initialization."
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePiece puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

32. Without one step: Pattern: BPE tokenization in 5 steps

Constraint

Discussion prompt

Run Pattern: BPE tokenization in 5 steps with this step confiscated:

Merge: select the pair with the highest frequency (BPE) or highest PMI score (WordPiece). Replace all occurrences. Record the merge rule.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Initialize: for each word in the training corpus, split into characters and append </w>. Store word frequencies.
  2. Pair stats: count frequency of every adjacent symbol pair, weighted by word frequency.
  3. Merge: select the pair with the highest frequency (BPE) or highest PMI score (WordPiece). Replace all occurrences. Record the merge rule.
  4. Repeat steps 2–3 until the target vocabulary size |V| is reached. Larger |V| → shorter sequences, more parameters.
  5. Tokenize new text: apply all recorded merge rules in the order learned, earliest first, to the character sequence of the new word.

33. Pattern: BPE tokenization in 5 steps

Pattern

  1. Initialize: for each word in the training corpus, split into characters and append </w>. Store word frequencies.
  2. Pair stats: count frequency of every adjacent symbol pair, weighted by word frequency.
  3. Merge: select the pair with the highest frequency (BPE) or highest PMI score (WordPiece). Replace all occurrences. Record the merge rule.
  4. Repeat steps 2–3 until the target vocabulary size |V| is reached. Larger |V| → shorter sequences, more parameters.
  5. Tokenize new text: apply all recorded merge rules in the order learned, earliest first, to the character sequence of the new word.

Key distinctions: BPE = frequency criterion; WordPiece = PMI/likelihood criterion; SentencePiece = framework (no whitespace assumption, treats raw bytes, uses BPE or Unigram LM underneath). All three produce subword vocabularies — choice affects boundary placement, not the guarantee of zero OOV.

34. Where does it stop working: Pattern: BPE tokenization in 5 steps

Edge cases

Discussion prompt

Pattern: BPE tokenization in 5 steps works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

Key distinctions: BPE = frequency criterion; WordPiece = PMI/likelihood criterion; SentencePiece = framework (no whitespace assumption, treats raw bytes, uses BPE or Unigram LM underneath). All three produce subword vocabularies — choice affects boundary placement, not the guarantee of zero OOV.

35. Rule out three: Check 1: BPE merge selection

Elimination

Eliminate the wrong options

Pair frequencies: (a, b) = 10+5 = 15; (b, </w>) = 10; (b, c) = 5+8 = 13; (c, </w>) = 5+8 = 13. Which pair does BPE choose?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. (a, b) — frequency 15
  • B. (b, c) — frequency 13
  • C. (c, </w>) — frequency 13
  • D. (b, </w>) — frequency 10

Survives elimination: A

Why: (a, b) appears in both 'ab' (freq 10) and 'abc' (freq 5), giving total frequency 15 — the highest of any pair. BPE always selects the globally most-frequent pair.

36. Check 1: BPE merge selection

Check

Corpus: {"ab": 10, "abc": 5, "bc": 8}. Character sequences after init: a b </w> (10), a b c </w> (5), b c </w> (8). What pair does BPE merge first?

Check your understanding

Pair frequencies: (a, b) = 10+5 = 15; (b, </w>) = 10; (b, c) = 5+8 = 13; (c, </w>) = 5+8 = 13. Which pair does BPE choose?

  • A. (a, b) — frequency 15 (correct)
  • B. (b, c) — frequency 13
  • C. (c, </w>) — frequency 13
  • D. (b, </w>) — frequency 10

Answer: A

Why: (a, b) appears in both 'ab' (freq 10) and 'abc' (freq 5), giving total frequency 15 — the highest of any pair. BPE always selects the globally most-frequent pair.

Why B tempts people
(b, c) has frequency 13 (from 'abc' ×5 and 'bc' ×8), which is less than 15. This would be selected on the second merge, not the first.
Why C tempts people
(c, </w>) also has frequency 13 — tied with (b, c) but still below 15. Ties would be broken by index, but the max is (a, b).
Why D tempts people
(b, </w>) appears only in 'ab' with frequency 10, the lowest of the four candidates.

37. Answer it before you see the options: Check 2: WordPiece vs BPE criterion

Prediction

Predict first

BPE chooses pair X (freq 100 > 20). Which pair does WordPiece choose?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: Z — PMI score for Z = 20/(22×25) = 0.036 > score for X = 0.001

Why: WordPiece uses score = freq(ab) / (freq(a)·freq(b)). score(X) = 100/100000 = 0.001; score(Z) = 20/550 ≈ 0.036. Z has much higher PMI despite lower raw frequency, so WordPiece merges Z first while BPE merges X first — demonstrating the algorithms diverge.

38. Check 2: WordPiece vs BPE criterion

Check

Candidates: pair X has freq(XY)=100, freq(X)=500, freq(Y)=200; pair Z has freq(ZW)=20, freq(Z)=22, freq(W)=25. BPE and WordPiece must each pick one pair to merge.

Check your understanding

BPE chooses pair X (freq 100 > 20). Which pair does WordPiece choose?

  • A. Also X — PMI score for X = 100/(500×200) = 0.001 vs Z = 20/(22×25) = 0.036; X wins
  • B. Z — PMI score for Z = 20/(22×25) = 0.036 > score for X = 0.001 (correct)
  • C. Both algorithms always agree since they start from the same corpus
  • D. Neither — WordPiece requires a full language model pass and cannot decide from counts alone

Answer: B

Why: WordPiece uses score = freq(ab) / (freq(a)·freq(b)). score(X) = 100/100000 = 0.001; score(Z) = 20/550 ≈ 0.036. Z has much higher PMI despite lower raw frequency, so WordPiece merges Z first while BPE merges X first — demonstrating the algorithms diverge.

Why A tempts people
Gets the PMI arithmetic right but draws the wrong conclusion — 0.001 < 0.036, so X loses, not wins.
Why C tempts people
Factually wrong: BPE and WordPiece use different selection criteria and will produce different merge sequences on the same corpus, especially early in training.
Why D tempts people
WordPiece's likelihood criterion reduces to a PMI proxy computable from pair counts alone. It does not require a separately trained language model at merge-selection time.

39. Rule out three: Check 3: vocab size tradeoffs

Elimination

Eliminate the wrong options

Which change reduces embedding parameter count by the most with the smallest increase in sequence length?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Switch from word-level (300k vocab) to BPE 32k
  • B. Switch from BPE 32k to character-level (~100 tokens)
  • C. Switch from BPE 32k to BPE 100k
  • D. Switch from BPE 100k to word-level (300k)

Survives elimination: A

Why: Going from 300k to 32k reduces the embedding table by ~9.4× (300k/32k). Sequence length increases modestly from word-level (10000 tokens) to BPE-32k (12000 tokens) — only 20% longer. Option B trades a 320× embedding reduction for a 3.5× sequence-length increase, worsening attention cost enormously. Option C increases embedding size. Option D increases both.

40. Check 3: vocab size tradeoffs

Check

A team wants to reduce the memory footprint of their transformer's embedding table. They consider three options.

Check your understanding

Which change reduces embedding parameter count by the most with the smallest increase in sequence length?

  • A. Switch from word-level (300k vocab) to BPE 32k (correct)
  • B. Switch from BPE 32k to character-level (~100 tokens)
  • C. Switch from BPE 32k to BPE 100k
  • D. Switch from BPE 100k to word-level (300k)

Answer: A

Why: Going from 300k to 32k reduces the embedding table by ~9.4× (300k/32k). Sequence length increases modestly from word-level (10000 tokens) to BPE-32k (12000 tokens) — only 20% longer. Option B trades a 320× embedding reduction for a 3.5× sequence-length increase, worsening attention cost enormously. Option C increases embedding size. Option D increases both.

Why B tempts people
Character-level shrinks the embedding table maximally (300k → ~100, a 3000× reduction) but blows up sequence length by 4.2×, multiplying attention cost by 17.6×. The net memory impact is worse, not better.
Why C tempts people
BPE 100k has a larger vocabulary than BPE 32k, so this increases the embedding table (76.8M → already larger than 24.6M), the opposite of the goal.
Why D tempts people
Switching from BPE 100k to word-level increases vocabulary size from 100k to 300k, tripling the embedding table. This makes memory worse.

41. Guess the shape of the answer: Worked example: full BPE implementation in…

Estimation

Predict first

After step 1: (e, s) freq=9 → es. After step 2: (es, t) freq=9 → est. After step 3: (est, </w>) freq=9 → est</w>. After step 4: (l, o) freq=7 → lo. After step 5: (lo, w) freq=7 → low. After step 6: (n, e) freq=6 → ne.

Commit before you compute: what does Worked example: full BPE implementation in Python come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Run 6 merge steps and print the trace

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Each step prints the winning pair, its frequency, and the resulting new symbol — this is the merge file that would be saved for later tokenization.

42. Worked example: full BPE implementation in Python

Worked example

Implement get_stats: count adjacent-pair frequencies weighted by word frequency

Why: Each word contributes its frequency to every pair of adjacent symbols it contains. This is O(|corpus|) per merge step.

def get_stats(vocab):
    pairs = {}
    for word, freq in vocab.items():
        symbols = list(word)
        for i in range(len(symbols) - 1):
            pair = (symbols[i], symbols[i+1])
            pairs[pair] = pairs.get(pair, 0) + freq
    return pairs

def merge_vocab(pair, vocab):
    new_vocab = {}
    a, b = pair
    merged = a + b
    for word, freq in vocab.items():
        out, i = [], 0
        wl = list(word)
        while i < len(wl):
            if i < len(wl)-1 and wl[i]==a and wl[i+1]==b:
                out.append(merged); i += 2
            else:
                out.append(wl[i]); i += 1
        new_vocab[tuple(out)] = freq
    return new_vocab
FunctionInputOutput
get_statsvocab dict (tuple→int)pairs dict (pair→int)
merge_vocabpair + vocabnew vocab with pair merged
Overall loopN merge stepslist of N merge rules

Run 6 merge steps and print the trace

Why: Each step prints the winning pair, its frequency, and the resulting new symbol — this is the merge file that would be saved for later tokenization.

After step 1: (e, s) freq=9 → es. After step 2: (es, t) freq=9 → est. After step 3: (est, </w>) freq=9 → est</w>. After step 4: (l, o) freq=7 → lo. After step 5: (lo, w) freq=7 → low. After step 6: (n, e) freq=6 → ne.

43. Where the cost goes: Worked example: full BPE implementation in Python

Cost model

Annotate

In Worked example: full BPE implementation in Python, before reading the notes: mark where the time actually goes. Which line dominates?

  • Each word contributes its frequency to every pair of adjacent symbols it contains. This is O(|corpus|) per merge step.
  • Each step prints the winning pair, its frequency, and the resulting new symbol — this is the merge file that would be saved for later tokenization.

44. Byte-level BPE: handling everything (GPT-2, GPT-4)

Concept

GPT-2 introduced byte-level BPE: the base vocabulary is all 256 possible bytes, not Unicode characters. Every string — code, math, emoji, binary data — decomposes into bytes without any UNK token.

PropertyChar-level BPEByte-level BPE
Base vocab size~130 (printable ASCII + Unicode)256 (all bytes)
OOV possible?Yes (rare Unicode)Never
Handles emoji?SometimesAlways
Handles code?MostlyAlways — byte-precise

The merge procedure is identical to standard BPE; only the initialization changes. GPT-2 uses 50 257-token vocabulary (256 bytes + 50 000 BPE merges + 1 special token). This is why GPT-family models have no OOV problem regardless of input language or domain.

45. What each one costs: Byte-level BPE: handling everything (GPT-2, GPT-4)

Trade off

Comparison matrix

From Byte-level BPE: handling everything (GPT-2, GPT-4): every row here is a choice with a cost. Fill the Char-level BPE column, then say which row you would actually pick and what you give up for it.

PropertyChar-level BPEByte-level BPE
Base vocab size~130 (printable ASCII + Unicode)256 (all bytes)
OOV possible?Yes (rare Unicode)Never
Handles emoji?SometimesAlways
Handles code?MostlyAlways — byte-precise

46. Plan first: Your turn: implement BPE and analyze tokenization

Step zero

Discussion prompt

Your turn: implement BPE and analyze tokenization — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.

Hint: It starts with: Milestone 1 — BPE training: given corpus dict {word: freq}, run N…

Answer:

  1. Milestone 1 — BPE training: given corpus dict {word: freq}, run N merge steps, return list of (pair, merged_token) tuples
  2. Milestone 2 — tokenize: apply the merge list (in order) to a new word, return list of subword tokens
  3. Milestone 3 — sequence length comparison: tokenize the sentence 'subword tokenization handles rare words elegantly' under char-level, your BPE, and a simulated…
  4. Full solution

47. Your turn: implement BPE and analyze tokenization

Worked example

Task: Implement full BPE training + tokenize a test sentence; compare sequence lengths across three schemes.

Milestone 1 — BPE training: given corpus dict {word: freq}, run N merge steps, return list of (pair, merged_token) tuples

Why: The function signature: train_bpe(corpus: dict[str, int], num_merges: int) -> list[tuple]. Use get_stats and merge_vocab from the worked example.

Milestone 2 — tokenize: apply the merge list (in order) to a new word, return list of subword tokens

Why: Function: tokenize(word: str, merges: list[tuple]) -> list[str]. Initialize as list(word) + ['</w>'], then apply each (a, b) merge rule as a scan-and-replace pass.

Milestone 3 — sequence length comparison: tokenize the sentence 'subword tokenization handles rare words elegantly' under char-level, your BPE, and a simulated word-level scheme

Why: Print a table with scheme name and token count. Verify BPE is between the two extremes.

Full solution

Why: Reference implementation — verify yours matches these results before moving to the experiment homework.

# --- train ---
def train_bpe(corpus, num_merges):
    vocab = {tuple(list(w) + ["</w>"]): f for w, f in corpus.items()}
    merges = []
    for _ in range(num_merges):
        pairs = get_stats(vocab)
        if not pairs: break
        best = max(pairs, key=pairs.get)
        vocab = merge_vocab(best, vocab)
        merges.append(best)
    return merges

# --- tokenize ---
def tokenize(word, merges):
    toks = list(word) + ["</w>"]
    for (a, b) in merges:
        out, i = [], 0
        while i < len(toks):
            if i < len(toks)-1 and toks[i]==a and toks[i+1]==b:
                out.append(a+b); i += 2
            else:
                out.append(toks[i]); i += 1
        toks = out
    return toks

corpus = {"low":5,"lower":2,"newest":6,"widest":3}
merges = train_bpe(corpus, 10)
print(tokenize("lowest", merges))   # ['low', 'est</w>']
print(tokenize("newest", merges))   # ['newest</w>']
Input wordBPE tokensToken count
lowest['low', 'est</w>']2
newest['newest</w>']1
widest['wi', 'd', 'est</w>']3
lower['low', 'e', 'r', '</w>']4

48. Fill in: BPE tokens for Your turn: implement BPE and analyze…

Comparison

Comparison matrix

From Your turn: implement BPE and analyze tokenization: refill the BPE tokens column from what you know. The rest of the table is as it appeared.

Input wordBPE tokensToken count
lowest['low', 'est</w>']2
newest['newest</w>']1
widest['wi', 'd', 'est</w>']3
lower['low', 'e', 'r', '</w>']4

49. Connect it up: Lesson 97: Subword Tokenization — BPE, WordPiece, SentencePiece

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why subword? — the two failure modes · BPE algorithm — build and apply · WordPiece & SentencePiece — variants · Vocabulary size tradeoffs — the parameter budget. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

50. Lesson 97 recap: Subword Tokenization

Recap

AlgorithmSelection criterionUsed byOOV?
BPEmax freq(ab)GPT-2, RoBERTa, LLaMANo (byte-level: never)
WordPiecemax freq(ab)/(freq(a)·freq(b))BERT, DistilBERTNo
SentencePiece+BPEmax freq(ab), raw bytesLLaMA, PaLMNever
SentencePiece+Unigrammax log-likelihood dropT5, mT5Never

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 97 — Subword Tokenization — Barron · USAAIO Round 2 Preparation, 2026
  2. Sennrich et al. 'Neural Machine Translation of Rare Words with Subword Units' (ACL 2016) — arXiv:1508.07909 — original BPE paper
  3. Schuster & Nakamura 'Japanese and Korean Voice Search' (ICASSP 2012) — IEEE — WordPiece algorithm
  4. Kudo & Richardson 'SentencePiece: A simple and language independent subword tokenizer' (EMNLP 2018) — arXiv:1808.06226
  5. BPE from-scratch implementation verified with Python 3 — all merge traces and seq-length numbers are real execution output, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108