USAAIO Lesson 105, from Phase 3. It covers BLEU, built on n-gram precision with a brevity penalty; ROUGE, which is recall-oriented n-gram overlap, along with ROUGE-L via the longest common subsequence; and BERTScore, which uses the cosine similarity of contextual embeddings. It also covers where each metric fails. All the computations were implemented from scratch in Python and verified by real execution, giving p_1 = 5/6, p_2 = 3/5, and BLEU-4 = 0.4209 on the machine-translation example, and a ROUGE-L F1 of 0.8333 on the LCS example. The lesson runs to 25 slides.
Subject: Machine Learning · 48 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 105 · Phase 3
Three generations of NLG evaluation: n-gram precision (BLEU), n-gram recall (ROUGE), and contextual embedding similarity (BERTScore). Every formula derived, every number executed.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 105: BLEU, ROUGE, and BERTScore: without looking back, what was the main idea of Text Decoding Strategies, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
greedy decoding, beam search with width-k expansion, length normalization, temperature scaling, top-k masking, and top-p (nucleus) sampling — all verified with real PyTorch log-softmax arithmetic. Greedy suboptimality proved by counter-example.
Section
Part 1 of 4
Concept
BLEU (Papineni 2002) measures how many n-grams in the hypothesis also appear in at least one reference. It is precision-oriented: it asks 'of what the model produced, how much was also in the reference?'
\[ p_n = \frac{\sum_{\text{hyp}} \min(\text{count}_{\text{hyp}}(g_n),\, \max_r \text{count}_{r}(g_n))}{\sum_{\text{hyp}} \text{count}_{\text{hyp}}(g_n)} \]
The min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.
\[ \text{BLEU-4} = BP \cdot \exp\!\left(\frac{1}{4}\sum_{n=1}^{4} \log p_n\right) \]
Counterexample
Discussion prompt
The min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.
\[ BP = \begin{cases} 1 & \text{if } |\text{hyp}| \geq |\text{ref}| \\ e^{1 - |\text{ref}|/|\text{hyp}|} & \text{otherwise} \end{cases} \]
Example: hyp = 'the cat' (2 tokens), ref = 'the cat sat on the mat' (6 tokens). BP = exp(1 - 6/2) = exp(-2) = 0.1353. Even with p_1 = p_2 = 1.0, BLEU is slashed.
Analogy
Discussion prompt
Explain Brevity Penalty: punishing short outputs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.
Ranking
Put in order
Put the moves of Worked Example: BLEU-4 from scratch into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Clip each hyp n-gram count at the max it appears in the reference, sum clipped counts over denominator (raw hyp n-gram count).
Worked example
Hyp: 'it is a guide to action which ensures that the military always obeys the commands of the party' (18 tokens). Ref: 'it is a guide to action that ensures that the military will forever heed party commands' (17 tokens).
Compute BP
Why: hyp_len=18, ref_len=17. Since 18 > 17, BP = 1.0 — no penalty.
Compute modified n-gram precisions p_1 through p_4
Why: Clip each hyp n-gram count at the max it appears in the reference, sum clipped counts over denominator (raw hyp n-gram count).
from collections import Counter
import math
def get_ngrams(tokens, n):
return [tuple(tokens[i:i+n]) for i in range(len(tokens)-n+1)]
def modified_precision(hyp, refs, n):
hyp_ng = Counter(get_ngrams(hyp, n))
max_ref = Counter()
for ref in refs:
ref_ng = Counter(get_ngrams(ref, n))
for ng in hyp_ng:
max_ref[ng] = max(max_ref[ng], ref_ng[ng])
clipped = sum(min(c, max_ref[ng]) for ng, c in hyp_ng.items())
return clipped, sum(hyp_ng.values())
hyp = "it is a guide to action which ensures that the military always obeys the commands of the party".split()
ref = "it is a guide to action that ensures that the military will forever heed party commands".split()
for n in range(1, 5):
cn, tn = modified_precision(hyp, [ref], n)
print(f'p_{n}: {cn}/{tn} = {cn/tn:.4f}')
bp = 1.0 # hyp_len 18 >= ref_len 17
log_sum = sum(math.log(cn/tn) for n in range(1,5)
for cn, tn in [modified_precision(hyp, [ref], n)])
bleu = bp * math.exp(log_sum / 4)
print(f'BLEU-4 = {bleu:.4f}')| n | clipped | total | p_n |
|---|---|---|---|
| 1 | 12 | 18 | 0.6667 |
| 2 | 8 | 17 | 0.4706 |
| 3 | 6 | 16 | 0.3750 |
| 4 | 4 | 15 | 0.2667 |
Combine: BLEU-4 = BP × exp((log 0.6667 + log 0.4706 + log 0.3750 + log 0.2667) / 4)
Why: = 1.0 × exp((-0.4055 - 0.7537 - 0.9808 - 1.3218) / 4) = exp(-0.8655) = 0.4209
Verify: BLEU-4 = 0.4209
Why: All four p_n > 0 and BP = 1.0, so the geometric mean is well-defined. A score of 0.42 is considered reasonable for a single-reference MT evaluation.
Comparison
Comparison matrix
From Worked Example: BLEU-4 from scratch: refill the clipped column from what you know. The rest of the table is as it appeared.
| n | clipped | total | p_n |
|---|---|---|---|
| 1 | 12 | 18 | 0.6667 |
| 2 | 8 | 17 | 0.4706 |
| 3 | 6 | 16 | 0.3750 |
| 4 | 4 | 15 | 0.2667 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student computes p_1=0.83, p_2=0.60, p_3=0.25, p_4=0.0 and writes BLEU-4 = 0.83 × 0.60 × 0.25 × 0.0 then claims that 'BLEU-4 is very low but nonzero.'
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.
If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.
Why: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.
Trap
A student computes p_1=0.83, p_2=0.60, p_3=0.25, p_4=0.0 and writes BLEU-4 = 0.83 × 0.60 × 0.25 × 0.0 then claims that 'BLEU-4 is very low but nonzero.'
Skip the p_4 term or average only p_1 through p_3
Why: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.
If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.
Return BLEU-4 = 0.0 whenever p_n = 0 for any n ≤ 4
Why: The geometric mean collapses: exp((-∞ + finite + finite + finite)/4) = exp(-∞) = 0. Example: hyp='the cat sat on the mat', ref='the cat is on the mat' gives p_4=0/3=0, so BLEU-4=0.0 despite p_1=5/6, p_2=3/5, p_3=1/4.
Break the constraint
Discussion prompt
The rule this trap just fixed:
If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.
Section
Part 2 of 4
Concept
ROUGE (Lin 2004) asks the opposite question from BLEU: 'of the n-grams in the reference, how many did the model produce?' This makes ROUGE suited to summarization, where coverage matters more than precision.
\[ \text{ROUGE-}n = \frac{\sum_{g_n \in \text{ref}} \min(\text{count}_{\text{hyp}}(g_n),\, \text{count}_{\text{ref}}(g_n))}{\sum_{g_n \in \text{ref}} \text{count}_{\text{ref}}(g_n)} \]
Matching
Match the pairs
From ROUGE: recall over reference n-grams — match each one to what it actually does. The descriptions have been shuffled.
Why: ROUGE-1, ROUGE-2, ROUGE-L are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.
Concept
ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.
\[ \text{ROUGE-L} = \frac{2 \cdot P_{\text{lcs}} \cdot R_{\text{lcs}}}{P_{\text{lcs}} + R_{\text{lcs}}} \quad P_{\text{lcs}} = \frac{|\text{LCS}|}{|\text{hyp}|} \quad R_{\text{lcs}} = \frac{|\text{LCS}|}{|\text{ref}|} \]
LCS is computed with a standard O(mn) DP table — same algorithm as Edit Distance. A subsequence like ['the','cat','on','the','mat'] skips 'is'/'sat' without breaking.
Analogy
Discussion prompt
Explain ROUGE-L: LCS instead of contiguous n-grams by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.
Ranking
Put in order
Put the moves of Worked Example: ROUGE-L via LCS DP into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. dp[i][j] = LCS length of hyp[:i] and ref[:j].
Worked example
Hyp: 'the cat sat on the mat' (6 tokens). Ref: 'the cat is on the mat' (6 tokens). Goal: compute ROUGE-L P, R, F1.
Build the LCS DP table
Why: dp[i][j] = LCS length of hyp[:i] and ref[:j]. Increment if tokens match, otherwise take max of neighbors.
def lcs_length(a, b):
m, n = len(a), len(b)
dp = [[0]*(n+1) for _ in range(m+1)]
for i in range(1, m+1):
for j in range(1, n+1):
if a[i-1] == b[j-1]:
dp[i][j] = dp[i-1][j-1] + 1
else:
dp[i][j] = max(dp[i-1][j], dp[i][j-1])
return dp[m][n]
hyp = 'the cat sat on the mat'.split()
ref = 'the cat is on the mat'.split()
lcs = lcs_length(hyp, ref)
p = lcs / len(hyp)
r = lcs / len(ref)
f1 = 2*p*r / (p+r)
print(f'LCS length: {lcs}')
print(f'P={p:.4f} R={r:.4f} F1={f1:.4f}')| dp row (hyp\ref) | the | cat | is | on | the | mat |
|---|---|---|---|---|---|---|
| the | 1 | 1 | 1 | 1 | 1 | 1 |
| cat | 1 | 2 | 2 | 2 | 2 | 2 |
| sat | 1 | 2 | 2 | 2 | 2 | 2 |
| on | 1 | 2 | 2 | 3 | 3 | 3 |
| the | 1 | 2 | 2 | 3 | 4 | 4 |
| mat | 1 | 2 | 2 | 3 | 4 | 5 |
Read LCS = dp[6][6] = 5; the matched subsequence is ['the','cat','on','the','mat']
Why: 'sat' in hyp and 'is' in ref have no counterpart — they are skipped. Everything else aligns in order.
ROUGE-L: P = 5/6 = 0.8333, R = 5/6 = 0.8333, F1 = 2×0.8333×0.8333/(0.8333+0.8333) = 0.8333
Why: Both strings are length 6 and share LCS=5, so P=R=F1. Compare: BLEU-4 was 0.0 on this same pair due to p_4=0/3.
Error analysis
Annotate
Walk the callouts on Worked Example: ROUGE-L via LCS DP. Each one is a place this is easy to get subtly wrong.
Section
Part 3 of 4
Concept
BERTScore (Zhang 2020) encodes hyp and ref through a pretrained BERT model and computes token-level cosine similarity across both sequences, then aggregates with a greedy matching strategy.
\[ P_{\text{BERT}} = \frac{1}{|\hat{x}|} \sum_{\hat{x}_i \in \hat{x}} \max_{x_j \in x} \cos(\mathbf{e}_{\hat{x}_i},\, \mathbf{e}_{x_j}) \]
\[ R_{\text{BERT}} = \frac{1}{|x|} \sum_{x_j \in x} \max_{\hat{x}_i \in \hat{x}} \cos(\mathbf{e}_{x_j},\, \mathbf{e}_{\hat{x}_i}) \]
F1_BERT = harmonic mean of P_BERT and R_BERT. Because contextual embeddings place 'feline' and 'cat' nearby in embedding space, BERTScore correctly assigns high similarity where BLEU-4 and ROUGE-1 score near 0.
Counterexample
Discussion prompt
BERTScore (Zhang 2020) encodes hyp and ref through a pretrained BERT model and computes token-level cosine similarity across both sequences, then aggregates with a greedy matching strategy.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
BLEU and ROUGE treat token matching as binary: 'cat' ≠ 'feline' even though they are synonyms. BERTScore uses continuous similarity — any two tokens get a score in [−1, 1] based on their contextual embedding dot product.
| Pair | BLEU-4 ≈ | ROUGE-1 ≈ | BERTScore F1 ≈ |
|---|---|---|---|
| 'the cat sat on the mat' vs ref | 0.00 | 0.83 | 0.95+ |
| 'a feline rested on the rug' vs ref | 0.00 | 0.33 | 0.87+ |
| Perfect match | 1.00 | 1.00 | 1.00 |
BERTScore requires pretrained model weights (typical: bert-base-uncased, roberta-large) and GPU inference — unlike BLEU/ROUGE which are O(N) string ops. This compute cost is the main reason BLEU remains the default MT benchmark.
Comparison
Comparison matrix
From Why BERTScore correlates better with humans: refill the BERTScore F1 ≈ column from what you know. The rest of the table is as it appeared.
| Pair | BLEU-4 ≈ | ROUGE-1 ≈ | BERTScore F1 ≈ |
|---|---|---|---|
| 'the cat sat on the mat' vs ref | 0.00 | 0.83 | 0.95+ |
| 'a feline rested on the rug' vs ref | 0.00 | 0.33 | 0.87+ |
| Perfect match | 1.00 | 1.00 | 1.00 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
A student evaluates two summaries solely by ROUGE-1 recall and concludes the higher-scoring one is the better summary.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Reasoning: 'higher recall = better coverage = better summary.'
ROUGE-1 counts unigrams regardless of order — a bag of words. A summary with high ROUGE-1 can be incoherent or miss the key claim if it happens to share many stopwords with the reference.
Why: Reasoning: 'higher recall = better coverage = better summary.'
Trap
A student evaluates two summaries solely by ROUGE-1 recall and concludes the higher-scoring one is the better summary.
Select the summary with ROUGE-1 = 0.78 over the one with ROUGE-1 = 0.61
Why: Reasoning: 'higher recall = better coverage = better summary.'
ROUGE-1 counts unigrams regardless of order — a bag of words. A summary with high ROUGE-1 can be incoherent or miss the key claim if it happens to share many stopwords with the reference.
Report ROUGE-1, ROUGE-2, and ROUGE-L together; consider ROUGE-2 as the primary signal
Why: ROUGE-2 bigrams are harder to match without understanding — they filter out stopword-padding. ROUGE-L further penalizes disordered output. For higher-stakes evaluation (e.g. USAAIO), add BERTScore to catch meaning-preserving paraphrases.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.; A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.; ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.Section
Part 4 of 4
Concept
| Task | Primary metric | Why |
|---|---|---|
| Machine translation | BLEU-4 | Reference translations are dense; precision matters; fast to compute; community convention |
| Abstractive summarization | ROUGE-1/2/L | Recall of key content matters; multiple valid wordings OK |
| Open-ended generation (dialogue, creative) | BERTScore (or human) | No fixed reference; meaning > surface form |
| Code generation | Exact match / pass@k | Syntax must be exact; n-gram overlap is meaningless |
All three metrics assume at least one reference output exists. For tasks where 'there is no single correct answer' (story generation, chatbots), automatic metrics correlate poorly with human judgment even with BERTScore.
Trade off
Comparison matrix
From Task-metric alignment: every row here is a choice with a cost. Fill the Primary metric column, then say which row you would actually pick and what you give up for it.
| Task | Primary metric | Why |
|---|---|---|
| Machine translation | BLEU-4 | Reference translations are dense; precision matters; fast to compute; community convention |
| Abstractive summarization | ROUGE-1/2/L | Recall of key content matters; multiple valid wordings OK |
| Open-ended generation (dialogue, creative) | BERTScore (or human) | No fixed reference; meaning > surface form |
| Code generation | Exact match / pass@k | Syntax must be exact; n-gram overlap is meaningless |
Constraint
Discussion prompt
Run Pattern: computing and interpreting NLG metrics with this step confiscated:
ROUGE-L: run O(mn) LCS DP → |LCS|. P = |LCS|/|hyp|, R = |LCS|/|ref|, F1 = harmonic mean.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
Pattern: computing and interpreting NLG metrics works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Hypothesis: 'dogs run fast' (3 tokens). Reference: 'the dogs run very fast' (5 tokens). All three hypothesis unigrams appear in the reference. What is the brevity penalty BP?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: BP = exp(1 − |ref|/|hyp|) when |hyp| < |ref|. Here |hyp|=3, |ref|=5, so BP = exp(1 − 5/3) = exp(−2/3) ≈ 0.5134. The penalty is applied because the hypothesis is shorter than the reference.
Check
Work through the arithmetic before clicking.
Check your understanding
Hypothesis: 'dogs run fast' (3 tokens). Reference: 'the dogs run very fast' (5 tokens). All three hypothesis unigrams appear in the reference. What is the brevity penalty BP?
Answer: A
Why: BP = exp(1 − |ref|/|hyp|) when |hyp| < |ref|. Here |hyp|=3, |ref|=5, so BP = exp(1 − 5/3) = exp(−2/3) ≈ 0.5134. The penalty is applied because the hypothesis is shorter than the reference.
Prediction
Predict first
Reference: 'the cat sat on the mat' (6 unigrams). Hypothesis: 'the cat' (2 unigrams, both in reference). What is ROUGE-1 recall?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: 2/6 ≈ 0.333
Why: ROUGE-1 recall = (matched ref unigrams) / (total ref unigrams) = 2 / 6 ≈ 0.333. Both 'the' and 'cat' appear in the hypothesis, but the reference has 6 total unigrams, so recall is 2/6.
Check
Recall the denominator difference between BLEU and ROUGE.
Check your understanding
Reference: 'the cat sat on the mat' (6 unigrams). Hypothesis: 'the cat' (2 unigrams, both in reference). What is ROUGE-1 recall?
Answer: A
Why: ROUGE-1 recall = (matched ref unigrams) / (total ref unigrams) = 2 / 6 ≈ 0.333. Both 'the' and 'cat' appear in the hypothesis, but the reference has 6 total unigrams, so recall is 2/6.
Elimination
Eliminate the wrong options
Hypothesis: 'a feline rested on the rug'. Reference: 'the cat sat on the mat'. BLEU-4 = 0 and ROUGE-1 recall ≈ 0.33. Which statement best explains why BERTScore F1 would be substantially higher?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: B
Why: BERTScore's core mechanism is pairwise cosine similarity between BERT contextual embeddings. 'Feline' and 'cat' are contextually similar tokens (high cosine similarity), so the greedy-max over the similarity matrix assigns each hypothesis token a high score even without a surface-level string match. This is the exact limitation that motivates BERTScore over BLEU/ROUGE.
Check
Think about what BERTScore's embedding similarity captures that BLEU/ROUGE cannot.
Check your understanding
Hypothesis: 'a feline rested on the rug'. Reference: 'the cat sat on the mat'. BLEU-4 = 0 and ROUGE-1 recall ≈ 0.33. Which statement best explains why BERTScore F1 would be substantially higher?
Answer: B
Why: BERTScore's core mechanism is pairwise cosine similarity between BERT contextual embeddings. 'Feline' and 'cat' are contextually similar tokens (high cosine similarity), so the greedy-max over the similarity matrix assigns each hypothesis token a high score even without a surface-level string match. This is the exact limitation that motivates BERTScore over BLEU/ROUGE.
Concept
Project brief: build bleu_rouge.py that (1) implements BLEU-4 with modified n-gram precision and brevity penalty, (2) implements ROUGE-L via LCS DP, and (3) shows a failure-mode example where BLEU=0 but ROUGE-L > 0.
count_ngrams(tokens, n) returning a Counter of n-gram tuples.modified_precision(hyp, refs, n) returning (clipped, total). Test: hyp='it is a guide to action which ensures that the military always obeys the commands of the party', ref as in the worked example → p_1=12/18=0.6667.brevity_penalty(hyp_len, ref_len) and bleu4(hyp, refs). Test: same sentence pair → BLEU-4=0.4209.lcs_length(a, b) with O(mn) DP and rouge_l(hyp, ref) returning (P, R, F1). Test: hyp='the cat sat on the mat', ref='the cat is on the mat' → LCS=5, F1=0.8333.Worked example
from collections import Counter
import math
# ── N-gram utilities ─────────────────────────────────────────
def count_ngrams(tokens, n):
return Counter(tuple(tokens[i:i+n])
for i in range(len(tokens)-n+1))
def modified_precision(hyp, refs, n):
hyp_ng = count_ngrams(hyp, n)
if not hyp_ng:
return 0, 0
max_ref = Counter()
for ref in refs:
ref_ng = count_ngrams(ref, n)
for ng in hyp_ng:
max_ref[ng] = max(max_ref[ng], ref_ng[ng])
clipped = sum(min(c, max_ref[ng]) for ng, c in hyp_ng.items())
return clipped, sum(hyp_ng.values())
# ── BLEU-4 ───────────────────────────────────────────────────
def brevity_penalty(hyp_len, ref_len):
return 1.0 if hyp_len >= ref_len else math.exp(1 - ref_len/hyp_len)
def bleu4(hyp_str, ref_strs):
hyp = hyp_str.split()
refs = [r.split() for r in ref_strs]
ref_len = min((len(r) for r in refs),
key=lambda x: abs(x - len(hyp)))
bp = brevity_penalty(len(hyp), ref_len)
log_sum = 0.0
for n in range(1, 5):
cn, tn = modified_precision(hyp, refs, n)
if tn == 0 or cn == 0:
return 0.0
log_sum += math.log(cn / tn)
return bp * math.exp(log_sum / 4)
# ── ROUGE-L ──────────────────────────────────────────────────
def lcs_length(a, b):
m, n = len(a), len(b)
dp = [[0]*(n+1) for _ in range(m+1)]
for i in range(1, m+1):
for j in range(1, n+1):
if a[i-1] == b[j-1]:
dp[i][j] = dp[i-1][j-1] + 1
else:
dp[i][j] = max(dp[i-1][j], dp[i][j-1])
return dp[m][n]
def rouge_l(hyp_str, ref_str):
h, r = hyp_str.split(), ref_str.split()
lcs = lcs_length(h, r)
p = lcs / len(h) if h else 0.0
rc = lcs / len(r) if r else 0.0
f1 = 2*p*rc/(p+rc) if (p+rc) else 0.0
return p, rc, f1
# ── Demo ─────────────────────────────────────────────────────
if __name__ == '__main__':
MT_HYP = ('it is a guide to action which ensures that the '
'military always obeys the commands of the party')
MT_REF = ('it is a guide to action that ensures that the '
'military will forever heed party commands')
b = bleu4(MT_HYP, [MT_REF])
print(f'BLEU-4 (MT pair): {b:.4f}') # 0.4209
H = 'the cat sat on the mat'
R = 'the cat is on the mat'
p, rc, f1 = rouge_l(H, R)
print(f'ROUGE-L F1: {f1:.4f}') # 0.8333
print(f'BLEU-4 same pair: {bleu4(H,[R]):.4f}') # 0.0 (p_4=0)| Call | Expected output | Actual output |
|---|---|---|
| bleu4(MT_HYP, [MT_REF]) | 0.4209 | 0.4209 |
| rouge_l('the cat sat on the mat', 'the cat is on the mat') F1 | 0.8333 | 0.8333 |
| bleu4('the cat sat on the mat', ['the cat is on the mat']) | 0.0000 | 0.0000 |
The third row is the failure-mode: BLEU-4 = 0 because p_4 = 0/3 = 0 (no shared 4-gram). ROUGE-L = 0.8333 because LCS = 5 tokens still capture the shared structure. This motivates reporting both metrics.
Comparison
Comparison matrix
From Full program: bleu_rouge.py: refill the Expected output column from what you know. The rest of the table is as it appeared.
| Call | Expected output | Actual output |
|---|---|---|
| bleu4(MT_HYP, [MT_REF]) | 0.4209 | 0.4209 |
| rouge_l('the cat sat on the mat', 'the cat is on the mat') F1 | 0.8333 | 0.8333 |
| bleu4('the cat sat on the mat', ['the cat is on the mat']) | 0.0000 | 0.0000 |
Concept
'the the the the the the', ref = 'the cat sat on the mat'. What does modified_precision return for n=1? Verify the clip prevents inflation.'a feline rested on the rug', ref = same as above. Show BLEU-4=0, ROUGE-1≈0.33, ROUGE-L F1≈0.33 — and explain why BERTScore would score this higher.bleu4. Confirm it picks the best-matching reference length for BP and takes the per-reference max counts for clipping.bleu4('', [ref]) returns 0.0 without crashing (BP → 0, and n-gram counts are empty).Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — BLEU — Bilingual Evaluation Understudy · ROUGE — Recall-Oriented Understudy · BERTScore — Contextual Embedding Similarity · Metric Selection and Failure Modes. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.