Lesson 105: BLEU, ROUGE, and BERTScore

USAAIO Lesson 105, from Phase 3. It covers BLEU, built on n-gram precision with a brevity penalty; ROUGE, which is recall-oriented n-gram overlap, along with ROUGE-L via the longest common subsequence; and BERTScore, which uses the cosine similarity of contextual embeddings. It also covers where each metric fails. All the computations were implemented from scratch in Python and verified by real execution, giving p_1 = 5/6, p_2 = 3/5, and BLEU-4 = 0.4209 on the machine-translation example, and a ROUGE-L F1 of 0.8333 on the LCS example. The lesson runs to 25 slides.

Subject: Machine Learning · 48 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. BLEU · ROUGE & BERTScore

Title

USAAIO · Lesson 105 · Phase 3

Three generations of NLG evaluation: n-gram precision (BLEU), n-gram recall (ROUGE), and contextual embedding similarity (BERTScore). Every formula derived, every number executed.

2. By the end of this lesson you can

Objectives

  1. Compute modified n-gram precision and apply the brevity penalty to produce a BLEU-4 score by hand
  2. Distinguish BLEU's precision orientation from ROUGE's recall orientation and state why each suits different tasks
  3. Implement ROUGE-L using the LCS dynamic-programming table and derive P, R, and F1
  4. Explain BERTScore's token-level cosine similarity aggregation and why it captures synonymy that BLEU/ROUGE cannot
  5. Identify failure modes of all three metrics and choose the right metric for translation, summarization, and open-ended generation

3. What survived from Text Decoding Strategies?

Warm-up

Discussion prompt

Before we open Lesson 105: BLEU, ROUGE, and BERTScore: without looking back, what was the main idea of Text Decoding Strategies, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

greedy decoding, beam search with width-k expansion, length normalization, temperature scaling, top-k masking, and top-p (nucleus) sampling — all verified with real PyTorch log-softmax arithmetic. Greedy suboptimality proved by counter-example.

4. BLEU — Bilingual Evaluation Understudy

Section

Part 1 of 4

5. BLEU: precision of n-gram overlap

Concept

BLEU (Papineni 2002) measures how many n-grams in the hypothesis also appear in at least one reference. It is precision-oriented: it asks 'of what the model produced, how much was also in the reference?'

\[ p_n = \frac{\sum_{\text{hyp}} \min(\text{count}_{\text{hyp}}(g_n),\, \max_r \text{count}_{r}(g_n))}{\sum_{\text{hyp}} \text{count}_{\text{hyp}}(g_n)} \]

The min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.

\[ \text{BLEU-4} = BP \cdot \exp\!\left(\frac{1}{4}\sum_{n=1}^{4} \log p_n\right) \]

6. Break it if you can: BLEU: precision of n-gram overlap

Counterexample

Discussion prompt

The min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

7. Brevity Penalty: punishing short outputs

Concept

A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.

\[ BP = \begin{cases} 1 & \text{if } |\text{hyp}| \geq |\text{ref}| \\ e^{1 - |\text{ref}|/|\text{hyp}|} & \text{otherwise} \end{cases} \]

Example: hyp = 'the cat' (2 tokens), ref = 'the cat sat on the mat' (6 tokens). BP = exp(1 - 6/2) = exp(-2) = 0.1353. Even with p_1 = p_2 = 1.0, BLEU is slashed.

8. By analogy: Brevity Penalty: punishing short outputs

Analogy

Discussion prompt

Explain Brevity Penalty: punishing short outputs by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.

9. What has to happen first: Worked Example: BLEU-4 from scratch

Ranking

Put in order

Put the moves of Worked Example: BLEU-4 from scratch into the order they have to happen.

  1. Compute modified n-gram precisions p_1 through p_4
  2. Combine: BLEU-4 = BP × exp((log 0.6667 + log 0.4706 + log 0.3750 + log 0.2667) / 4)
  3. Verify: BLEU-4 = 0.4209

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Clip each hyp n-gram count at the max it appears in the reference, sum clipped counts over denominator (raw hyp n-gram count).

10. Worked Example: BLEU-4 from scratch

Worked example

Hyp: 'it is a guide to action which ensures that the military always obeys the commands of the party' (18 tokens). Ref: 'it is a guide to action that ensures that the military will forever heed party commands' (17 tokens).

Compute BP

Why: hyp_len=18, ref_len=17. Since 18 > 17, BP = 1.0 — no penalty.

Compute modified n-gram precisions p_1 through p_4

Why: Clip each hyp n-gram count at the max it appears in the reference, sum clipped counts over denominator (raw hyp n-gram count).

from collections import Counter
import math

def get_ngrams(tokens, n):
    return [tuple(tokens[i:i+n]) for i in range(len(tokens)-n+1)]

def modified_precision(hyp, refs, n):
    hyp_ng = Counter(get_ngrams(hyp, n))
    max_ref = Counter()
    for ref in refs:
        ref_ng = Counter(get_ngrams(ref, n))
        for ng in hyp_ng:
            max_ref[ng] = max(max_ref[ng], ref_ng[ng])
    clipped = sum(min(c, max_ref[ng]) for ng, c in hyp_ng.items())
    return clipped, sum(hyp_ng.values())

hyp = "it is a guide to action which ensures that the military always obeys the commands of the party".split()
ref = "it is a guide to action that ensures that the military will forever heed party commands".split()

for n in range(1, 5):
    cn, tn = modified_precision(hyp, [ref], n)
    print(f'p_{n}: {cn}/{tn} = {cn/tn:.4f}')

bp = 1.0  # hyp_len 18 >= ref_len 17
log_sum = sum(math.log(cn/tn) for n in range(1,5)
              for cn, tn in [modified_precision(hyp, [ref], n)])
bleu = bp * math.exp(log_sum / 4)
print(f'BLEU-4 = {bleu:.4f}')
nclippedtotalp_n
112180.6667
28170.4706
36160.3750
44150.2667

Combine: BLEU-4 = BP × exp((log 0.6667 + log 0.4706 + log 0.3750 + log 0.2667) / 4)

Why: = 1.0 × exp((-0.4055 - 0.7537 - 0.9808 - 1.3218) / 4) = exp(-0.8655) = 0.4209

Verify: BLEU-4 = 0.4209

Why: All four p_n > 0 and BP = 1.0, so the geometric mean is well-defined. A score of 0.42 is considered reasonable for a single-reference MT evaluation.

11. Fill in: clipped for Worked Example: BLEU-4 from scratch

Comparison

Comparison matrix

From Worked Example: BLEU-4 from scratch: refill the clipped column from what you know. The rest of the table is as it appeared.

nclippedtotalp_n
112180.6667
28170.4706
36160.3750
44150.2667

12. Something is wrong here: treating any p_n = 0 as a recoverable edge case

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student computes p_1=0.83, p_2=0.60, p_3=0.25, p_4=0.0 and writes BLEU-4 = 0.83 × 0.60 × 0.25 × 0.0 then claims that 'BLEU-4 is very low but nonzero.'

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.

If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.

Why: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.

13. Trap: treating any p_n = 0 as a recoverable edge case

Trap

The trap

A student computes p_1=0.83, p_2=0.60, p_3=0.25, p_4=0.0 and writes BLEU-4 = 0.83 × 0.60 × 0.25 × 0.0 then claims that 'BLEU-4 is very low but nonzero.'

Skip the p_4 term or average only p_1 through p_3

Why: Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.

The fix

If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.

Return BLEU-4 = 0.0 whenever p_n = 0 for any n ≤ 4

Why: The geometric mean collapses: exp((-∞ + finite + finite + finite)/4) = exp(-∞) = 0. Example: hyp='the cat sat on the mat', ref='the cat is on the mat' gives p_4=0/3=0, so BLEU-4=0.0 despite p_1=5/6, p_2=3/5, p_3=1/4.

14. Break it on purpose: treating any p_n = 0 as a recoverable edge…

Break the constraint

Discussion prompt

The rule this trap just fixed:

If any p_n = 0, then log(p_n) = -∞ and BLEU-4 = 0.0 exactly — no exceptions.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Reasoning: 'zero precision for 4-grams is a minor issue; average the rest.' This misunderstands the geometric-mean formula.

15. ROUGE — Recall-Oriented Understudy

Section

Part 2 of 4

16. ROUGE: recall over reference n-grams

Concept

ROUGE (Lin 2004) asks the opposite question from BLEU: 'of the n-grams in the reference, how many did the model produce?' This makes ROUGE suited to summarization, where coverage matters more than precision.

\[ \text{ROUGE-}n = \frac{\sum_{g_n \in \text{ref}} \min(\text{count}_{\text{hyp}}(g_n),\, \text{count}_{\text{ref}}(g_n))}{\sum_{g_n \in \text{ref}} \text{count}_{\text{ref}}(g_n)} \]

ROUGE-1
Unigram recall. Measures word coverage; loose but correlates with fluency.
ROUGE-2
Bigram recall. Captures phrase overlap; stronger signal for content fidelity.
ROUGE-L
Longest Common Subsequence recall. Respects word order without requiring contiguous matches.

17. Which is which: ROUGE: recall over reference n-grams

Matching

Match the pairs

From ROUGE: recall over reference n-grams — match each one to what it actually does. The descriptions have been shuffled.

  • c1. ROUGE-1
  • c2. ROUGE-2
  • c3. ROUGE-L
  • b1. Unigram recall. Measures word coverage; loose but correlates with fluency.
  • b2. Bigram recall. Captures phrase overlap; stronger signal for content fidelity.
  • b3. Longest Common Subsequence recall. Respects word order without requiring contiguous matches.

Why: ROUGE-1, ROUGE-2, ROUGE-L are easy to tell apart while they are sitting next to their descriptions and much harder afterwards, which is what this checks.

18. ROUGE-L: LCS instead of contiguous n-grams

Concept

ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.

\[ \text{ROUGE-L} = \frac{2 \cdot P_{\text{lcs}} \cdot R_{\text{lcs}}}{P_{\text{lcs}} + R_{\text{lcs}}} \quad P_{\text{lcs}} = \frac{|\text{LCS}|}{|\text{hyp}|} \quad R_{\text{lcs}} = \frac{|\text{LCS}|}{|\text{ref}|} \]

LCS is computed with a standard O(mn) DP table — same algorithm as Edit Distance. A subsequence like ['the','cat','on','the','mat'] skips 'is'/'sat' without breaking.

19. By analogy: ROUGE-L: LCS instead of contiguous n-grams

Analogy

Discussion prompt

Explain ROUGE-L: LCS instead of contiguous n-grams by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.

20. What has to happen first: Worked Example: ROUGE-L via LCS DP

Ranking

Put in order

Put the moves of Worked Example: ROUGE-L via LCS DP into the order they have to happen.

  1. Build the LCS DP table
  2. Read LCS = dp[6][6] = 5; the matched subsequence is ['the','cat','on','the','mat']
  3. ROUGE-L: P = 5/6 = 0.8333, R = 5/6 = 0.8333, F1 = 2×0.8333×0.8333/(0.8333+0.8333) = 0.8333

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. dp[i][j] = LCS length of hyp[:i] and ref[:j].

21. Worked Example: ROUGE-L via LCS DP

Worked example

Hyp: 'the cat sat on the mat' (6 tokens). Ref: 'the cat is on the mat' (6 tokens). Goal: compute ROUGE-L P, R, F1.

Build the LCS DP table

Why: dp[i][j] = LCS length of hyp[:i] and ref[:j]. Increment if tokens match, otherwise take max of neighbors.

def lcs_length(a, b):
    m, n = len(a), len(b)
    dp = [[0]*(n+1) for _ in range(m+1)]
    for i in range(1, m+1):
        for j in range(1, n+1):
            if a[i-1] == b[j-1]:
                dp[i][j] = dp[i-1][j-1] + 1
            else:
                dp[i][j] = max(dp[i-1][j], dp[i][j-1])
    return dp[m][n]

hyp = 'the cat sat on the mat'.split()
ref = 'the cat is on the mat'.split()

lcs = lcs_length(hyp, ref)
p = lcs / len(hyp)
r = lcs / len(ref)
f1 = 2*p*r / (p+r)
print(f'LCS length: {lcs}')
print(f'P={p:.4f}  R={r:.4f}  F1={f1:.4f}')
dp row (hyp\ref)thecatisonthemat
the111111
cat122222
sat122222
on122333
the122344
mat122345

Read LCS = dp[6][6] = 5; the matched subsequence is ['the','cat','on','the','mat']

Why: 'sat' in hyp and 'is' in ref have no counterpart — they are skipped. Everything else aligns in order.

ROUGE-L: P = 5/6 = 0.8333, R = 5/6 = 0.8333, F1 = 2×0.8333×0.8333/(0.8333+0.8333) = 0.8333

Why: Both strings are length 6 and share LCS=5, so P=R=F1. Compare: BLEU-4 was 0.0 on this same pair due to p_4=0/3.

22. Inspect it line by line: Worked Example: ROUGE-L via LCS DP

Error analysis

Annotate

Walk the callouts on Worked Example: ROUGE-L via LCS DP. Each one is a place this is easy to get subtly wrong.

  • dp[i][j] = LCS length of hyp[:i] and ref[:j]. Increment if tokens match, otherwise take max of neighbors.
  • 'sat' in hyp and 'is' in ref have no counterpart — they are skipped. Everything else aligns in order.
  • Both strings are length 6 and share LCS=5, so P=R=F1. Compare: BLEU-4 was 0.0 on this same pair due to p_4=0/3.

23. BERTScore — Contextual Embedding Similarity

Section

Part 3 of 4

24. BERTScore: token-level cosine similarity

Concept

BERTScore (Zhang 2020) encodes hyp and ref through a pretrained BERT model and computes token-level cosine similarity across both sequences, then aggregates with a greedy matching strategy.

\[ P_{\text{BERT}} = \frac{1}{|\hat{x}|} \sum_{\hat{x}_i \in \hat{x}} \max_{x_j \in x} \cos(\mathbf{e}_{\hat{x}_i},\, \mathbf{e}_{x_j}) \]

\[ R_{\text{BERT}} = \frac{1}{|x|} \sum_{x_j \in x} \max_{\hat{x}_i \in \hat{x}} \cos(\mathbf{e}_{x_j},\, \mathbf{e}_{\hat{x}_i}) \]

F1_BERT = harmonic mean of P_BERT and R_BERT. Because contextual embeddings place 'feline' and 'cat' nearby in embedding space, BERTScore correctly assigns high similarity where BLEU-4 and ROUGE-1 score near 0.

25. Break it if you can: BERTScore: token-level cosine similarity

Counterexample

Discussion prompt

BERTScore (Zhang 2020) encodes hyp and ref through a pretrained BERT model and computes token-level cosine similarity across both sequences, then aggregates with a greedy matching strategy.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

26. Why BERTScore correlates better with humans

Concept

BLEU and ROUGE treat token matching as binary: 'cat' ≠ 'feline' even though they are synonyms. BERTScore uses continuous similarity — any two tokens get a score in [−1, 1] based on their contextual embedding dot product.

PairBLEU-4 ≈ROUGE-1 ≈BERTScore F1 ≈
'the cat sat on the mat' vs ref0.000.830.95+
'a feline rested on the rug' vs ref0.000.330.87+
Perfect match1.001.001.00

BERTScore requires pretrained model weights (typical: bert-base-uncased, roberta-large) and GPU inference — unlike BLEU/ROUGE which are O(N) string ops. This compute cost is the main reason BLEU remains the default MT benchmark.

27. Fill in: BERTScore F1 ≈ for Why BERTScore correlates better with humans

Comparison

Comparison matrix

From Why BERTScore correlates better with humans: refill the BERTScore F1 ≈ column from what you know. The rest of the table is as it appeared.

PairBLEU-4 ≈ROUGE-1 ≈BERTScore F1 ≈
'the cat sat on the mat' vs ref0.000.830.95+
'a feline rested on the rug' vs ref0.000.330.87+
Perfect match1.001.001.00

28. Something is wrong here: treating ROUGE-1 as sufficient for summarization…

Anomaly

Predict first

A student writes this, and it looks reasonable:

A student evaluates two summaries solely by ROUGE-1 recall and concludes the higher-scoring one is the better summary.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Reasoning: 'higher recall = better coverage = better summary.'

ROUGE-1 counts unigrams regardless of order — a bag of words. A summary with high ROUGE-1 can be incoherent or miss the key claim if it happens to share many stopwords with the reference.

Why: Reasoning: 'higher recall = better coverage = better summary.'

29. Trap: treating ROUGE-1 as sufficient for summarization quality

Trap

The trap

A student evaluates two summaries solely by ROUGE-1 recall and concludes the higher-scoring one is the better summary.

Select the summary with ROUGE-1 = 0.78 over the one with ROUGE-1 = 0.61

Why: Reasoning: 'higher recall = better coverage = better summary.'

The fix

ROUGE-1 counts unigrams regardless of order — a bag of words. A summary with high ROUGE-1 can be incoherent or miss the key claim if it happens to share many stopwords with the reference.

Report ROUGE-1, ROUGE-2, and ROUGE-L together; consider ROUGE-2 as the primary signal

Why: ROUGE-2 bigrams are harder to match without understanding — they filter out stopword-padding. ROUGE-L further penalizes disordered output. For higher-stakes evaluation (e.g. USAAIO), add BERTScore to catch meaning-preserving paraphrases.

30. Which of these survive contact with Lesson 105: BLEU, ROUGE, and BERTScore?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
The min clip prevents gaming the metric by repeating a common word. BLEU-4 takes the geometric mean of p_1 through p_4, then multiplies by a brevity penalty BP.; A model that outputs only 'the' has perfect 1-gram precision if 'the' appears in the reference. The brevity penalty BP suppresses this.; ROUGE-L uses the longest common subsequence (LCS) — the longest sequence of tokens that appears (in order, not necessarily contiguous) in both hypothesis and reference.
Breaks
A student computes p_1=0.83, p_2=0.60, p_3=0.25, p_4=0.0 and writes BLEU-4 = 0.83 × 0.60 × 0.25 × 0.0 then claims that 'BLEU-4 is very low but nonzero.'; A student evaluates two summaries solely by ROUGE-1 recall and concludes the higher-scoring one is the better summary.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 105: BLEU, ROUGE, and BERTScore puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

31. Metric Selection and Failure Modes

Section

Part 4 of 4

32. Task-metric alignment

Concept

TaskPrimary metricWhy
Machine translationBLEU-4Reference translations are dense; precision matters; fast to compute; community convention
Abstractive summarizationROUGE-1/2/LRecall of key content matters; multiple valid wordings OK
Open-ended generation (dialogue, creative)BERTScore (or human)No fixed reference; meaning > surface form
Code generationExact match / pass@kSyntax must be exact; n-gram overlap is meaningless

All three metrics assume at least one reference output exists. For tasks where 'there is no single correct answer' (story generation, chatbots), automatic metrics correlate poorly with human judgment even with BERTScore.

33. What each one costs: Task-metric alignment

Trade off

Comparison matrix

From Task-metric alignment: every row here is a choice with a cost. Fill the Primary metric column, then say which row you would actually pick and what you give up for it.

TaskPrimary metricWhy
Machine translationBLEU-4Reference translations are dense; precision matters; fast to compute; community convention
Abstractive summarizationROUGE-1/2/LRecall of key content matters; multiple valid wordings OK
Open-ended generation (dialogue, creative)BERTScore (or human)No fixed reference; meaning > surface form
Code generationExact match / pass@kSyntax must be exact; n-gram overlap is meaningless

34. Without one step: Pattern: computing and interpreting NLG metrics

Constraint

Discussion prompt

Run Pattern: computing and interpreting NLG metrics with this step confiscated:

ROUGE-L: run O(mn) LCS DP → |LCS|. P = |LCS|/|hyp|, R = |LCS|/|ref|, F1 = harmonic mean.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Tokenize hypothesis and reference(s) consistently (same lowercasing, punctuation rules).
  2. BLEU-n: for each n ∈ {1..N}, count clipped matches / total hyp n-grams → p_n. If any p_n = 0, BLEU = 0. Apply BP = min(1, exp(1 − |ref| / |hyp|)).
  3. BLEU-N = BP × exp((1/N) × Σ log p_n). Typical N = 4 for translation.
  4. ROUGE-n: count ref n-gram overlaps / total ref n-grams (recall). Report both n=1 and n=2.
  5. ROUGE-L: run O(mn) LCS DP → |LCS|. P = |LCS|/|hyp|, R = |LCS|/|ref|, F1 = harmonic mean.
  6. BERTScore: encode both with pretrained BERT; compute pairwise cosine similarity matrix; greedy-match to get P_BERT and R_BERT; take F1.
  7. Sanity check: a perfect reference-match should give BLEU=1.0, ROUGE=1.0, BERTScore=1.0. A length-zero hypothesis gives BLEU=0 (BP=0).
  8. Choose the right metric for the task — translation → BLEU, summarization → ROUGE, open-ended → BERTScore or human eval.

35. Pattern: computing and interpreting NLG metrics

Pattern

  1. Tokenize hypothesis and reference(s) consistently (same lowercasing, punctuation rules).
  2. BLEU-n: for each n ∈ {1..N}, count clipped matches / total hyp n-grams → p_n. If any p_n = 0, BLEU = 0. Apply BP = min(1, exp(1 − |ref| / |hyp|)).
  3. BLEU-N = BP × exp((1/N) × Σ log p_n). Typical N = 4 for translation.
  4. ROUGE-n: count ref n-gram overlaps / total ref n-grams (recall). Report both n=1 and n=2.
  5. ROUGE-L: run O(mn) LCS DP → |LCS|. P = |LCS|/|hyp|, R = |LCS|/|ref|, F1 = harmonic mean.
  6. BERTScore: encode both with pretrained BERT; compute pairwise cosine similarity matrix; greedy-match to get P_BERT and R_BERT; take F1.
  7. Sanity check: a perfect reference-match should give BLEU=1.0, ROUGE=1.0, BERTScore=1.0. A length-zero hypothesis gives BLEU=0 (BP=0).
  8. Choose the right metric for the task — translation → BLEU, summarization → ROUGE, open-ended → BERTScore or human eval.

36. Where does it stop working: Pattern: computing and interpreting NLG…

Edge cases

Discussion prompt

Pattern: computing and interpreting NLG metrics works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Tokenize hypothesis and reference(s) consistently (same lowercasing, punctuation rules).
  2. BLEU-n: for each n ∈ {1..N}, count clipped matches / total hyp n-grams → p_n. If any p_n = 0, BLEU = 0. Apply BP = min(1, exp(1 − |ref| / |hyp|)).
  3. BLEU-N = BP × exp((1/N) × Σ log p_n). Typical N = 4 for translation.
  4. ROUGE-n: count ref n-gram overlaps / total ref n-grams (recall). Report both n=1 and n=2.
  5. ROUGE-L: run O(mn) LCS DP → |LCS|. P = |LCS|/|hyp|, R = |LCS|/|ref|, F1 = harmonic mean.
  6. BERTScore: encode both with pretrained BERT; compute pairwise cosine similarity matrix; greedy-match to get P_BERT and R_BERT; take F1.
  7. Sanity check: a perfect reference-match should give BLEU=1.0, ROUGE=1.0, BERTScore=1.0. A length-zero hypothesis gives BLEU=0 (BP=0).
  8. Choose the right metric for the task — translation → BLEU, summarization → ROUGE, open-ended → BERTScore or human eval.

37. Rule out three: Check 1: BLEU brevity penalty

Elimination

Eliminate the wrong options

Hypothesis: 'dogs run fast' (3 tokens). Reference: 'the dogs run very fast' (5 tokens). All three hypothesis unigrams appear in the reference. What is the brevity penalty BP?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. BP = exp(1 − 5/3) ≈ 0.5134
  • B. BP = exp(1 − 3/5) = exp(0.4) ≈ 1.492
  • C. BP = 1.0 (hypothesis is shorter, no penalty)
  • D. BP = 3/5 = 0.6

Survives elimination: A

Why: BP = exp(1 − |ref|/|hyp|) when |hyp| < |ref|. Here |hyp|=3, |ref|=5, so BP = exp(1 − 5/3) = exp(−2/3) ≈ 0.5134. The penalty is applied because the hypothesis is shorter than the reference.

38. Check 1: BLEU brevity penalty

Check

Work through the arithmetic before clicking.

Check your understanding

Hypothesis: 'dogs run fast' (3 tokens). Reference: 'the dogs run very fast' (5 tokens). All three hypothesis unigrams appear in the reference. What is the brevity penalty BP?

  • A. BP = exp(1 − 5/3) ≈ 0.5134 (correct)
  • B. BP = exp(1 − 3/5) = exp(0.4) ≈ 1.492
  • C. BP = 1.0 (hypothesis is shorter, no penalty)
  • D. BP = 3/5 = 0.6

Answer: A

Why: BP = exp(1 − |ref|/|hyp|) when |hyp| < |ref|. Here |hyp|=3, |ref|=5, so BP = exp(1 − 5/3) = exp(−2/3) ≈ 0.5134. The penalty is applied because the hypothesis is shorter than the reference.

Why B tempts people
Inverted the ratio — wrote exp(1 − |hyp|/|ref|) = exp(1 − 3/5), which would give a value > 1 (impossible for a penalty). The formula always uses |ref|/|hyp| in the exponent.
Why C tempts people
Confused the direction: BP = 1 when |hyp| ≥ |ref| (hypothesis is at least as long). Here |hyp| < |ref|, so the penalty applies.
Why D tempts people
Used the ratio |hyp|/|ref| linearly without the exponential. The actual formula is exponential to ensure BP approaches 0 smoothly as hypothesis length → 0.

39. Answer it before you see the options: Check 2: ROUGE orientation

Prediction

Predict first

Reference: 'the cat sat on the mat' (6 unigrams). Hypothesis: 'the cat' (2 unigrams, both in reference). What is ROUGE-1 recall?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: 2/6 ≈ 0.333

Why: ROUGE-1 recall = (matched ref unigrams) / (total ref unigrams) = 2 / 6 ≈ 0.333. Both 'the' and 'cat' appear in the hypothesis, but the reference has 6 total unigrams, so recall is 2/6.

40. Check 2: ROUGE orientation

Check

Recall the denominator difference between BLEU and ROUGE.

Check your understanding

Reference: 'the cat sat on the mat' (6 unigrams). Hypothesis: 'the cat' (2 unigrams, both in reference). What is ROUGE-1 recall?

  • A. 2/6 ≈ 0.333 (correct)
  • B. 2/2 = 1.0
  • C. 4/6 ≈ 0.667
  • D. 2/4 = 0.5

Answer: A

Why: ROUGE-1 recall = (matched ref unigrams) / (total ref unigrams) = 2 / 6 ≈ 0.333. Both 'the' and 'cat' appear in the hypothesis, but the reference has 6 total unigrams, so recall is 2/6.

Why B tempts people
Computed BLEU-1 precision (2 clipped matches / 2 hyp tokens = 1.0) rather than ROUGE-1 recall. Precision denominator is hyp length; recall denominator is ref length.
Why C tempts people
Counted the remaining 4 reference unigrams not matched ('sat','on','the','mat') instead of the 2 that were matched. Confusion between overlap and non-overlap.
Why D tempts people
Used 4 as the denominator — perhaps counting unique reference unigrams excluding duplicates. ROUGE uses the total count of reference tokens, not the unique vocabulary size.

41. Rule out three: Check 3: BERTScore advantage

Elimination

Eliminate the wrong options

Hypothesis: 'a feline rested on the rug'. Reference: 'the cat sat on the mat'. BLEU-4 = 0 and ROUGE-1 recall ≈ 0.33. Which statement best explains why BERTScore F1 would be substantially higher?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. BERTScore counts subword tokens, so 'feline' and 'feli-' share a prefix with some reference subword.
  • B. BERTScore computes pairwise cosine similarity between contextual embeddings, so near-synonyms like 'feline'/'cat' and 'rested'/'sat' each contribute high token-level similarity scores.
  • C. BERTScore applies a larger n (n=6) for n-gram matching, capturing 'on the' as a bigram.
  • D. BERTScore ignores stopwords ('a','the') and only scores content words, inflating the score for the four matching content words.

Survives elimination: B

Why: BERTScore's core mechanism is pairwise cosine similarity between BERT contextual embeddings. 'Feline' and 'cat' are contextually similar tokens (high cosine similarity), so the greedy-max over the similarity matrix assigns each hypothesis token a high score even without a surface-level string match. This is the exact limitation that motivates BERTScore over BLEU/ROUGE.

42. Check 3: BERTScore advantage

Check

Think about what BERTScore's embedding similarity captures that BLEU/ROUGE cannot.

Check your understanding

Hypothesis: 'a feline rested on the rug'. Reference: 'the cat sat on the mat'. BLEU-4 = 0 and ROUGE-1 recall ≈ 0.33. Which statement best explains why BERTScore F1 would be substantially higher?

  • A. BERTScore counts subword tokens, so 'feline' and 'feli-' share a prefix with some reference subword.
  • B. BERTScore computes pairwise cosine similarity between contextual embeddings, so near-synonyms like 'feline'/'cat' and 'rested'/'sat' each contribute high token-level similarity scores. (correct)
  • C. BERTScore applies a larger n (n=6) for n-gram matching, capturing 'on the' as a bigram.
  • D. BERTScore ignores stopwords ('a','the') and only scores content words, inflating the score for the four matching content words.

Answer: B

Why: BERTScore's core mechanism is pairwise cosine similarity between BERT contextual embeddings. 'Feline' and 'cat' are contextually similar tokens (high cosine similarity), so the greedy-max over the similarity matrix assigns each hypothesis token a high score even without a surface-level string match. This is the exact limitation that motivates BERTScore over BLEU/ROUGE.

Why A tempts people
BERTScore operates on BERT's wordpiece subwords, but the benefit comes from embedding similarity, not prefix overlap. 'Feline' and 'cat' share no common subword prefix.
Why C tempts people
BERTScore does not use n-gram matching at all — it uses token-level embedding similarity. There is no n parameter analogous to BLEU/ROUGE.
Why D tempts people
BERTScore does not remove stopwords. It scores every token, including 'a' and 'the'. The improvement over BLEU/ROUGE comes from continuous embedding similarity, not stopword removal.

43. Your Turn: Implement BLEU and ROUGE-L from scratch

Concept

Project brief: build bleu_rouge.py that (1) implements BLEU-4 with modified n-gram precision and brevity penalty, (2) implements ROUGE-L via LCS DP, and (3) shows a failure-mode example where BLEU=0 but ROUGE-L > 0.

  1. Milestone 1: implement count_ngrams(tokens, n) returning a Counter of n-gram tuples.
  2. Milestone 2: implement modified_precision(hyp, refs, n) returning (clipped, total). Test: hyp='it is a guide to action which ensures that the military always obeys the commands of the party', ref as in the worked example → p_1=12/18=0.6667.
  3. Milestone 3: implement brevity_penalty(hyp_len, ref_len) and bleu4(hyp, refs). Test: same sentence pair → BLEU-4=0.4209.
  4. Milestone 4: implement lcs_length(a, b) with O(mn) DP and rouge_l(hyp, ref) returning (P, R, F1). Test: hyp='the cat sat on the mat', ref='the cat is on the mat' → LCS=5, F1=0.8333.
  5. Milestone 5: demonstrate a sentence pair where BLEU-4=0 but ROUGE-L F1 > 0 (hint: one of the worked examples above qualifies).

44. Full program: bleu_rouge.py

Worked example

from collections import Counter
import math

# ── N-gram utilities ─────────────────────────────────────────
def count_ngrams(tokens, n):
    return Counter(tuple(tokens[i:i+n])
                   for i in range(len(tokens)-n+1))

def modified_precision(hyp, refs, n):
    hyp_ng = count_ngrams(hyp, n)
    if not hyp_ng:
        return 0, 0
    max_ref = Counter()
    for ref in refs:
        ref_ng = count_ngrams(ref, n)
        for ng in hyp_ng:
            max_ref[ng] = max(max_ref[ng], ref_ng[ng])
    clipped = sum(min(c, max_ref[ng]) for ng, c in hyp_ng.items())
    return clipped, sum(hyp_ng.values())

# ── BLEU-4 ───────────────────────────────────────────────────
def brevity_penalty(hyp_len, ref_len):
    return 1.0 if hyp_len >= ref_len else math.exp(1 - ref_len/hyp_len)

def bleu4(hyp_str, ref_strs):
    hyp = hyp_str.split()
    refs = [r.split() for r in ref_strs]
    ref_len = min((len(r) for r in refs),
                  key=lambda x: abs(x - len(hyp)))
    bp = brevity_penalty(len(hyp), ref_len)
    log_sum = 0.0
    for n in range(1, 5):
        cn, tn = modified_precision(hyp, refs, n)
        if tn == 0 or cn == 0:
            return 0.0
        log_sum += math.log(cn / tn)
    return bp * math.exp(log_sum / 4)

# ── ROUGE-L ──────────────────────────────────────────────────
def lcs_length(a, b):
    m, n = len(a), len(b)
    dp = [[0]*(n+1) for _ in range(m+1)]
    for i in range(1, m+1):
        for j in range(1, n+1):
            if a[i-1] == b[j-1]:
                dp[i][j] = dp[i-1][j-1] + 1
            else:
                dp[i][j] = max(dp[i-1][j], dp[i][j-1])
    return dp[m][n]

def rouge_l(hyp_str, ref_str):
    h, r = hyp_str.split(), ref_str.split()
    lcs = lcs_length(h, r)
    p = lcs / len(h) if h else 0.0
    rc = lcs / len(r) if r else 0.0
    f1 = 2*p*rc/(p+rc) if (p+rc) else 0.0
    return p, rc, f1

# ── Demo ─────────────────────────────────────────────────────
if __name__ == '__main__':
    MT_HYP = ('it is a guide to action which ensures that the '
              'military always obeys the commands of the party')
    MT_REF = ('it is a guide to action that ensures that the '
              'military will forever heed party commands')
    b = bleu4(MT_HYP, [MT_REF])
    print(f'BLEU-4 (MT pair): {b:.4f}')  # 0.4209

    H = 'the cat sat on the mat'
    R = 'the cat is on the mat'
    p, rc, f1 = rouge_l(H, R)
    print(f'ROUGE-L F1:        {f1:.4f}')  # 0.8333
    print(f'BLEU-4 same pair:  {bleu4(H,[R]):.4f}')  # 0.0 (p_4=0)
CallExpected outputActual output
bleu4(MT_HYP, [MT_REF])0.42090.4209
rouge_l('the cat sat on the mat', 'the cat is on the mat') F10.83330.8333
bleu4('the cat sat on the mat', ['the cat is on the mat'])0.00000.0000

The third row is the failure-mode: BLEU-4 = 0 because p_4 = 0/3 = 0 (no shared 4-gram). ROUGE-L = 0.8333 because LCS = 5 tokens still capture the shared structure. This motivates reporting both metrics.

45. Fill in: Expected output for Full program: bleu_rouge.py

Comparison

Comparison matrix

From Full program: bleu_rouge.py: refill the Expected output column from what you know. The rest of the table is as it appeared.

CallExpected outputActual output
bleu4(MT_HYP, [MT_REF])0.42090.4209
rouge_l('the cat sat on the mat', 'the cat is on the mat') F10.83330.8333
bleu4('the cat sat on the mat', ['the cat is on the mat'])0.00000.0000

46. Show it off: corner cases to test

Concept

47. Connect it up: Lesson 105: BLEU, ROUGE, and BERTScore

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — BLEU — Bilingual Evaluation Understudy · ROUGE — Recall-Oriented Understudy · BERTScore — Contextual Embedding Similarity · Metric Selection and Failure Modes. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

48. Lesson 105 recap: BLEU, ROUGE, BERTScore

Recap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 105 — BLEU, ROUGE, BERTScore — Barron · USAAIO Round 2 Preparation, 2026
  2. Papineni et al. 'BLEU: a Method for Automatic Evaluation of Machine Translation' (ACL 2002) — ACL Anthology
  3. Lin, C.-Y. 'ROUGE: A Package for Automatic Evaluation of Summaries' (ACL 2004 Workshop) — ACL Anthology
  4. Zhang et al. 'BERTScore: Evaluating Text Generation with BERT' (ICLR 2020) — arXiv:1904.09675
  5. All BLEU/ROUGE computations implemented from scratch and verified with Python 3 / numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108