USAAIO Lesson 108, from Phase 3. It contrasts extractive with abstractive summarization, then covers BART's denoising pretraining - text infilling, token deletion, and sentence permutation - and Pegasus's sentence-level masking. It defines perplexity as exponentiated cross-entropy and traces the log-probabilities concretely on a three-token toy language model, giving 1.26 for the good model, 1.42 for the bad one, and 3 for the uniform one. A TinyBART encoder-decoder of 22,996 parameters is then fine-tuned on a copy task, taking perplexity from 11.02 to 1.42 over 50 epochs. It closes with calibration and expected calibration error, using a five-bin reliability table in which the calibrated model scores 0.0296 against 0.1862 for the overconfident one. All the numbers were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 31 slides.
Subject: Machine Learning · 56 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 108 · Phase 3
Abstractive vs extractive summarization, BART denoising pretraining, Pegasus gap-sentence masking, perplexity as exponentiated NLL, and ECE calibration — every number verified in PyTorch.
Objectives
Section
Part 1 of 4
Concept
Extractive summarization selects and concatenates complete sentences (or phrases) from the source document. No new tokens are generated — the summary is a subset of the input.
Classic extractive systems: TextRank (graph-based sentence scoring), BertSum (Lesson 90 BERT encoder + sentence-selection head).
Counterexample
Discussion prompt
Classic extractive systems: TextRank (graph-based sentence scoring), BertSum (Lesson 90 BERT encoder + sentence-selection head).
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Abstractive summarization generates tokens that may never appear in the source. The model must compress, paraphrase, and synthesize — the same task a human writer performs.
| method | generates new tokens | ROUGE-1 F1 (vs source) | hallucination risk |
|---|---|---|---|
| extractive | no | high (copies n-grams) | none |
| abstractive | yes | lower (paraphrase) | present |
| hybrid | yes (constrained) | moderate | reduced |
Comparison
Comparison matrix
From Abstractive summarization: generate new text: refill the generates new tokens column from what you know. The rest of the table is as it appeared.
| method | generates new tokens | ROUGE-1 F1 (vs source) | hallucination risk |
|---|---|---|---|
| extractive | no | high (copies n-grams) | none |
| abstractive | yes | lower (paraphrase) | present |
| hybrid | yes (constrained) | moderate | reduced |
Estimation
Predict first
Reference: the cat sat on the mat near the window (9 tokens). Compare an extractive summary (literal copy of first 6 tokens) and an abstractive paraphrase. Compute ROUGE-1 and ROUGE-2 F1 for each.
Commit before you compute: what does ROUGE-1 and ROUGE-2 trace — extractive vs abstractive come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Extractive ROUGE-1=0.8000, ROUGE-2=0.7692 — abstractive ROUGE-1=0.2667, ROUGE-2=0.1538
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The extractive copy shares 6 of 9 reference unigrams verbatim (precision=1.0, recall=6/9=0.667).
Worked example
Reference: the cat sat on the mat near the window (9 tokens). Compare an extractive summary (literal copy of first 6 tokens) and an abstractive paraphrase. Compute ROUGE-1 and ROUGE-2 F1 for each.
def rouge_n_f1(ref, hyp, n):
def ngrams(s, n):
toks = s.lower().split()
return [tuple(toks[i:i+n]) for i in range(len(toks)-n+1)]
rb = set(ngrams(ref, n))
hb = ngrams(hyp, n)
if not hb or not rb:
return 0.0
overlap = sum(1 for g in hb if g in rb)
prec = overlap / len(hb)
rec = overlap / len(rb)
return 2*prec*rec/(prec+rec) if prec+rec else 0.0
ref = 'the cat sat on the mat near the window'
ext = 'the cat sat on the mat' # extractive copy
abs_ = 'a feline rested beside the window' # abstractive
for name, hyp in [('extractive', ext), ('abstractive', abs_)]:
r1 = rouge_n_f1(ref, hyp, 1)
r2 = rouge_n_f1(ref, hyp, 2)
print(f'{name}: ROUGE-1={r1:.4f} ROUGE-2={r2:.4f}')
# extractive: ROUGE-1=0.8000 ROUGE-2=0.7692
# abstractive: ROUGE-1=0.2667 ROUGE-2=0.1538Extractive ROUGE-1=0.8000, ROUGE-2=0.7692 — abstractive ROUGE-1=0.2667, ROUGE-2=0.1538
Why: The extractive copy shares 6 of 9 reference unigrams verbatim (precision=1.0, recall=6/9=0.667). The abstractive only shares 'the' and 'window' (2 tokens). ROUGE rewards surface n-gram overlap, penalizing paraphrase even when semantically faithful.
| summary | shared unigrams | ROUGE-1 F1 | ROUGE-2 F1 |
|---|---|---|---|
| extractive ('the cat sat on the mat') | 6/9 of ref | 0.8000 | 0.7692 |
| abstractive ('a feline rested beside the window') | 2/9 of ref | 0.2667 | 0.1538 |
| perfect copy of reference | 9/9 | 1.0000 | 1.0000 |
Trade off
Comparison matrix
From ROUGE-1 and ROUGE-2 trace — extractive vs abstractive: every row here is a choice with a cost. Fill the ROUGE-2 F1 column, then say which row you would actually pick and what you give up for it.
| summary | shared unigrams | ROUGE-1 F1 | ROUGE-2 F1 |
|---|---|---|---|
| extractive ('the cat sat on the mat') | 6/9 of ref | 0.8000 | 0.7692 |
| abstractive ('a feline rested beside the window') | 2/9 of ref | 0.2667 | 0.1538 |
| perfect copy of reference | 9/9 | 1.0000 | 1.0000 |
Anomaly
Predict first
A student writes this, and it looks reasonable:
ROUGE is the gold standard for summarization quality. A model achieving ROUGE-1=0.80 is always better than one scoring ROUGE-1=0.27.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: ROUGE is a surface n-gram metric.
ROUGE measures n-gram overlap. Use it alongside factual consistency checks (BertScore, QA-based) and human preference studies — especially for abstractive systems where paraphrase is correct but ROUGE penalizes it.
Why: ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.
Trap
ROUGE is the gold standard for summarization quality. A model achieving ROUGE-1=0.80 is always better than one scoring ROUGE-1=0.27.
Rank models by ROUGE alone; discard any system with ROUGE below 0.5
Why: ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.
ROUGE measures n-gram overlap. Use it alongside factual consistency checks (BertScore, QA-based) and human preference studies — especially for abstractive systems where paraphrase is correct but ROUGE penalizes it.
Combine ROUGE with BertScore and hallucination detection; never rank abstractive against extractive by ROUGE alone
Why: The USAAIO exam expects you to know ROUGE's limits: it cannot distinguish a correct paraphrase from a wrong one, and an extractive oracle trivially maximizes it.
Break the constraint
Discussion prompt
The rule this trap just fixed:
The USAAIO exam expects you to know ROUGE's limits: it cannot distinguish a correct paraphrase from a wrong one, and an extractive oracle trivially maximizes it.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.
Section
Part 2 of 4
Concept
BART (Lewis et al., 2020) is a seq2seq denoising autoencoder: the encoder reads a corrupted version of the text and the decoder learns to reconstruct the original — exactly like a denoising autoencoder, but at the token sequence level.
Architecture: standard Transformer encoder (like BERT, Lesson 90 — bidirectional attention, no causal mask) stacked with a standard Transformer decoder (like GPT, Lesson 91 — causal mask, cross-attention to encoder output). Same shape as the original 'Attention is All You Need' (Lesson 88).
| component | attention type | analogous model | role |
|---|---|---|---|
| encoder | bidirectional (no mask) | BERT (L90) | read corrupted input |
| decoder | causal + cross-attention | GPT (L91) | generate clean output |
| cross-attention | decoder queries encoder | standard seq2seq | condition on source |
Concept
BART pretraining applies one or more noise functions to the source text. The decoder must reconstruct the original token sequence regardless of how it was corrupted.
[MASK] — the decoder must predict the number of missing tokens as well as their identity. Stronger than BERT's independent masking.| corruption | example input to encoder | decoder target |
|---|---|---|
| text infilling | The [MASK] over the lazy dog | The quick brown fox jumped over the lazy dog |
| token deletion | The quick brown over the lazy dog | The quick brown fox jumped over the lazy dog |
| sentence permutation | jumped over the lazy dog The quick brown fox | The quick brown fox jumped over the lazy dog |
| document rotation | over the lazy dog The quick brown fox jumped | The quick brown fox jumped over the lazy dog |
Analogy
Discussion prompt
Explain BART's four corruption strategies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
BART pretraining applies one or more noise functions to the source text. The decoder must reconstruct the original token sequence regardless of how it was corrupted.
Concept
Pegasus (Zhang et al., 2020) replaces entire sentences with [MASK] tokens, selecting the sentences that are most important to the rest of the document — the same sentences a human would include in a summary.
Importance is scored by ROUGE-1 recall between the candidate sentence and the rest of the document. The top-k scoring sentences are masked (Gap Sentence Generation, GSG). The model learns to generate exactly the content humans use in extractive summaries — aligning pretraining directly with the downstream task.
| model | pretraining objective | unit masked | mask selection |
|---|---|---|---|
| BERT (L90) | masked language model | individual tokens (15%) | random |
| BART | denoising autoencoder | spans or sentences | random / span-sampled |
| Pegasus | gap sentence generation (GSG) | whole sentences | highest ROUGE-1 vs rest of doc |
Missing information
Discussion prompt
Build a minimal BART-shaped encoder-decoder with vocab_size=20, d_model=32, n_heads=2, 1 encoder layer, 1 decoder layer. Trace shapes through the forward pass.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Embedding tables: 20×32=640 each (×2 shared, plus position). TransformerEncoderLayer(32,2,ff=64): self-attn 32×32×4+biases + ff 32×64×2+biases ≈ 6,400. Decoder adds cross-attention block ≈ 8,480. Head: 32×20+20=660. Verified: 22,996 params total.
Worked example
Build a minimal BART-shaped encoder-decoder with vocab_size=20, d_model=32, n_heads=2, 1 encoder layer, 1 decoder layer. Trace shapes through the forward pass.
import torch, torch.nn as nn
class TinyBART(nn.Module):
def __init__(self, V=20, d=32, h=2, max_len=10):
super().__init__()
self.emb = nn.Embedding(V, d)
self.pos = nn.Embedding(max_len, d)
enc_l = nn.TransformerEncoderLayer(d, h, dim_feedforward=64,
batch_first=True, dropout=0.0)
self.enc = nn.TransformerEncoder(enc_l, num_layers=1)
dec_l = nn.TransformerDecoderLayer(d, h, dim_feedforward=64,
batch_first=True, dropout=0.0)
self.dec = nn.TransformerDecoder(dec_l, num_layers=1)
self.head = nn.Linear(d, V)
def forward(self, src, tgt):
S, T = src.shape[1], tgt.shape[1]
enc_out = self.enc(self.emb(src) + self.pos(torch.arange(S)))
mask = nn.Transformer.generate_square_subsequent_mask(T)
dec_out = self.dec(self.emb(tgt)+self.pos(torch.arange(T)),
enc_out, tgt_mask=mask, tgt_is_causal=True)
return self.head(dec_out)
torch.manual_seed(42)
model = TinyBART()
print(sum(p.numel() for p in model.parameters()), 'params') # 22996
src = torch.randint(0, 20, (2, 5)) # B=2, src_len=5
tgt = torch.randint(0, 20, (2, 4)) # B=2, tgt_len=4
out = model(src, tgt)
print(out.shape) # torch.Size([2, 4, 20])22,996 parameters; output shape (2, 4, 20) — batch × target_len × vocab_size
Why: Embedding tables: 20×32=640 each (×2 shared, plus position). TransformerEncoderLayer(32,2,ff=64): self-attn 32×32×4+biases + ff 32×64×2+biases ≈ 6,400. Decoder adds cross-attention block ≈ 8,480. Head: 32×20+20=660. Verified: 22,996 params total.
| tensor | shape | note |
|---|---|---|
| src tokens | (2, 5) | B=2 sequences, src_len=5 |
| tgt tokens | (2, 4) | B=2 sequences, tgt_len=4 |
| enc_out | (2, 5, 32) | encoder reads all 5 src positions |
| dec_out | (2, 4, 32) | decoder: causal over tgt, cross-attends enc_out |
| logits (head) | (2, 4, 20) | one distribution over V=20 per tgt position |
Comparison
Comparison matrix
From TinyBART — encoder-decoder forward pass: refill the note column from what you know. The rest of the table is as it appeared.
| tensor | shape | note |
|---|---|---|
| src tokens | (2, 5) | B=2 sequences, src_len=5 |
| tgt tokens | (2, 4) | B=2 sequences, tgt_len=4 |
| enc_out | (2, 5, 32) | encoder reads all 5 src positions |
| dec_out | (2, 4, 32) | decoder: causal over tgt, cross-attends enc_out |
| logits (head) | (2, 4, 20) | one distribution over V=20 per tgt position |
Section
Part 3 of 4
Concept
\[ \text{PPL}(x_1,\ldots,x_N) = \exp\!\left(-\frac{1}{N}\sum_{i=1}^{N} \log P(x_i \mid x_{<i})\right) \]
PPL is the geometric mean inverse probability of each token given its context. Intuitively: a model with PPL=100 is as confused as a model that assigns equal probability to 100 choices at every step.
Estimation
Predict first
Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.
Commit before you compute: what does Perplexity trace on a 3-token toy LM come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Both transitions are high-probability (the LM was designed that way).
Worked example
Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.
import torch, math
# Row = context token (0=the,1=cat,2=sat), Col = next token
logits = torch.tensor([
[0.1, 2.0, 0.3], # after 'the': cat very likely
[0.2, 0.1, 2.5], # after 'cat': sat very likely
[1.5, 0.2, 0.1], # after 'sat': the most likely
])
log_p = torch.log_softmax(logits, dim=-1)
print(log_p.numpy().round(4))
# [[-2.1869 -0.2869 -1.9869]
# [-2.4748 -2.5748 -0.1748]
# [-0.4181 -1.7181 -1.8181]]
def ppl(seq, log_p):
lps = [log_p[seq[i], seq[i+1]].item() for i in range(len(seq)-1)]
return math.exp(-sum(lps)/len(lps))
good = [0, 1, 2] # the -> cat -> sat
bad = [2, 0, 1] # sat -> the -> cat
print(f'PPL(the cat sat) = {ppl(good, log_p):.4f}') # 1.2596
print(f'PPL(sat the cat) = {ppl(bad, log_p):.4f}') # 1.4226
print(f'PPL(uniform V=3) = 3')log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596
Why: Both transitions are high-probability (the LM was designed that way). PPL=1.26 is well below the uniform baseline of 3, confirming the LM has learned the 'the cat sat' pattern. The bad sequence scores PPL=1.42 — still low because each transition happens to be moderately likely under the toy LM.
| sequence | log P(step 1) | log P(step 2) | mean NLL | PPL |
|---|---|---|---|---|
| the → cat → sat (good) | -0.2869 | -0.1748 | 0.2308 | 1.2596 |
| sat → the → cat (bad) | -0.4181 | -1.7181 | 1.0681 | 2.9095 |
| uniform over V=3 | — | — | log 3 | 3.0000 |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.
Concept
Sorting
Sort into buckets
These are the pieces of Lesson 108: Summarization, BART, Perplexity & Calibration, out of order. Put each one back under the part of the lesson it belongs to.
Anomaly
Predict first
A student writes this, and it looks reasonable:
PPL is computed over the raw token count N. A short 3-token sequence has higher PPL than a 100-token sequence from the same model, because the 100-token sequence averages more steps.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent).
PPL is already length-normalized by the 1/N factor. The meaningful comparison is between two models evaluated on the same test set, not between a short and a long sequence.
Why: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent). A 3-token sequence and a 100-token sequence produce comparable PPL values when drawn from the same distribution — length does not systematically inflate or deflate PPL.
Trap
PPL is computed over the raw token count N. A short 3-token sequence has higher PPL than a 100-token sequence from the same model, because the 100-token sequence averages more steps.
Compare PPL(3-token summary) vs PPL(100-token paragraph) to judge model quality
Why: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent). A 3-token sequence and a 100-token sequence produce comparable PPL values when drawn from the same distribution — length does not systematically inflate or deflate PPL.
PPL is already length-normalized by the 1/N factor. The meaningful comparison is between two models evaluated on the same test set, not between a short and a long sequence.
Only compare PPL values when both models share the same vocabulary and the same test set
Why: Different vocabularies change the difficulty of each prediction step; different test sets change the distribution. PPL is a valid ranking within a fixed evaluation protocol.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
torch.manual_seed(42) for reproducibility.Section
Part 4 of 4
Concept
A model is calibrated when its predicted probability matches the empirical frequency of being correct. If the model says 70% on 1,000 examples, roughly 700 should be correct. Modern deep networks are typically overconfident — they output 95% when they are right only 70% of the time.
\[ \text{ECE} = \sum_{b=1}^{B} \frac{|\mathcal{B}_b|}{N} \left| \text{acc}(\mathcal{B}_b) - \text{conf}(\mathcal{B}_b) \right| \]
ECE bins predictions by confidence, then computes the weighted average absolute gap between accuracy and confidence within each bin. ECE = 0 = perfectly calibrated; ECE = 0.1 means confidence is off by 10 percentage points on average.
Estimation
Predict first
Simulate two models on n=200 examples with 5 confidence bins. Compute ECE for each. Predict which model has lower ECE before running.
Commit before you compute: what does ECE computation — calibrated vs overconfident come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: ECE calibrated=0.0296; ECE overconfident=0.1862 — a 6× gap
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The calibrated model's confidence tracks empirical accuracy within each bin (small |acc-conf|).
Worked example
Simulate two models on n=200 examples with 5 confidence bins. Compute ECE for each. Predict which model has lower ECE before running.
import numpy as np
def ece(conf, correct, n_bins=5):
bins = np.linspace(0, 1, n_bins + 1)
total = 0.0
for b in range(n_bins):
mask = (conf >= bins[b]) & (conf < bins[b+1])
if mask.sum() == 0:
continue
bin_acc = correct[mask].mean()
bin_conf = conf[mask].mean()
total += (mask.sum() / len(conf)) * abs(bin_acc - bin_conf)
return total
np.random.seed(0)
n = 200
conf_cal = np.random.uniform(0.5, 1.0, n)
corr_cal = (np.random.rand(n) < conf_cal).astype(float)
conf_over = np.clip(conf_cal * 1.25, 0, 1.0) # inflated confidence
corr_over = (np.random.rand(n) < conf_cal*0.75).astype(float)
print(f'ECE calibrated: {ece(conf_cal, corr_cal):.4f}') # 0.0296
print(f'ECE overconfident: {ece(conf_over, corr_over):.4f}') # 0.1862ECE calibrated=0.0296; ECE overconfident=0.1862 — a 6× gap
Why: The calibrated model's confidence tracks empirical accuracy within each bin (small |acc-conf|). The overconfident model inflates confidence by 25% while accuracy stays the same, producing large within-bin gaps that sum to ECE≈0.19.
| bin | n | avg conf | avg acc | |diff| |
|---|---|---|---|---|
| [0.4, 0.6) | 41 | 0.5505 | 0.4634 | 0.0871 |
| [0.6, 0.8) | 77 | 0.7067 | 0.7013 | 0.0054 |
| [0.8, 1.0) | 82 | 0.8909 | 0.9146 | 0.0237 |
| ECE (weighted avg) | 200 | — | — | 0.0296 |
Trade off
Comparison matrix
From ECE computation — calibrated vs overconfident: every row here is a choice with a cost. Fill the n column, then say which row you would actually pick and what you give up for it.
| bin | n | avg conf | avg acc | |diff| |
|---|---|---|---|---|
| [0.4, 0.6) | 41 | 0.5505 | 0.4634 | 0.0871 |
| [0.6, 0.8) | 77 | 0.7067 | 0.7013 | 0.0054 |
| [0.8, 1.0) | 82 | 0.8909 | 0.9146 | 0.0237 |
| ECE (weighted avg) | 200 | — | — | 0.0296 |
Constraint
Discussion prompt
Run The Lesson 108 recipe with this step confiscated:
Perplexity: PPL = exp(-1/N ∑ log P(x_i|x_<i)). Lower = better. Uniform baseline = V. Only compare models with the same vocabulary and test set.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
Pattern
Edge cases
Discussion prompt
The Lesson 108 recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
Elimination
Eliminate the wrong options
Which BART corruption strategy requires the decoder to also predict how many tokens were removed, not just which tokens?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Text infilling replaces a contiguous span of k tokens with a single [MASK] token. The decoder sees one mask but must generate k tokens — it must learn to infer the span length from context. Token deletion drops tokens entirely but the decoder knows the original length from the target; sentence permutation and rotation do not remove tokens.
Check
Think about what makes each BART corruption strategy unique.
Check your understanding
Which BART corruption strategy requires the decoder to also predict how many tokens were removed, not just which tokens?
Answer: A
Why: Text infilling replaces a contiguous span of k tokens with a single [MASK] token. The decoder sees one mask but must generate k tokens — it must learn to infer the span length from context. Token deletion drops tokens entirely but the decoder knows the original length from the target; sentence permutation and rotation do not remove tokens.
Prediction
Predict first
A language model has vocabulary size V=10,000. On a test set, it achieves PPL=100. Which statement is correct?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: The model is effectively choosing uniformly among 100 tokens at each step — still far from a perfect predictor but much better than random.
Why: PPL=100 means the model's average per-step uncertainty equals a uniform distribution over 100 tokens. The uniform baseline for this model is PPL=V=10,000, so the model is 100× better than random (10,000/100=100×). PPL is not token accuracy — a model can have PPL<10 while still predicting incorrectly on many tokens because probabilities don't have to be 1.0.
Check
Apply the PPL formula and the uniform baseline.
Check your understanding
A language model has vocabulary size V=10,000. On a test set, it achieves PPL=100. Which statement is correct?
Answer: B
Why: PPL=100 means the model's average per-step uncertainty equals a uniform distribution over 100 tokens. The uniform baseline for this model is PPL=V=10,000, so the model is 100× better than random (10,000/100=100×). PPL is not token accuracy — a model can have PPL<10 while still predicting incorrectly on many tokens because probabilities don't have to be 1.0.
Elimination
Eliminate the wrong options
A model outputs 0.95 confidence on 500 examples. 400 of those are correct (80% accuracy). Which fixes this overconfidence best?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: B
Why: Temperature scaling is the standard post-hoc calibration method. Dividing logits by T>1 softens the softmax distribution (reduces confidence) without changing the argmax prediction or retraining. It is fit on a held-out validation set by minimizing NLL, and is the most practical fix for overconfidence identified in Guo et al. 2017.
Check
Reason about the reliability diagram before selecting.
Check your understanding
A model outputs 0.95 confidence on 500 examples. 400 of those are correct (80% accuracy). Which fixes this overconfidence best?
Answer: B
Why: Temperature scaling is the standard post-hoc calibration method. Dividing logits by T>1 softens the softmax distribution (reduces confidence) without changing the argmax prediction or retraining. It is fit on a held-out validation set by minimizing NLL, and is the most practical fix for overconfidence identified in Guo et al. 2017.
Section
Project
Concept
Build and fine-tune TinyBART — a minimal encoder-decoder transformer — on a synthetic copy task. Track PPL as training progresses to see the loss curve a real summarization fine-tune would produce.
| # | milestone | key tool |
|---|---|---|
| 1 | Implement TinyBART (encoder + decoder + causal mask), trace shapes through forward pass | nn.TransformerEncoder, nn.TransformerDecoder, generate_square_subsequent_mask |
| 2 | Write a perplexity function: given model logits and labels, return exp(cross-entropy) | log_softmax, gather, mean, exp |
| 3 | Train 50 epochs on copy task, log PPL every 10 epochs — target PPL < 2.0 at epoch 50 | Adam lr=1e-3, CrossEntropyLoss |
Build rules: print tensor shapes after encoder and decoder; verify causal mask shape is (T, T); log both loss and PPL; use torch.manual_seed(42) for reproducibility.
Counterexample
Discussion prompt
Build and fine-tune TinyBART — a minimal encoder-decoder transformer — on a synthetic copy task. Track PPL as training progresses to see the loss curve a real summarization fine-tune would produce.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Build rules: print tensor shapes after encoder and decoder; verify causal mask shape is (T, T); log both loss and PPL; use torch.manual_seed(42) for reproducibility.
Worked example
Your turn: given model output logits of shape (B, T, V) and labels (B, T), compute per-sequence PPL. Predict the PPL at random initialization for V=10.
Hint: use log_softmax(logits, dim=-1), then gather to select log-prob at each label index, then mean(dim=1) for per-sequence NLL, then exp.
import torch, math
def ppl_from_logits(logits, labels):
# logits: (B, T, V), labels: (B, T)
log_p = torch.log_softmax(logits, dim=-1) # (B, T, V)
lp_label = log_p.gather(2, labels.unsqueeze(-1)) # (B, T, 1)
lp_label = lp_label.squeeze(-1) # (B, T)
nll = -lp_label.mean(dim=1) # (B,)
return nll.exp() # (B,) PPL per seq
torch.manual_seed(42)
logits = torch.randn(2, 4, 10) # random-init, V=10
labels = torch.randint(0, 10, (2, 4))
ppl = ppl_from_logits(logits, labels)
print(ppl) # tensor([24.2912, 19.5234]) (random init ~V=10 range)| tensor | shape | operation |
|---|---|---|
| logits | (2, 4, 10) | raw model output |
| log_p | (2, 4, 10) | log_softmax over vocab dim |
| lp_label | (2, 4) | gather log-prob at label index |
| nll per seq | (2,) | mean over T positions |
| PPL | (2,) | exp(nll): [24.29, 19.52] |
Discrimination
Sort into buckets
Sort these by shape, from memory, without looking back at Milestone 2 — perplexity from logits. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Worked example
Your turn: train TinyBART for 50 epochs on 32 random sequences (V=10, seq_len=5). Predict: will PPL converge below 2.0? Estimate at how many epochs.
Hint: the copy task is easy (label = src); Adam at lr=1e-3 converges faster than the transformer default 3e-4. Teacher-forcing (feeding ground truth as decoder input) keeps the loss stable.
import torch, torch.nn as nn, math
torch.manual_seed(42)
VOCAB = 10
model = TinyBART(V=VOCAB, d=32, h=2, max_len=8) # from Milestone 1
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
src_data = torch.randint(0, VOCAB, (32, 5))
labels = src_data # copy task: predict src
for epoch in range(51):
opt.zero_grad()
logits = model(src_data, src_data) # teacher-force tgt = src
loss = loss_fn(logits.reshape(-1, VOCAB), labels.reshape(-1))
loss.backward(); opt.step()
if epoch % 10 == 0:
ppl = math.exp(loss.item())
print(f'epoch {epoch:3d} loss={loss.item():.4f} PPL={ppl:.2f}')
# epoch 0 loss=2.3997 PPL=11.02
# epoch 10 loss=1.7515 PPL=5.76
# epoch 20 loss=1.2176 PPL=3.38
# epoch 30 loss=0.8125 PPL=2.25
# epoch 40 loss=0.5339 PPL=1.71
# epoch 50 loss=0.3535 PPL=1.42| epoch | loss | PPL | interpretation |
|---|---|---|---|
| 0 | 2.3997 | 11.02 | random init; above V=10 average because labels are concentrated |
| 10 | 1.7515 | 5.76 | model starts copying short patterns |
| 20 | 1.2176 | 3.38 | approaching meaningful compression |
| 30 | 0.8125 | 2.25 | task nearly learned |
| 40 | 0.5339 | 1.71 | PPL < 2 — well below random |
| 50 | 0.3535 | 1.42 | strong copy; further epochs approach PPL=1 |
Concept
import torch, torch.nn as nn, math
class TinyBART(nn.Module):
def __init__(self, V=10, d=32, h=2, max_len=8):
super().__init__()
self.emb = nn.Embedding(V, d)
self.pos = nn.Embedding(max_len, d)
self.enc = nn.TransformerEncoder(
nn.TransformerEncoderLayer(d, h, dim_feedforward=64,
batch_first=True, dropout=0.0), 1)
self.dec = nn.TransformerDecoder(
nn.TransformerDecoderLayer(d, h, dim_feedforward=64,
batch_first=True, dropout=0.0), 1)
self.head = nn.Linear(d, V)
def forward(self, src, tgt):
S, T = src.shape[1], tgt.shape[1]
enc_out = self.enc(self.emb(src)+self.pos(torch.arange(S)))
mask = nn.Transformer.generate_square_subsequent_mask(T)
dec_out = self.dec(self.emb(tgt)+self.pos(torch.arange(T)),
enc_out, tgt_mask=mask, tgt_is_causal=True)
return self.head(dec_out)
torch.manual_seed(42)
VOCAB = 10
model = TinyBART(V=VOCAB)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
src = torch.randint(0, VOCAB, (32, 5))
for ep in range(51):
opt.zero_grad()
loss = loss_fn(model(src, src).reshape(-1, VOCAB), src.reshape(-1))
loss.backward(); opt.step()
if ep % 10 == 0:
print(f'ep {ep:3d} PPL={math.exp(loss.item()):.2f}')
# ep 0 PPL=11.02
# ep 10 PPL=5.76
# ep 20 PPL=3.38
# ep 30 PPL=2.25
# ep 40 PPL=1.71
# ep 50 PPL=1.42| design choice | value | rationale |
|---|---|---|
| vocab size V | 10 | small → copy task learned quickly |
| d_model | 32 | minimal but valid transformer dim |
| n_heads | 2 | d_k = 32/2 = 16 per head |
| learning rate | 1e-3 | copy task is simple; faster lr than summarization |
| epochs | 50 | PPL crosses 2.0 at epoch ~35 |
| params | 22,996 (V=20 variant) | verified with named_parameters() |
Comparison
Comparison matrix
From The full program: refill the rationale column from what you know. The rest of the table is as it appeared.
| design choice | value | rationale |
|---|---|---|
| vocab size V | 10 | small → copy task learned quickly |
| d_model | 32 | minimal but valid transformer dim |
| n_heads | 2 | d_k = 32/2 = 16 per head |
| learning rate | 1e-3 | copy task is simple; faster lr than summarization |
| epochs | 50 | PPL crosses 2.0 at epoch ~35 |
| params | 22,996 (V=20 variant) | verified with named_parameters() |
Concept
Out loud, slides closed: (1) explain why extractive summarization always outscores abstractive on ROUGE-1 against the source; (2) walk through all four BART corruption strategies and state which one is unique because the decoder must infer span length; (3) compute PPL by hand given two log-prob values.
Homework (from the lesson plan): fine-tune BART-base for abstractive summarization on CNN/DailyMail (use Hugging Face datasets); implement a perplexity function that works for any autoregressive language model; compare extractive (first-3-sentences baseline) vs abstractive (BART) using ROUGE and at least one factual-consistency metric.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Extractive vs Abstractive — what ROUGE reveals · BART — denoising encoder-decoder pretraining · Perplexity — exponentiated cross-entropy · Calibration — does confidence match accuracy? · Your turn: TinyBART copy-task fine-tuning. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
| concept | the one thing to remember |
|---|---|
| extractive | selects sentences — ROUGE inflated, no hallucination |
| abstractive | generates new text — lower ROUGE, hallucination risk |
| BART | encoder reads corrupted, decoder reconstructs clean |
| Pegasus | mask highest-ROUGE-1 sentences to align pretraining with summarization |
| perplexity | PPL=exp(-1/N ∑ log P); lower=better; uniform baseline=V |
| calibration / ECE | ECE = weighted |acc-conf|; overconfident → high ECE; fix: temperature scaling |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.