Lesson 108: Summarization, BART, Perplexity & Calibration

USAAIO Lesson 108, from Phase 3. It contrasts extractive with abstractive summarization, then covers BART's denoising pretraining - text infilling, token deletion, and sentence permutation - and Pegasus's sentence-level masking. It defines perplexity as exponentiated cross-entropy and traces the log-probabilities concretely on a three-token toy language model, giving 1.26 for the good model, 1.42 for the bad one, and 3 for the uniform one. A TinyBART encoder-decoder of 22,996 parameters is then fine-tuned on a copy task, taking perplexity from 11.02 to 1.42 over 50 epochs. It closes with calibration and expected calibration error, using a five-bin reliability table in which the calibrated model scores 0.0296 against 0.1862 for the overconfident one. All the numbers were verified with torch 2.7.1+cpu and numpy 2.2.6. The lesson runs to 31 slides.

Subject: Machine Learning · 56 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Summarization, BART Perplexity & Calibration

Title

USAAIO · Lesson 108 · Phase 3

Abstractive vs extractive summarization, BART denoising pretraining, Pegasus gap-sentence masking, perplexity as exponentiated NLL, and ECE calibration — every number verified in PyTorch.

2. By the end of this lesson you can

Objectives

  1. Distinguish extractive from abstractive summarization and state the ROUGE-1/2 asymmetry that reveals the difference
  2. Describe BART's four corruption strategies (text infilling, token deletion, sentence permutation, document rotation) and explain why a denoising objective trains a strong encoder-decoder
  3. Explain Pegasus gap-sentence masking — how sentence importance is scored, why it aligns pretraining with the summarization objective
  4. Compute perplexity by hand from a token-level log-prob table and interpret low/high/uniform-baseline values
  5. Explain calibration and ECE — read a reliability diagram, identify an overconfident model, and state when perplexity is a misleading metric

3. Extractive vs Abstractive — what ROUGE reveals

Section

Part 1 of 4

4. Extractive summarization: copy, do not generate

Concept

Extractive summarization selects and concatenates complete sentences (or phrases) from the source document. No new tokens are generated — the summary is a subset of the input.

Classic extractive systems: TextRank (graph-based sentence scoring), BertSum (Lesson 90 BERT encoder + sentence-selection head).

5. Break it if you can: Extractive summarization: copy, do not generate

Counterexample

Discussion prompt

Classic extractive systems: TextRank (graph-based sentence scoring), BertSum (Lesson 90 BERT encoder + sentence-selection head).

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

6. Abstractive summarization: generate new text

Concept

Abstractive summarization generates tokens that may never appear in the source. The model must compress, paraphrase, and synthesize — the same task a human writer performs.

methodgenerates new tokensROUGE-1 F1 (vs source)hallucination risk
extractivenohigh (copies n-grams)none
abstractiveyeslower (paraphrase)present
hybridyes (constrained)moderatereduced

7. Fill in: generates new tokens for Abstractive summarization: generate new text

Comparison

Comparison matrix

From Abstractive summarization: generate new text: refill the generates new tokens column from what you know. The rest of the table is as it appeared.

methodgenerates new tokensROUGE-1 F1 (vs source)hallucination risk
extractivenohigh (copies n-grams)none
abstractiveyeslower (paraphrase)present
hybridyes (constrained)moderatereduced

8. Guess the shape of the answer: ROUGE-1 and ROUGE-2 trace — extractive vs…

Estimation

Predict first

Reference: the cat sat on the mat near the window (9 tokens). Compare an extractive summary (literal copy of first 6 tokens) and an abstractive paraphrase. Compute ROUGE-1 and ROUGE-2 F1 for each.

Commit before you compute: what does ROUGE-1 and ROUGE-2 trace — extractive vs abstractive come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Extractive ROUGE-1=0.8000, ROUGE-2=0.7692 — abstractive ROUGE-1=0.2667, ROUGE-2=0.1538

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The extractive copy shares 6 of 9 reference unigrams verbatim (precision=1.0, recall=6/9=0.667).

9. ROUGE-1 and ROUGE-2 trace — extractive vs abstractive

Worked example

Reference: the cat sat on the mat near the window (9 tokens). Compare an extractive summary (literal copy of first 6 tokens) and an abstractive paraphrase. Compute ROUGE-1 and ROUGE-2 F1 for each.

def rouge_n_f1(ref, hyp, n):
    def ngrams(s, n):
        toks = s.lower().split()
        return [tuple(toks[i:i+n]) for i in range(len(toks)-n+1)]
    rb = set(ngrams(ref, n))
    hb = ngrams(hyp, n)
    if not hb or not rb:
        return 0.0
    overlap = sum(1 for g in hb if g in rb)
    prec = overlap / len(hb)
    rec  = overlap / len(rb)
    return 2*prec*rec/(prec+rec) if prec+rec else 0.0

ref = 'the cat sat on the mat near the window'
ext = 'the cat sat on the mat'          # extractive copy
abs_ = 'a feline rested beside the window'  # abstractive

for name, hyp in [('extractive', ext), ('abstractive', abs_)]:
    r1 = rouge_n_f1(ref, hyp, 1)
    r2 = rouge_n_f1(ref, hyp, 2)
    print(f'{name}: ROUGE-1={r1:.4f}  ROUGE-2={r2:.4f}')
# extractive: ROUGE-1=0.8000  ROUGE-2=0.7692
# abstractive: ROUGE-1=0.2667  ROUGE-2=0.1538

Extractive ROUGE-1=0.8000, ROUGE-2=0.7692 — abstractive ROUGE-1=0.2667, ROUGE-2=0.1538

Why: The extractive copy shares 6 of 9 reference unigrams verbatim (precision=1.0, recall=6/9=0.667). The abstractive only shares 'the' and 'window' (2 tokens). ROUGE rewards surface n-gram overlap, penalizing paraphrase even when semantically faithful.

summaryshared unigramsROUGE-1 F1ROUGE-2 F1
extractive ('the cat sat on the mat')6/9 of ref0.80000.7692
abstractive ('a feline rested beside the window')2/9 of ref0.26670.1538
perfect copy of reference9/91.00001.0000

10. What each one costs: ROUGE-1 and ROUGE-2 trace — extractive vs…

Trade off

Comparison matrix

From ROUGE-1 and ROUGE-2 trace — extractive vs abstractive: every row here is a choice with a cost. Fill the ROUGE-2 F1 column, then say which row you would actually pick and what you give up for it.

summaryshared unigramsROUGE-1 F1ROUGE-2 F1
extractive ('the cat sat on the mat')6/9 of ref0.80000.7692
abstractive ('a feline rested beside the window')2/9 of ref0.26670.1538
perfect copy of reference9/91.00001.0000

11. Something is wrong here: high ROUGE means better summary

Anomaly

Predict first

A student writes this, and it looks reasonable:

ROUGE is the gold standard for summarization quality. A model achieving ROUGE-1=0.80 is always better than one scoring ROUGE-1=0.27.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: ROUGE is a surface n-gram metric.

ROUGE measures n-gram overlap. Use it alongside factual consistency checks (BertScore, QA-based) and human preference studies — especially for abstractive systems where paraphrase is correct but ROUGE penalizes it.

Why: ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.

12. Trap: high ROUGE means better summary

Trap

The trap

ROUGE is the gold standard for summarization quality. A model achieving ROUGE-1=0.80 is always better than one scoring ROUGE-1=0.27.

Rank models by ROUGE alone; discard any system with ROUGE below 0.5

Why: ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.

The fix

ROUGE measures n-gram overlap. Use it alongside factual consistency checks (BertScore, QA-based) and human preference studies — especially for abstractive systems where paraphrase is correct but ROUGE penalizes it.

Combine ROUGE with BertScore and hallucination detection; never rank abstractive against extractive by ROUGE alone

Why: The USAAIO exam expects you to know ROUGE's limits: it cannot distinguish a correct paraphrase from a wrong one, and an extractive oracle trivially maximizes it.

13. Break it on purpose: high ROUGE means better summary

Break the constraint

Discussion prompt

The rule this trap just fixed:

The USAAIO exam expects you to know ROUGE's limits: it cannot distinguish a correct paraphrase from a wrong one, and an extractive oracle trivially maximizes it.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

ROUGE is a surface n-gram metric. An extractive system that copies the first three sentences scores very high ROUGE but may be less informative than a concise abstractive summary. ROUGE also cannot detect hallucinations.

14. BART — denoising encoder-decoder pretraining

Section

Part 2 of 4

15. BART: bidirectional encoder + autoregressive decoder

Concept

BART (Lewis et al., 2020) is a seq2seq denoising autoencoder: the encoder reads a corrupted version of the text and the decoder learns to reconstruct the original — exactly like a denoising autoencoder, but at the token sequence level.

Architecture: standard Transformer encoder (like BERT, Lesson 90 — bidirectional attention, no causal mask) stacked with a standard Transformer decoder (like GPT, Lesson 91 — causal mask, cross-attention to encoder output). Same shape as the original 'Attention is All You Need' (Lesson 88).

componentattention typeanalogous modelrole
encoderbidirectional (no mask)BERT (L90)read corrupted input
decodercausal + cross-attentionGPT (L91)generate clean output
cross-attentiondecoder queries encoderstandard seq2seqcondition on source

16. BART's four corruption strategies

Concept

BART pretraining applies one or more noise functions to the source text. The decoder must reconstruct the original token sequence regardless of how it was corrupted.

  1. Text infilling: replace a contiguous span of tokens with a single [MASK] — the decoder must predict the number of missing tokens as well as their identity. Stronger than BERT's independent masking.
  2. Token deletion: delete random tokens with no mask placeholder — the decoder must infer which positions are missing
  3. Sentence permutation: shuffle the order of sentences in the document; decoder reconstructs the coherent order
  4. Document rotation: rotate the document to begin at a random token; decoder identifies and restores the true start
corruptionexample input to encoderdecoder target
text infillingThe [MASK] over the lazy dogThe quick brown fox jumped over the lazy dog
token deletionThe quick brown over the lazy dogThe quick brown fox jumped over the lazy dog
sentence permutationjumped over the lazy dog The quick brown foxThe quick brown fox jumped over the lazy dog
document rotationover the lazy dog The quick brown fox jumpedThe quick brown fox jumped over the lazy dog

17. By analogy: BART's four corruption strategies

Analogy

Discussion prompt

Explain BART's four corruption strategies by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

BART pretraining applies one or more noise functions to the source text. The decoder must reconstruct the original token sequence regardless of how it was corrupted.

18. PEGASUS: sentence-level masking for summarization

Concept

Pegasus (Zhang et al., 2020) replaces entire sentences with [MASK] tokens, selecting the sentences that are most important to the rest of the document — the same sentences a human would include in a summary.

Importance is scored by ROUGE-1 recall between the candidate sentence and the rest of the document. The top-k scoring sentences are masked (Gap Sentence Generation, GSG). The model learns to generate exactly the content humans use in extractive summaries — aligning pretraining directly with the downstream task.

modelpretraining objectiveunit maskedmask selection
BERT (L90)masked language modelindividual tokens (15%)random
BARTdenoising autoencoderspans or sentencesrandom / span-sampled
Pegasusgap sentence generation (GSG)whole sentenceshighest ROUGE-1 vs rest of doc

19. What has to be given first: TinyBART — encoder-decoder forward pass

Missing information

Discussion prompt

Build a minimal BART-shaped encoder-decoder with vocab_size=20, d_model=32, n_heads=2, 1 encoder layer, 1 decoder layer. Trace shapes through the forward pass.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Embedding tables: 20×32=640 each (×2 shared, plus position). TransformerEncoderLayer(32,2,ff=64): self-attn 32×32×4+biases + ff 32×64×2+biases ≈ 6,400. Decoder adds cross-attention block ≈ 8,480. Head: 32×20+20=660. Verified: 22,996 params total.

20. TinyBART — encoder-decoder forward pass

Worked example

Build a minimal BART-shaped encoder-decoder with vocab_size=20, d_model=32, n_heads=2, 1 encoder layer, 1 decoder layer. Trace shapes through the forward pass.

import torch, torch.nn as nn

class TinyBART(nn.Module):
    def __init__(self, V=20, d=32, h=2, max_len=10):
        super().__init__()
        self.emb = nn.Embedding(V, d)
        self.pos = nn.Embedding(max_len, d)
        enc_l = nn.TransformerEncoderLayer(d, h, dim_feedforward=64,
                                           batch_first=True, dropout=0.0)
        self.enc = nn.TransformerEncoder(enc_l, num_layers=1)
        dec_l = nn.TransformerDecoderLayer(d, h, dim_feedforward=64,
                                            batch_first=True, dropout=0.0)
        self.dec = nn.TransformerDecoder(dec_l, num_layers=1)
        self.head = nn.Linear(d, V)

    def forward(self, src, tgt):
        S, T = src.shape[1], tgt.shape[1]
        enc_out = self.enc(self.emb(src) + self.pos(torch.arange(S)))
        mask = nn.Transformer.generate_square_subsequent_mask(T)
        dec_out = self.dec(self.emb(tgt)+self.pos(torch.arange(T)),
                           enc_out, tgt_mask=mask, tgt_is_causal=True)
        return self.head(dec_out)

torch.manual_seed(42)
model = TinyBART()
print(sum(p.numel() for p in model.parameters()), 'params')  # 22996
src = torch.randint(0, 20, (2, 5))  # B=2, src_len=5
tgt = torch.randint(0, 20, (2, 4))  # B=2, tgt_len=4
out = model(src, tgt)
print(out.shape)                     # torch.Size([2, 4, 20])

22,996 parameters; output shape (2, 4, 20) — batch × target_len × vocab_size

Why: Embedding tables: 20×32=640 each (×2 shared, plus position). TransformerEncoderLayer(32,2,ff=64): self-attn 32×32×4+biases + ff 32×64×2+biases ≈ 6,400. Decoder adds cross-attention block ≈ 8,480. Head: 32×20+20=660. Verified: 22,996 params total.

tensorshapenote
src tokens(2, 5)B=2 sequences, src_len=5
tgt tokens(2, 4)B=2 sequences, tgt_len=4
enc_out(2, 5, 32)encoder reads all 5 src positions
dec_out(2, 4, 32)decoder: causal over tgt, cross-attends enc_out
logits (head)(2, 4, 20)one distribution over V=20 per tgt position

21. Fill in: note for TinyBART — encoder-decoder forward pass

Comparison

Comparison matrix

From TinyBART — encoder-decoder forward pass: refill the note column from what you know. The rest of the table is as it appeared.

tensorshapenote
src tokens(2, 5)B=2 sequences, src_len=5
tgt tokens(2, 4)B=2 sequences, tgt_len=4
enc_out(2, 5, 32)encoder reads all 5 src positions
dec_out(2, 4, 32)decoder: causal over tgt, cross-attends enc_out
logits (head)(2, 4, 20)one distribution over V=20 per tgt position

22. Perplexity — exponentiated cross-entropy

Section

Part 3 of 4

23. Perplexity: the formula and its meaning

Concept

\[ \text{PPL}(x_1,\ldots,x_N) = \exp\!\left(-\frac{1}{N}\sum_{i=1}^{N} \log P(x_i \mid x_{<i})\right) \]

PPL is the geometric mean inverse probability of each token given its context. Intuitively: a model with PPL=100 is as confused as a model that assigns equal probability to 100 choices at every step.

24. Guess the shape of the answer: Perplexity trace on a 3-token toy LM

Estimation

Predict first

Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.

Commit before you compute: what does Perplexity trace on a 3-token toy LM come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Both transitions are high-probability (the LM was designed that way).

25. Perplexity trace on a 3-token toy LM

Worked example

Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.

import torch, math

# Row = context token (0=the,1=cat,2=sat), Col = next token
logits = torch.tensor([
    [0.1, 2.0, 0.3],    # after 'the':  cat very likely
    [0.2, 0.1, 2.5],   # after 'cat':  sat very likely
    [1.5, 0.2, 0.1],   # after 'sat':  the most likely
])
log_p = torch.log_softmax(logits, dim=-1)
print(log_p.numpy().round(4))
# [[-2.1869 -0.2869 -1.9869]
#  [-2.4748 -2.5748 -0.1748]
#  [-0.4181 -1.7181 -1.8181]]

def ppl(seq, log_p):
    lps = [log_p[seq[i], seq[i+1]].item() for i in range(len(seq)-1)]
    return math.exp(-sum(lps)/len(lps))

good = [0, 1, 2]   # the -> cat -> sat
bad  = [2, 0, 1]   # sat -> the -> cat
print(f'PPL(the cat sat) = {ppl(good, log_p):.4f}')   # 1.2596
print(f'PPL(sat the cat) = {ppl(bad,  log_p):.4f}')   # 1.4226
print(f'PPL(uniform V=3) = 3')

log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596

Why: Both transitions are high-probability (the LM was designed that way). PPL=1.26 is well below the uniform baseline of 3, confirming the LM has learned the 'the cat sat' pattern. The bad sequence scores PPL=1.42 — still low because each transition happens to be moderately likely under the toy LM.

sequencelog P(step 1)log P(step 2)mean NLLPPL
the → cat → sat (good)-0.2869-0.17480.23081.2596
sat → the → cat (bad)-0.4181-1.71811.06812.9095
uniform over V=3——log 33.0000

26. Work backwards from the answer: Perplexity trace on a 3-token toy LM

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

log P(cat|the)=-0.2869, log P(sat|cat)=-0.1748; mean NLL=0.2308; PPL=exp(0.2308)=1.2596

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Vocab: {0:the, 1:cat, 2:sat}. A toy LM defines a 3×3 log-prob table where row=context token, col=next token. Compute PPL for the cat sat (good) vs sat the cat (bad). Predict before calculating.

27. When perplexity is a good (and bad) metric

Concept

28. Where does each piece belong: Lesson 108: Summarization, BART, Perplexity…

Sorting

Sort into buckets

These are the pieces of Lesson 108: Summarization, BART, Perplexity & Calibration, out of order. Put each one back under the part of the lesson it belongs to.

Extractive vs Abstractive — what ROUGE reveals
Extractive summarization: copy, do not generate; Abstractive summarization: generate new text; ROUGE-1 and ROUGE-2 trace — extractive vs abstractive
BART — denoising encoder-decoder pretraining
BART: bidirectional encoder + autoregressive decoder; BART's four corruption strategies; PEGASUS: sentence-level masking for summarization
Perplexity — exponentiated cross-entropy
Perplexity: the formula and its meaning; Perplexity trace on a 3-token toy LM; When perplexity is a good (and bad) metric
s1
Extractive vs Abstractive — what ROUGE reveals is where Lesson 108: Summarization, BART, Perplexity & Calibration puts Extractive summarization: copy, do not generate, Abstractive summarization: generate new text, ROUGE-1 and ROUGE-2 trace — extractive vs abstractive. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
BART — denoising encoder-decoder pretraining is where Lesson 108: Summarization, BART, Perplexity & Calibration puts BART: bidirectional encoder + autoregressive decoder, BART's four corruption strategies, PEGASUS: sentence-level masking for summarization. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Perplexity — exponentiated cross-entropy is where Lesson 108: Summarization, BART, Perplexity & Calibration puts Perplexity: the formula and its meaning, Perplexity trace on a 3-token toy LM, When perplexity is a good (and bad) metric. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

29. Something is wrong here: confusing PPL(text) with PPL(token sequence length N)

Anomaly

Predict first

A student writes this, and it looks reasonable:

PPL is computed over the raw token count N. A short 3-token sequence has higher PPL than a 100-token sequence from the same model, because the 100-token sequence averages more steps.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent).

PPL is already length-normalized by the 1/N factor. The meaningful comparison is between two models evaluated on the same test set, not between a short and a long sequence.

Why: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent). A 3-token sequence and a 100-token sequence produce comparable PPL values when drawn from the same distribution — length does not systematically inflate or deflate PPL.

30. Trap: confusing PPL(text) with PPL(token sequence length N)

Trap

The trap

PPL is computed over the raw token count N. A short 3-token sequence has higher PPL than a 100-token sequence from the same model, because the 100-token sequence averages more steps.

Compare PPL(3-token summary) vs PPL(100-token paragraph) to judge model quality

Why: This comparison is invalid. PPL is already normalized by N (the /N inside the exponent). A 3-token sequence and a 100-token sequence produce comparable PPL values when drawn from the same distribution — length does not systematically inflate or deflate PPL.

The fix

PPL is already length-normalized by the 1/N factor. The meaningful comparison is between two models evaluated on the same test set, not between a short and a long sequence.

Only compare PPL values when both models share the same vocabulary and the same test set

Why: Different vocabularies change the difficulty of each prediction step; different test sets change the distribution. PPL is a valid ranking within a fixed evaluation protocol.

31. Which of these survive contact with Lesson 108: Summarization, BART, Perplexity…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Classic extractive systems: TextRank (graph-based sentence scoring), BertSum (Lesson 90 BERT encoder + sentence-selection head).; BART pretraining applies one or more noise functions to the source text. The decoder must reconstruct the original token sequence regardless of how it was corrupted.; Build rules: print tensor shapes after encoder and decoder; verify causal mask shape is (T, T); log both loss and PPL; use torch.manual_seed(42) for reproducibility.
Breaks
ROUGE is the gold standard for summarization quality. A model achieving ROUGE-1=0.80 is always better than one scoring ROUGE-1=0.27.; Compare PPL(3-token summary) vs PPL(100-token paragraph) to judge model quality
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 108: Summarization, BART, Perplexity & Calibration puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

32. Calibration — does confidence match accuracy?

Section

Part 4 of 4

33. Calibration and Expected Calibration Error (ECE)

Concept

A model is calibrated when its predicted probability matches the empirical frequency of being correct. If the model says 70% on 1,000 examples, roughly 700 should be correct. Modern deep networks are typically overconfident — they output 95% when they are right only 70% of the time.

\[ \text{ECE} = \sum_{b=1}^{B} \frac{|\mathcal{B}_b|}{N} \left| \text{acc}(\mathcal{B}_b) - \text{conf}(\mathcal{B}_b) \right| \]

ECE bins predictions by confidence, then computes the weighted average absolute gap between accuracy and confidence within each bin. ECE = 0 = perfectly calibrated; ECE = 0.1 means confidence is off by 10 percentage points on average.

34. Guess the shape of the answer: ECE computation — calibrated vs overconfident

Estimation

Predict first

Simulate two models on n=200 examples with 5 confidence bins. Compute ECE for each. Predict which model has lower ECE before running.

Commit before you compute: what does ECE computation — calibrated vs overconfident come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: ECE calibrated=0.0296; ECE overconfident=0.1862 — a 6× gap

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. The calibrated model's confidence tracks empirical accuracy within each bin (small |acc-conf|).

35. ECE computation — calibrated vs overconfident

Worked example

Simulate two models on n=200 examples with 5 confidence bins. Compute ECE for each. Predict which model has lower ECE before running.

import numpy as np

def ece(conf, correct, n_bins=5):
    bins = np.linspace(0, 1, n_bins + 1)
    total = 0.0
    for b in range(n_bins):
        mask = (conf >= bins[b]) & (conf < bins[b+1])
        if mask.sum() == 0:
            continue
        bin_acc  = correct[mask].mean()
        bin_conf = conf[mask].mean()
        total += (mask.sum() / len(conf)) * abs(bin_acc - bin_conf)
    return total

np.random.seed(0)
n = 200
conf_cal  = np.random.uniform(0.5, 1.0, n)
corr_cal  = (np.random.rand(n) < conf_cal).astype(float)

conf_over = np.clip(conf_cal * 1.25, 0, 1.0)   # inflated confidence
corr_over = (np.random.rand(n) < conf_cal*0.75).astype(float)

print(f'ECE calibrated:    {ece(conf_cal,  corr_cal):.4f}')   # 0.0296
print(f'ECE overconfident: {ece(conf_over, corr_over):.4f}')  # 0.1862

ECE calibrated=0.0296; ECE overconfident=0.1862 — a 6× gap

Why: The calibrated model's confidence tracks empirical accuracy within each bin (small |acc-conf|). The overconfident model inflates confidence by 25% while accuracy stays the same, producing large within-bin gaps that sum to ECE≈0.19.

binnavg confavg acc|diff|
[0.4, 0.6)410.55050.46340.0871
[0.6, 0.8)770.70670.70130.0054
[0.8, 1.0)820.89090.91460.0237
ECE (weighted avg)200——0.0296

36. What each one costs: ECE computation — calibrated vs overconfident

Trade off

Comparison matrix

From ECE computation — calibrated vs overconfident: every row here is a choice with a cost. Fill the n column, then say which row you would actually pick and what you give up for it.

binnavg confavg acc|diff|
[0.4, 0.6)410.55050.46340.0871
[0.6, 0.8)770.70670.70130.0054
[0.8, 1.0)820.89090.91460.0237
ECE (weighted avg)200——0.0296

37. Without one step: The Lesson 108 recipe

Constraint

Discussion prompt

Run The Lesson 108 recipe with this step confiscated:

Perplexity: PPL = exp(-1/N ∑ log P(x_i|x_<i)). Lower = better. Uniform baseline = V. Only compare models with the same vocabulary and test set.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Extractive vs abstractive: extractive copies sentences → high ROUGE, no hallucination; abstractive generates new text → lower ROUGE, hallucination risk…
  2. BART pretraining: encoder reads corrupted text; decoder reconstructs original. Four corruptions: text infilling (span→[MASK]), token deletion (no…
  3. Pegasus GSG: mask the k sentences with highest ROUGE-1 recall vs the rest. Aligns pretraining with summarization — the model learns to generate exactly the…
  4. Perplexity: PPL = exp(-1/N ∑ log P(x_i|x_<i)). Lower = better. Uniform baseline = V. Only compare models with the same vocabulary and test set.
  5. Calibration / ECE: bin predictions by confidence; ECE = weighted |acc - conf| across bins. Overconfident models have high ECE. Fix with temperature scaling…
  6. When PPL misleads: PPL ignores factual accuracy, penalizes paraphrases, and cannot compare models across vocabulary sizes. Always pair with downstream task…

38. The Lesson 108 recipe

Pattern

  1. Extractive vs abstractive: extractive copies sentences → high ROUGE, no hallucination; abstractive generates new text → lower ROUGE, hallucination risk. Use ROUGE + factual-consistency together.
  2. BART pretraining: encoder reads corrupted text; decoder reconstructs original. Four corruptions: text infilling (span→[MASK]), token deletion (no placeholder), sentence permutation, document rotation.
  3. Pegasus GSG: mask the k sentences with highest ROUGE-1 recall vs the rest. Aligns pretraining with summarization — the model learns to generate exactly the most important sentences.
  4. Perplexity: PPL = exp(-1/N ∑ log P(x_i|x_<i)). Lower = better. Uniform baseline = V. Only compare models with the same vocabulary and test set.
  5. Calibration / ECE: bin predictions by confidence; ECE = weighted |acc - conf| across bins. Overconfident models have high ECE. Fix with temperature scaling (divide logits by T>1 before softmax).
  6. When PPL misleads: PPL ignores factual accuracy, penalizes paraphrases, and cannot compare models across vocabulary sizes. Always pair with downstream task metrics.

39. Where does it stop working: The Lesson 108 recipe

Edge cases

Discussion prompt

The Lesson 108 recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Extractive vs abstractive: extractive copies sentences → high ROUGE, no hallucination; abstractive generates new text → lower ROUGE, hallucination risk…
  2. BART pretraining: encoder reads corrupted text; decoder reconstructs original. Four corruptions: text infilling (span→[MASK]), token deletion (no…
  3. Pegasus GSG: mask the k sentences with highest ROUGE-1 recall vs the rest. Aligns pretraining with summarization — the model learns to generate exactly the…
  4. Perplexity: PPL = exp(-1/N ∑ log P(x_i|x_<i)). Lower = better. Uniform baseline = V. Only compare models with the same vocabulary and test set.
  5. Calibration / ECE: bin predictions by confidence; ECE = weighted |acc - conf| across bins. Overconfident models have high ECE. Fix with temperature scaling…
  6. When PPL misleads: PPL ignores factual accuracy, penalizes paraphrases, and cannot compare models across vocabulary sizes. Always pair with downstream task…

40. Rule out three: Check yourself — BART corruption

Elimination

Eliminate the wrong options

Which BART corruption strategy requires the decoder to also predict how many tokens were removed, not just which tokens?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Text infilling (span masking)
  • B. Token deletion
  • C. Sentence permutation
  • D. Document rotation

Survives elimination: A

Why: Text infilling replaces a contiguous span of k tokens with a single [MASK] token. The decoder sees one mask but must generate k tokens — it must learn to infer the span length from context. Token deletion drops tokens entirely but the decoder knows the original length from the target; sentence permutation and rotation do not remove tokens.

41. Check yourself — BART corruption

Check

Think about what makes each BART corruption strategy unique.

Check your understanding

Which BART corruption strategy requires the decoder to also predict how many tokens were removed, not just which tokens?

  • A. Text infilling (span masking) (correct)
  • B. Token deletion
  • C. Sentence permutation
  • D. Document rotation

Answer: A

Why: Text infilling replaces a contiguous span of k tokens with a single [MASK] token. The decoder sees one mask but must generate k tokens — it must learn to infer the span length from context. Token deletion drops tokens entirely but the decoder knows the original length from the target; sentence permutation and rotation do not remove tokens.

Why B tempts people
Token deletion drops tokens without a mask placeholder, but the target sequence length is fixed — the decoder recovers positions but does not need to predict span size from a single mask.
Why C tempts people
Sentence permutation shuffles order without changing token count. The decoder reconstructs order, not missing spans.
Why D tempts people
Document rotation circularly shifts the start token. All tokens are present; the decoder identifies and restores the original start position.

42. Answer it before you see the options: Check yourself — perplexity…

Prediction

Predict first

A language model has vocabulary size V=10,000. On a test set, it achieves PPL=100. Which statement is correct?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: The model is effectively choosing uniformly among 100 tokens at each step — still far from a perfect predictor but much better than random.

Why: PPL=100 means the model's average per-step uncertainty equals a uniform distribution over 100 tokens. The uniform baseline for this model is PPL=V=10,000, so the model is 100× better than random (10,000/100=100×). PPL is not token accuracy — a model can have PPL<10 while still predicting incorrectly on many tokens because probabilities don't have to be 1.0.

43. Check yourself — perplexity interpretation

Check

Apply the PPL formula and the uniform baseline.

Check your understanding

A language model has vocabulary size V=10,000. On a test set, it achieves PPL=100. Which statement is correct?

  • A. The model is 100× better than a uniform random predictor over the vocabulary.
  • B. The model is effectively choosing uniformly among 100 tokens at each step — still far from a perfect predictor but much better than random. (correct)
  • C. PPL=100 means the model gets 100 tokens correct out of every 10,000, i.e., 1% accuracy.
  • D. PPL=100 is meaningless unless the test set has at least 100 sentences.

Answer: B

Why: PPL=100 means the model's average per-step uncertainty equals a uniform distribution over 100 tokens. The uniform baseline for this model is PPL=V=10,000, so the model is 100× better than random (10,000/100=100×). PPL is not token accuracy — a model can have PPL<10 while still predicting incorrectly on many tokens because probabilities don't have to be 1.0.

Why A tempts people
This is backwards. PPL=100 means the model is as confident as a uniform predictor over 100 choices. The uniform baseline (PPL=V=10,000) is 100× worse, so the model is 100× better than random — but choice A reverses the direction of the comparison.
Why C tempts people
PPL is not accuracy. A model assigns probability distributions; PPL measures the average negative log-probability of the correct token, not the fraction of correctly predicted tokens.
Why D tempts people
PPL is well-defined regardless of the number of sentences. It is computed over individual token positions and averaged. Sentence count does not set a validity threshold.

44. Rule out three: Check yourself — calibration ECE

Elimination

Eliminate the wrong options

A model outputs 0.95 confidence on 500 examples. 400 of those are correct (80% accuracy). Which fixes this overconfidence best?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Re-train with a smaller learning rate to reduce loss.
  • B. Apply temperature scaling: divide logits by T>1 before softmax to flatten the distribution.
  • C. Add dropout during inference to reduce confidence.
  • D. Increase model capacity so it fits the training set better.

Survives elimination: B

Why: Temperature scaling is the standard post-hoc calibration method. Dividing logits by T>1 softens the softmax distribution (reduces confidence) without changing the argmax prediction or retraining. It is fit on a held-out validation set by minimizing NLL, and is the most practical fix for overconfidence identified in Guo et al. 2017.

45. Check yourself — calibration ECE

Check

Reason about the reliability diagram before selecting.

Check your understanding

A model outputs 0.95 confidence on 500 examples. 400 of those are correct (80% accuracy). Which fixes this overconfidence best?

  • A. Re-train with a smaller learning rate to reduce loss.
  • B. Apply temperature scaling: divide logits by T>1 before softmax to flatten the distribution. (correct)
  • C. Add dropout during inference to reduce confidence.
  • D. Increase model capacity so it fits the training set better.

Answer: B

Why: Temperature scaling is the standard post-hoc calibration method. Dividing logits by T>1 softens the softmax distribution (reduces confidence) without changing the argmax prediction or retraining. It is fit on a held-out validation set by minimizing NLL, and is the most practical fix for overconfidence identified in Guo et al. 2017.

Why A tempts people
A smaller learning rate changes optimization dynamics but does not directly address the mismatch between confidence and accuracy already present in a trained model. Calibration is a post-training property.
Why C tempts people
MC Dropout during inference adds stochasticity and can estimate uncertainty, but it is not a calibration method per se — it can actually produce inconsistent confidence estimates and is more complex than temperature scaling.
Why D tempts people
Increasing capacity typically makes overconfidence worse, not better. Larger models tend to be more overconfident (Guo et al. 2017 showed that ResNets became less calibrated as depth increased).

46. Your turn: TinyBART copy-task fine-tuning

Section

Project

47. Project brief

Concept

Build and fine-tune TinyBART — a minimal encoder-decoder transformer — on a synthetic copy task. Track PPL as training progresses to see the loss curve a real summarization fine-tune would produce.

#milestonekey tool
1Implement TinyBART (encoder + decoder + causal mask), trace shapes through forward passnn.TransformerEncoder, nn.TransformerDecoder, generate_square_subsequent_mask
2Write a perplexity function: given model logits and labels, return exp(cross-entropy)log_softmax, gather, mean, exp
3Train 50 epochs on copy task, log PPL every 10 epochs — target PPL < 2.0 at epoch 50Adam lr=1e-3, CrossEntropyLoss

Build rules: print tensor shapes after encoder and decoder; verify causal mask shape is (T, T); log both loss and PPL; use torch.manual_seed(42) for reproducibility.

48. Break it if you can: Project brief

Counterexample

Discussion prompt

Build and fine-tune TinyBART — a minimal encoder-decoder transformer — on a synthetic copy task. Track PPL as training progresses to see the loss curve a real summarization fine-tune would produce.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Build rules: print tensor shapes after encoder and decoder; verify causal mask shape is (T, T); log both loss and PPL; use torch.manual_seed(42) for reproducibility.

49. Milestone 2 — perplexity from logits

Worked example

Your turn: given model output logits of shape (B, T, V) and labels (B, T), compute per-sequence PPL. Predict the PPL at random initialization for V=10.

Hint: use log_softmax(logits, dim=-1), then gather to select log-prob at each label index, then mean(dim=1) for per-sequence NLL, then exp.

import torch, math

def ppl_from_logits(logits, labels):
    # logits: (B, T, V), labels: (B, T)
    log_p = torch.log_softmax(logits, dim=-1)          # (B, T, V)
    lp_label = log_p.gather(2, labels.unsqueeze(-1))   # (B, T, 1)
    lp_label = lp_label.squeeze(-1)                    # (B, T)
    nll = -lp_label.mean(dim=1)                        # (B,)
    return nll.exp()                                   # (B,) PPL per seq

torch.manual_seed(42)
logits = torch.randn(2, 4, 10)   # random-init, V=10
labels = torch.randint(0, 10, (2, 4))
ppl = ppl_from_logits(logits, labels)
print(ppl)  # tensor([24.2912, 19.5234])  (random init ~V=10 range)
tensorshapeoperation
logits(2, 4, 10)raw model output
log_p(2, 4, 10)log_softmax over vocab dim
lp_label(2, 4)gather log-prob at label index
nll per seq(2,)mean over T positions
PPL(2,)exp(nll): [24.29, 19.52]

50. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at Milestone 2 — perplexity from logits. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(2, 4, 10)
logits; log_p
(2, 4)
lp_label
(2,)
nll per seq; PPL
g1
shape is "(2, 4, 10)" for logits, log_p — that is what the table on "Milestone 2 — perplexity from logits" records, and it is the single property separating this group from the rest.
g2
shape is "(2, 4)" for lp_label — that is what the table on "Milestone 2 — perplexity from logits" records, and it is the single property separating this group from the rest.
g3
shape is "(2,)" for nll per seq, PPL — that is what the table on "Milestone 2 — perplexity from logits" records, and it is the single property separating this group from the rest.

51. Milestone 3 — copy-task training curve

Worked example

Your turn: train TinyBART for 50 epochs on 32 random sequences (V=10, seq_len=5). Predict: will PPL converge below 2.0? Estimate at how many epochs.

Hint: the copy task is easy (label = src); Adam at lr=1e-3 converges faster than the transformer default 3e-4. Teacher-forcing (feeding ground truth as decoder input) keeps the loss stable.

import torch, torch.nn as nn, math

torch.manual_seed(42)
VOCAB = 10
model = TinyBART(V=VOCAB, d=32, h=2, max_len=8)  # from Milestone 1
opt   = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()

src_data = torch.randint(0, VOCAB, (32, 5))
labels   = src_data                      # copy task: predict src

for epoch in range(51):
    opt.zero_grad()
    logits = model(src_data, src_data)   # teacher-force tgt = src
    loss   = loss_fn(logits.reshape(-1, VOCAB), labels.reshape(-1))
    loss.backward(); opt.step()
    if epoch % 10 == 0:
        ppl = math.exp(loss.item())
        print(f'epoch {epoch:3d}  loss={loss.item():.4f}  PPL={ppl:.2f}')
# epoch   0  loss=2.3997  PPL=11.02
# epoch  10  loss=1.7515  PPL=5.76
# epoch  20  loss=1.2176  PPL=3.38
# epoch  30  loss=0.8125  PPL=2.25
# epoch  40  loss=0.5339  PPL=1.71
# epoch  50  loss=0.3535  PPL=1.42
epochlossPPLinterpretation
02.399711.02random init; above V=10 average because labels are concentrated
101.75155.76model starts copying short patterns
201.21763.38approaching meaningful compression
300.81252.25task nearly learned
400.53391.71PPL < 2 — well below random
500.35351.42strong copy; further epochs approach PPL=1

52. The full program

Concept

import torch, torch.nn as nn, math

class TinyBART(nn.Module):
    def __init__(self, V=10, d=32, h=2, max_len=8):
        super().__init__()
        self.emb = nn.Embedding(V, d)
        self.pos = nn.Embedding(max_len, d)
        self.enc = nn.TransformerEncoder(
            nn.TransformerEncoderLayer(d, h, dim_feedforward=64,
                                       batch_first=True, dropout=0.0), 1)
        self.dec = nn.TransformerDecoder(
            nn.TransformerDecoderLayer(d, h, dim_feedforward=64,
                                       batch_first=True, dropout=0.0), 1)
        self.head = nn.Linear(d, V)

    def forward(self, src, tgt):
        S, T = src.shape[1], tgt.shape[1]
        enc_out = self.enc(self.emb(src)+self.pos(torch.arange(S)))
        mask = nn.Transformer.generate_square_subsequent_mask(T)
        dec_out = self.dec(self.emb(tgt)+self.pos(torch.arange(T)),
                           enc_out, tgt_mask=mask, tgt_is_causal=True)
        return self.head(dec_out)

torch.manual_seed(42)
VOCAB = 10
model   = TinyBART(V=VOCAB)
opt     = torch.optim.Adam(model.parameters(), lr=1e-3)
loss_fn = nn.CrossEntropyLoss()
src     = torch.randint(0, VOCAB, (32, 5))
for ep in range(51):
    opt.zero_grad()
    loss = loss_fn(model(src, src).reshape(-1, VOCAB), src.reshape(-1))
    loss.backward(); opt.step()
    if ep % 10 == 0:
        print(f'ep {ep:3d}  PPL={math.exp(loss.item()):.2f}')
# ep   0  PPL=11.02
# ep  10  PPL=5.76
# ep  20  PPL=3.38
# ep  30  PPL=2.25
# ep  40  PPL=1.71
# ep  50  PPL=1.42
design choicevaluerationale
vocab size V10small → copy task learned quickly
d_model32minimal but valid transformer dim
n_heads2d_k = 32/2 = 16 per head
learning rate1e-3copy task is simple; faster lr than summarization
epochs50PPL crosses 2.0 at epoch ~35
params22,996 (V=20 variant)verified with named_parameters()

53. Fill in: rationale for The full program

Comparison

Comparison matrix

From The full program: refill the rationale column from what you know. The rest of the table is as it appeared.

design choicevaluerationale
vocab size V10small → copy task learned quickly
d_model32minimal but valid transformer dim
n_heads2d_k = 32/2 = 16 per head
learning rate1e-3copy task is simple; faster lr than summarization
epochs50PPL crosses 2.0 at epoch ~35
params22,996 (V=20 variant)verified with named_parameters()

54. Show it off

Concept

Out loud, slides closed: (1) explain why extractive summarization always outscores abstractive on ROUGE-1 against the source; (2) walk through all four BART corruption strategies and state which one is unique because the decoder must infer span length; (3) compute PPL by hand given two log-prob values.

Homework (from the lesson plan): fine-tune BART-base for abstractive summarization on CNN/DailyMail (use Hugging Face datasets); implement a perplexity function that works for any autoregressive language model; compare extractive (first-3-sentences baseline) vs abstractive (BART) using ROUGE and at least one factual-consistency metric.

55. Connect it up: Lesson 108: Summarization, BART, Perplexity & Calibration

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Extractive vs Abstractive — what ROUGE reveals · BART — denoising encoder-decoder pretraining · Perplexity — exponentiated cross-entropy · Calibration — does confidence match accuracy? · Your turn: TinyBART copy-task fine-tuning. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

56. What you can do now

Recap

conceptthe one thing to remember
extractiveselects sentences — ROUGE inflated, no hallucination
abstractivegenerates new text — lower ROUGE, hallucination risk
BARTencoder reads corrupted, decoder reconstructs clean
Pegasusmask highest-ROUGE-1 sentences to align pretraining with summarization
perplexityPPL=exp(-1/N ∑ log P); lower=better; uniform baseline=V
calibration / ECEECE = weighted |acc-conf|; overconfident → high ECE; fix: temperature scaling

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 108 — Abstractive Summarization, BART, Perplexity, Calibration — Barron · USAAIO Round 2 Preparation, 2026
  2. Lewis et al. 'BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension' (ACL 2020) — arXiv:1910.13461
  3. Zhang et al. 'PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization' (ICML 2020) — arXiv:1912.08777
  4. Guo et al. 'On Calibration of Modern Neural Networks' (ICML 2017) — arXiv:1706.04599
  5. TinyBART (22,996 params), ROUGE-1/2 trace, 3-token LM PPL, and ECE calibration table verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108