USAAIO Lesson 86, from Phase 3. It covers the encoder-decoder architecture, teacher forcing against autoregressive inference, beam search with a k=2 trace, and cross-entropy loss with label smoothing at eps=0.1, then trains a toy seq2seq Transformer from scratch in PyTorch. All the attention scores, causal masks, label-smoothing probabilities, cumulative beam-search log-probabilities, and training losses were verified with torch 2.7.1. The lesson runs to 30 slides.
Subject: Machine Learning · 61 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 86 · Phase 3
Architecture, teacher forcing, autoregressive inference, beam search, and label smoothing — the complete seq2seq pipeline from encoder to generated output.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 86: Transformer Encoder-Decoder & Seq2Seq: without looking back, what was the main idea of Transformer Decoder Block, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
the three sublayers of the transformer decoder block — causal masked self-attention (upper-triangular -inf mask), cross-attention (Q from decoder, K/V from encoder), and FFN — plus autoregressive generation.
Section
Part 1 of 4
Concept
The encoder-decoder Transformer (Vaswani 2017) splits the problem: the encoder converts the source sequence to a context representation; the decoder attends to that context while generating the target token by token.
| half | input | output | attention types |
|---|---|---|---|
| Encoder | source tokens + PE | memory (same shape) | self-attention only |
| Decoder | target tokens + PE | hidden states | self-attn + cross-attn to memory |
| Projection | decoder hidden | logits over vocab | linear + softmax |
Encoder self-attention is bidirectional (every token attends to every other). Decoder self-attention uses a causal mask so position t cannot see positions t+1, t+2, … (Lesson 83).
Comparison
Comparison matrix
From Two halves of seq2seq: refill the input column from what you know. The rest of the table is as it appeared.
| half | input | output | attention types |
|---|---|---|---|
| Encoder | source tokens + PE | memory (same shape) | self-attention only |
| Decoder | target tokens + PE | hidden states | self-attn + cross-attn to memory |
| Projection | decoder hidden | logits over vocab | linear + softmax |
Concept
Both encoder and decoder add sinusoidal positional encodings (Lesson 84) to their embeddings before the first layer. With d_model=16 and max_len=10 (verified):
\[ \text{PE}(\text{pos}, 2i) = \sin\!\left(\frac{\text{pos}}{10000^{2i/d}}\right),\quad \text{PE}(\text{pos}, 2i{+}1) = \cos\!\left(\frac{\text{pos}}{10000^{2i/d}}\right) \]
| pos | PE[pos, 0] | PE[pos, 1] | PE[pos, 2] | PE[pos, 3] |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| 1 | 0.8415 | 0.5403 | 0.3110 | 0.9504 |
| 5 | -0.9589 | 0.2837 | 0.9999 | -0.0103 |
Trade off
Comparison matrix
From Positional encoding in the encoder-decoder: every row here is a choice with a cost. Fill the PE[pos, 2] column, then say which row you would actually pick and what you give up for it.
| pos | PE[pos, 0] | PE[pos, 1] | PE[pos, 2] | PE[pos, 3] |
|---|---|---|---|---|
| 0 | 0.0000 | 1.0000 | 0.0000 | 1.0000 |
| 1 | 0.8415 | 0.5403 | 0.3110 | 0.9504 |
| 5 | -0.9589 | 0.2837 | 0.9999 | -0.0103 |
Estimation
Predict first
Trace single-head attention on 4 source tokens with d_k=8. Q, K, V are random (seed 0). This is the same scaled dot-product from Lesson 83 — used in both self-attention and cross-attention.
Commit before you compute: what does Scaled dot-product attention (single head, 4 tokens) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Scores scale down by sqrt(d_k)=2.828 to prevent softmax saturation (Lesson 83)
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Without the sqrt(d_k) divisor, dot products grow with d_k, pushing softmax into near-zero-gradient regions and slowing training.
Worked example
Trace single-head attention on 4 source tokens with d_k=8. Q, K, V are random (seed 0). This is the same scaled dot-product from Lesson 83 — used in both self-attention and cross-attention.
import torch, torch.nn.functional as F
torch.manual_seed(0)
seq_len, d_k = 4, 8
Q = torch.randn(seq_len, d_k)
K = torch.randn(seq_len, d_k)
V = torch.randn(seq_len, d_k)
scores = Q @ K.T / (d_k ** 0.5) # (4,4)
attn_w = F.softmax(scores, dim=-1)
out = attn_w @ V
print(attn_w.numpy().round(3))
print(out[0].numpy().round(4))Scores scale down by sqrt(d_k)=2.828 to prevent softmax saturation (Lesson 83)
Why: Without the sqrt(d_k) divisor, dot products grow with d_k, pushing softmax into near-zero-gradient regions and slowing training.
| query | attn to tok0 | attn to tok1 | attn to tok2 | attn to tok3 |
|---|---|---|---|---|
| tok0 | 0.121 | 0.741 | 0.106 | 0.032 |
| tok1 | 0.258 | 0.265 | 0.171 | 0.307 |
| tok2 | 0.336 | 0.017 | 0.187 | 0.460 |
| tok3 | 0.203 | 0.623 | 0.102 | 0.071 |
Pattern
Step through it
Step through Scaled dot-product attention (single head, 4 tokens) one row at a time. What is driving the change, and what would the row after the last one be?
Concept
A TransformerEncoderLayer leaves the shape unchanged: (batch, src_len, d_model) → (batch, src_len, d_model). The decoder takes both the target embedding AND the encoder memory.
| tensor | shape | meaning |
|---|---|---|
| encoder input (src) | (2, 6, 16) | batch=2, src_len=6, d_model=16 |
| encoder output (memory) | (2, 6, 16) | same shape — each src token has a context vector |
| decoder input (tgt) | (2, 5, 16) | batch=2, tgt_len=5, d_model=16 |
| decoder output | (2, 5, 16) | same shape — each tgt position's hidden state |
| projection logits | (2, 5, V) | one distribution over vocab per tgt position |
The encoder runs once per source. At inference the memory is cached; the decoder runs once per generated token (autoregressive). This asymmetry is why inference is slow compared to encoding.
Discrimination
Sort into buckets
Sort these by shape, from memory, without looking back at Encoder and decoder shapes in PyTorch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Section
Part 2 of 4
Concept
During training, instead of feeding the model's own (potentially wrong) predicted tokens back into the decoder, we feed the ground-truth target tokens shifted by one. This is teacher forcing.
| step t | decoder input token | decoder target label |
|---|---|---|
| 1 | [BOS] | y_1 |
| 2 | y_1 (ground truth) | y_2 |
| 3 | y_2 (ground truth) | y_3 |
| 4 | y_3 (ground truth) | [EOS] |
Teacher forcing lets the whole target sequence be processed in one parallel forward pass (with the causal mask blocking future tokens). Without it, early wrong predictions would corrupt every subsequent step.
Counterexample
Discussion prompt
During training, instead of feeding the model's own (potentially wrong) predicted tokens back into the decoder, we feed the ground-truth target tokens shifted by one. This is teacher forcing.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Missing information
Discussion prompt
For a target sequence [BOS=10, 3, 3, EOS=11], the decoder input is target[:-1] and the labels are target[1:]. The causal mask ensures position t sees only positions 0..t.
What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.
Hint: Anything you would have to invent to get started is a thing the problem must supply.
Answer:
Adding -inf before softmax zeroes out those attention weights — the decoder cannot peek at future ground-truth tokens, so the training objective matches inference.
Worked example
For a target sequence [BOS=10, 3, 3, EOS=11], the decoder input is target[:-1] and the labels are target[1:]. The causal mask ensures position t sees only positions 0..t.
import torch, torch.nn as nn
# target: [BOS=10, 3, 3, EOS=11]
target = torch.tensor([[10, 3, 3, 11]]) # (1, 4)
tgt_in = target[:, :-1] # (1, 3): [BOS, 3, 3]
tgt_out = target[:, 1:] # (1, 3): [3, 3, EOS]
# causal mask for tgt_len=3
tgt_len = tgt_in.size(1)
causal = torch.triu(torch.ones(tgt_len, tgt_len)*float('-inf'), diagonal=1)
print('tgt_in: ', tgt_in.tolist())
print('tgt_out:', tgt_out.tolist())
print('causal mask:\n', causal.numpy())causal[i,j] = -inf for j>i: position 0 sees only position 0; position 2 sees 0,1,2
Why: Adding -inf before softmax zeroes out those attention weights — the decoder cannot peek at future ground-truth tokens, so the training objective matches inference.
| decoder pos | sees positions | predicts label |
|---|---|---|
| 0 (BOS) | 0 only | 3 |
| 1 (3) | 0, 1 | 3 |
| 2 (3) | 0, 1, 2 | EOS=11 |
Pattern
Step through it
Step through Teacher forcing in the training loop one row at a time. What is driving the change, and what would the row after the last one be?
Concept
The training loss is cross-entropy over the decoder's output logits vs the target token IDs. Label smoothing (ε=0.1, Lesson 86 spec) regularizes by softening the one-hot target distribution.
\[ \tilde{y}_k = \begin{cases} 1 - \varepsilon + \varepsilon/V & k = \text{true class} \\ \varepsilon/V & \text{otherwise} \end{cases} \]
| ε | V | true-class prob | other-class prob | CE (top logit=3.0) |
|---|---|---|---|---|
| 0.0 (none) | 5 | 1.0000 | 0.0000 | 0.2876 |
| 0.1 | 5 | 0.9200 | 0.0200 | 0.4916 |
| interpretation | — | never 100% confident | nonzero everywhere | higher loss → prevents overconfidence |
Analogy
Discussion prompt
Explain Cross-entropy loss with label smoothing by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
The training loss is cross-entropy over the decoder's output logits vs the target token IDs. Label smoothing (ε=0.1, Lesson 86 spec) regularizes by softening the one-hot target distribution.
Anomaly
Predict first
A student writes this, and it looks reasonable:
At inference, feed the ground-truth token at each step — the same approach as training. This keeps the decoder input correct and avoids compounding errors.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: At inference the target is unknown — that is the whole point.
At inference, feed the model's own previous output (autoregressive). Teacher forcing is training only.
Why: At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.
Trap
At inference, feed the ground-truth token at each step — the same approach as training. This keeps the decoder input correct and avoids compounding errors.
Feed target[t] (ground truth) at step t even during inference
Why: At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.
At inference, feed the model's own previous output (autoregressive). Teacher forcing is training only.
Step t: input = [BOS, y_hat_1, ..., y_hat_{t-1}]; take argmax of last position logit to get y_hat_t
Why: Autoregressive decoding extends the hypothesis one token at a time using the model's own predictions. This matches the deployment setting — no oracle tokens available.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Autoregressive decoding extends the hypothesis one token at a time using the model's own predictions. This matches the deployment setting — no oracle tokens available.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.
Section
Part 3 of 4
Concept
At inference, start with [BOS]. At each step, run the full decoder forward pass on the current hypothesis and take the argmax of the last position's logits to extend by one token. Repeat until [EOS] or max length.
| step | decoder input (length t) | new token (greedy argmax) |
|---|---|---|
| 1 | [BOS] | y_hat_1 |
| 2 | [BOS, y_hat_1] | y_hat_2 |
| 3 | [BOS, y_hat_1, y_hat_2] | y_hat_3 or EOS |
| 4 | ... | keep going or stop at EOS |
Each step re-encodes the decoder input from scratch (naive) or uses a KV cache to avoid recomputing earlier positions (Lesson 87 topic). Cost without cache: O(T²) total — T steps, each with O(T) attention.
Comparison
Comparison matrix
From Autoregressive inference: refill the new token (greedy argmax) column from what you know. The rest of the table is as it appeared.
| step | decoder input (length t) | new token (greedy argmax) |
|---|---|---|
| 1 | [BOS] | y_hat_1 |
| 2 | [BOS, y_hat_1] | y_hat_2 |
| 3 | [BOS, y_hat_1, y_hat_2] | y_hat_3 or EOS |
| 4 | ... | keep going or stop at EOS |
Estimation
Predict first
The toy seq2seq (digit → [digit, digit, EOS]) trains with teacher forcing for 300 epochs. After training, we decode digit=3 autoregressively. Verified loss trajectory below.
Commit before you compute: what does Autoregressive decode of a trained toy seq2seq come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Generated [BOS=10, 3, 3, EOS=11] — model learned to echo digit twice then stop
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Autoregressive generation: each step extends the hypothesis by 1.
Worked example
The toy seq2seq (digit → [digit, digit, EOS]) trains with teacher forcing for 300 epochs. After training, we decode digit=3 autoregressively. Verified loss trajectory below.
import torch, torch.nn as nn
torch.manual_seed(99)
# ToySeq2Seq: src vocab 12, tgt vocab 12, d=16, nhead=2
# BOS=10, EOS=11 — training code in milestone section
# After 300 epochs, greedy decode:
BOS, EOS = 10, 11
# (model already trained, shown below)
model2.eval()
with torch.no_grad():
src_test = torch.tensor([[3]]) # digit 3
generated = [BOS]
for _ in range(4):
tgt_in = torch.tensor([generated])
tl = len(generated)
cm = torch.triu(torch.ones(tl,tl)*float('-inf'), diagonal=1)
logits = model2(src_test, tgt_in, tgt_mask=cm)
next_t = logits[0,-1].argmax().item()
if next_t == EOS: break
generated.append(next_t)
print(generated) # [10, 3, 3]Generated [BOS=10, 3, 3, EOS=11] — model learned to echo digit twice then stop
Why: Autoregressive generation: each step extends the hypothesis by 1. The causal mask grows by one row/col per step. The model's own output at step t becomes the input at step t+1.
| epoch | training loss |
|---|---|
| 0 | 2.7717 |
| 50 | 0.8280 |
| 100 | 0.4379 |
| 200 | 0.3123 |
| 299 | 0.0609 |
Pattern
Step through it
Step through Autoregressive decode of a trained toy seq2seq one row at a time. What is driving the change, and what would the row after the last one be?
Concept
Greedy decoding takes the single highest-probability token at each step. It is fast but locally optimal — a strong early choice can crowd out a better overall sequence.
Beam search maintains the top-k hypotheses (beams) simultaneously. At each step it expands all k beams by every vocab token, scores the k×V candidates by cumulative log-probability, and retains the top k.
\[ \text{score}(y_{1:t}) = \sum_{i=1}^{t} \log P(y_i \mid y_{<i},\, x) \]
Sorting
Sort into buckets
These are the pieces of Lesson 86: Transformer Encoder-Decoder & Seq2Seq, out of order. Put each one back under the part of the lesson it belongs to.
Estimation
Predict first
Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.
Commit before you compute: what does Beam search trace (k=2, vocab=5, 2 steps) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Beam search scores sequences, not individual tokens.
Worked example
Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.
import torch
step1 = torch.log(torch.tensor([0.05,0.10,0.15,0.45,0.25]))
topk1 = torch.topk(step1, k=2)
print('Step-1 top-2:', topk1.indices.tolist(), topk1.values.numpy().round(3))
# Step 2: from beam1 (tok=3) and beam2 (tok=4)
b1_lp = torch.log(torch.tensor([0.02,0.03,0.10,0.15,0.70]))
b2_lp = torch.log(torch.tensor([0.05,0.05,0.20,0.40,0.30]))
b1_top2 = torch.topk(b1_lp, k=2)
b2_top2 = torch.topk(b2_lp, k=2)
print('Beam1 step2:', b1_top2.indices.tolist(), b1_top2.values.numpy().round(3))
print('Beam2 step2:', b2_top2.indices.tolist(), b2_top2.values.numpy().round(3))Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates
Why: Beam search scores sequences, not individual tokens. A lower step-1 probability can be rescued by a very high step-2 probability — greedy would miss this.
| sequence | step-1 log-p | step-2 log-p | cum log-p | kept? |
|---|---|---|---|---|
| [3, 4] (beam1 → tok4) | -0.799 | -0.357 | -1.155 | yes (best) |
| [4, 3] (beam2 → tok3) | -1.386 | -0.916 | -2.303 | yes (2nd) |
| [4, 4] (beam2 → tok4) | -1.386 | -1.204 | -2.590 | dropped |
| [3, 3] (beam1 → tok3) | -0.799 | -1.897 | -2.696 | dropped |
Reverse engineer
Discussion prompt
Work backwards. The example finished here:
Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates
What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.
Hint: Every quantity in the result had to enter somewhere. Account for each one.
Answer:
Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Beam search (k>1) is strictly better than greedy (k=1) — it explores more of the search space, so its output is always higher quality.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Beam search is heuristic, not exact.
Beam search is better than greedy in practice, but it is not exact and has failure modes — especially without length normalization.
Why: Beam search is heuristic, not exact. It can still miss the global optimum. Large k also penalizes short sequences: beams tend to keep longer hypotheses alive because cumulative log-probs keep decreasing. Without length normalization, beam search degenerates toward very long outputs.
Trap
Beam search (k>1) is strictly better than greedy (k=1) — it explores more of the search space, so its output is always higher quality.
Prefer beam search with large k to maximize output quality
Why: Beam search is heuristic, not exact. It can still miss the global optimum. Large k also penalizes short sequences: beams tend to keep longer hypotheses alive because cumulative log-probs keep decreasing. Without length normalization, beam search degenerates toward very long outputs.
Beam search is better than greedy in practice, but it is not exact and has failure modes — especially without length normalization.
Use beam search with length normalization: divide cum log-prob by sequence length^alpha (alpha~0.6-0.7)
Why: Raw cumulative log-probs always decrease with length, so longer sequences score lower regardless of quality. Length normalization rescales scores so short and long sequences compete fairly.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.Constraint
Discussion prompt
Run The encoder-decoder seq2seq recipe with this step confiscated:
Inference (greedy): start [BOS]; repeat: forward pass on current hypothesis, argmax last logit, append token; stop at [EOS] or max_len
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear…target[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfidentPattern
TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear projection → logitstarget[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfidentEdge cases
Discussion prompt
The encoder-decoder seq2seq recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear…target[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfidentElimination
Eliminate the wrong options
During training with teacher forcing, what is fed as the decoder input at step t?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Teacher forcing feeds the true target token (shifted by one) at each decoder step. This allows the full target sequence to be processed in one parallel forward pass with a causal mask. The decoder input is target[:-1]; the labels are target[1:].
Check
Work out the mechanism before clicking.
Check your understanding
During training with teacher forcing, what is fed as the decoder input at step t?
Answer: A
Why: Teacher forcing feeds the true target token (shifted by one) at each decoder step. This allows the full target sequence to be processed in one parallel forward pass with a causal mask. The decoder input is target[:-1]; the labels are target[1:].
Prediction
Predict first
In the beam search trace (k=2), after step 2, which sequence is dropped first?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: [3, 3] with cum log-prob −2.696
Why: [3, 3] has cumulative log-prob −2.696, the lowest of the four candidates. After sorting descending by cum log-prob and keeping the top-2, [3, 3] and [4, 4] are dropped; [3, 4] (−1.155) and [4, 3] (−2.303) are retained.
Check
Use the verified trace numbers.
Check your understanding
In the beam search trace (k=2), after step 2, which sequence is dropped first?
Answer: D
Why: [3, 3] has cumulative log-prob −2.696, the lowest of the four candidates. After sorting descending by cum log-prob and keeping the top-2, [3, 3] and [4, 4] are dropped; [3, 4] (−1.155) and [4, 3] (−2.303) are retained.
Elimination
Eliminate the wrong options
Label smoothing ε=0.1 is applied to a problem with vocabulary size V=10. What is the smoothed probability assigned to the correct class?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The true-class smoothed probability is 1 − ε + ε/V = 1 − 0.1 + 0.1/10 = 0.9 + 0.01 = 0.91. Each other class gets ε/V = 0.01, and 0.91 + 9×0.01 = 1.0.
Check
Apply the formula.
Check your understanding
Label smoothing ε=0.1 is applied to a problem with vocabulary size V=10. What is the smoothed probability assigned to the correct class?
Answer: A
Why: The true-class smoothed probability is 1 − ε + ε/V = 1 − 0.1 + 0.1/10 = 0.9 + 0.01 = 0.91. Each other class gets ε/V = 0.01, and 0.91 + 9×0.01 = 1.0.
Section
Project
Concept
Build a full encoder-decoder Transformer for a toy seq2seq task: given a source digit 0-9, generate [digit, digit, EOS]. Three milestones: training loop with teacher forcing → autoregressive inference → beam search (k=2).
| # | milestone | key tool |
|---|---|---|
| 1 | Teacher-forcing training loop (300 epochs, loss → 0.06) | nn.Transformer, CrossEntropyLoss |
| 2 | Autoregressive greedy decode (src=3 → [3, 3, EOS]) | causal mask growing per step |
| 3 | Beam search k=2 on step-1 probs | torch.topk, cumulative log-probs |
Build rules: use batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.
Counterexample
Discussion prompt
Build rules: use batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Worked example
Your turn: set up ToySeq2Seq (src_vocab=12, tgt_vocab=12, d=16, nhead=2, ff=32). Train 300 epochs with Adam lr=5e-3. Predict the final loss.
Hint: decoder input = tgts_in (BOS prepended), labels = tgts_out (EOS appended). Causal mask shape = (3,3). loss = CrossEntropyLoss(logits.reshape(-1,VTGT), labels.reshape(-1)).
import torch, torch.nn as nn
torch.manual_seed(99)
BOS, EOS, VTGT = 10, 11, 12
srcs = torch.arange(10).unsqueeze(1) # (10,1)
tgts_in = torch.stack([torch.tensor([BOS,i,i]) for i in range(10)])
tgts_out= torch.stack([torch.tensor([i,i,EOS]) for i in range(10)])
causal = torch.triu(torch.ones(3,3)*float('-inf'), diagonal=1)
model2 = ToySeq2Seq() # defined above
opt = torch.optim.Adam(model2.parameters(), lr=5e-3)
for ep in range(300):
opt.zero_grad()
logits = model2(srcs, tgts_in, tgt_mask=causal)
loss = nn.CrossEntropyLoss()(logits.reshape(-1,VTGT), tgts_out.reshape(-1))
loss.backward(); opt.step()
print(f'final loss: {loss.item():.4f}')| epoch | loss |
|---|---|
| 0 | 2.7717 |
| 50 | 0.8280 |
| 100 | 0.4379 |
| 200 | 0.3123 |
| 299 | 0.0609 |
Pattern
Step through it
Step through Milestone 1 — teacher-forcing training loop one row at a time. What is driving the change, and what would the row after the last one be?
Worked example
Your turn: after training, decode src=3 step by step (no teacher forcing). Predict the output token at step 1.
Hint: start with generated=[BOS]; at each step extend the causal mask by 1 row+col; append logits[0,-1].argmax().item(); stop when EOS.
model2.eval()
with torch.no_grad():
src_t = torch.tensor([[3]])
gen = [BOS]
for _ in range(4):
tl = len(gen)
cm = torch.triu(torch.ones(tl,tl)*float('-inf'), diagonal=1)
lg = model2(src_t, torch.tensor([gen]), tgt_mask=cm)
nxt = lg[0,-1].argmax().item()
gen.append(nxt)
if nxt == EOS: break
print(gen) # [10, 3, 3, 11]| step | decoder input | next token predicted |
|---|---|---|
| 1 | [BOS=10] | 3 |
| 2 | [BOS, 3] | 3 |
| 3 | [BOS, 3, 3] | EOS=11 → stop |
Trade off
Comparison matrix
From Milestone 2 — autoregressive greedy decode: every row here is a choice with a cost. Fill the next token predicted column, then say which row you would actually pick and what you give up for it.
| step | decoder input | next token predicted |
|---|---|---|
| 1 | [BOS=10] | 3 |
| 2 | [BOS, 3] | 3 |
| 3 | [BOS, 3, 3] | EOS=11 → stop |
Worked example
Your turn: implement beam search for 2 steps on the given step-1 probs. Predict which sequences survive after step 2.
Hint: torch.topk(log_probs, k=2) at each step; accumulate log-probs by summing; after expanding all k beams, take global top-k.
import torch
step1 = torch.log(torch.tensor([0.05,0.10,0.15,0.45,0.25]))
tk1 = torch.topk(step1, k=2) # ids=[3,4], lps=[-0.799,-1.386]
b1 = torch.log(torch.tensor([0.02,0.03,0.10,0.15,0.70]))
b2 = torch.log(torch.tensor([0.05,0.05,0.20,0.40,0.30]))
cands = []
for i,src_lp in enumerate([tk1.values[0], tk1.values[1]]):
for tok,lp in zip(torch.topk(([b1,b2][i]),k=2).indices,
torch.topk(([b1,b2][i]),k=2).values):
cands.append((tok.item(), src_lp.item()+lp.item()))
cands.sort(key=lambda x:-x[1])
print(cands[:2]) # kept beams| sequence | cum log-prob | kept? |
|---|---|---|
| [3, 4] | -1.155 | yes |
| [4, 3] | -2.303 | yes |
| [4, 4] | -2.590 | dropped |
| [3, 3] | -2.696 | dropped |
Discrimination
Sort into buckets
Sort these by kept?, from memory, without looking back at Milestone 3 — beam search k=2. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.
Concept
import torch, torch.nn as nn
torch.manual_seed(99)
BOS,EOS,V = 10,11,12
class ToySeq2Seq(nn.Module):
def __init__(self,vsrc=V,vtgt=V,d=16,nh=2,ff=32,nl=1):
super().__init__()
self.se=nn.Embedding(vsrc,d); self.te=nn.Embedding(vtgt,d)
self.tr=nn.Transformer(d,nh,nl,nl,ff,batch_first=True)
self.pr=nn.Linear(d,vtgt)
def forward(self,s,t,tgt_mask=None):
return self.pr(self.tr(self.se(s),self.te(t),tgt_mask=tgt_mask))
srcs=torch.arange(10).unsqueeze(1)
ti=torch.stack([torch.tensor([BOS,i,i])for i in range(10)])
to=torch.stack([torch.tensor([i,i,EOS])for i in range(10)])
cm=torch.triu(torch.ones(3,3)*float('-inf'),diagonal=1)
m=ToySeq2Seq(); opt=torch.optim.Adam(m.parameters(),lr=5e-3)
for _ in range(300):
opt.zero_grad()
nn.CrossEntropyLoss()(m(srcs,ti,tgt_mask=cm).reshape(-1,V),to.reshape(-1)).backward()
opt.step()
m.eval()
with torch.no_grad():
g=[BOS]
for _ in range(4):
tl=len(g); cm2=torch.triu(torch.ones(tl,tl)*float('-inf'),diagonal=1)
nxt=m(torch.tensor([[3]]),torch.tensor([g]),tgt_mask=cm2)[0,-1].argmax().item()
g.append(nxt)
if nxt==EOS: break
print(g) # [10, 3, 3, 11]| component | choice | why |
|---|---|---|
| encoder | nn.TransformerEncoder (bidirectional) | full context over source — no causal mask |
| decoder | nn.TransformerDecoder + causal mask | can't peek at future target tokens |
| training | teacher forcing + CrossEntropyLoss | parallel over sequence, fast convergence |
| inference | autoregressive (own output → next input) | no ground truth available at deploy time |
| loss | label_smoothing=0.1 | prevents overconfident logits, improves generalization |
Final training loss 0.0609; greedy decode of digit 3 → [BOS, 3, 3, EOS]. Same five-step training loop as Lesson 40 — only the model (Transformer) and the shifted target trick change.
Comparison
Comparison matrix
From The full program: refill the choice column from what you know. The rest of the table is as it appeared.
| component | choice | why |
|---|---|---|
| encoder | nn.TransformerEncoder (bidirectional) | full context over source — no causal mask |
| decoder | nn.TransformerDecoder + causal mask | can't peek at future target tokens |
| training | teacher forcing + CrossEntropyLoss | parallel over sequence, fast convergence |
| inference | autoregressive (own output → next input) | no ground truth available at deploy time |
| loss | label_smoothing=0.1 | prevents overconfident logits, improves generalization |
Concept
Out loud, slides closed: (1) explain what the encoder produces and how the decoder uses it via cross-attention, (2) walk through one teacher-forcing training step showing tgt_in vs tgt_out, and (3) trace beam search for k=2 from [3, 4] at step 2.
Stretch (homework): add label_smoothing=0.1 to the training loss and compare final loss; implement length normalization for the beam search; extend the task to actual number-to-words translation (e.g., 42 → 'forty two'). Next up: Lesson 87 — KV cache and efficient inference.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Encoder-Decoder architecture · Teacher forcing & training · Autoregressive inference & beam search · Your turn: seq2seq from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
target[:-1] to decoder, predict target[1:], apply causal mask| concept | the one thing to remember |
|---|---|
| encoder-decoder | encoder → memory (cached); decoder cross-attends every step |
| teacher forcing | training only — feed ground truth, not model output |
| autoregressive | inference only — grow hypothesis one token at a time |
| beam search | keep top-k cumulative log-probs; normalize by length |
| label smoothing | 1−ε+ε/V for true class; prevents overconfident logits |
Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.