Lesson 86: Transformer Encoder-Decoder & Seq2Seq

USAAIO Lesson 86, from Phase 3. It covers the encoder-decoder architecture, teacher forcing against autoregressive inference, beam search with a k=2 trace, and cross-entropy loss with label smoothing at eps=0.1, then trains a toy seq2seq Transformer from scratch in PyTorch. All the attention scores, causal masks, label-smoothing probabilities, cumulative beam-search log-probabilities, and training losses were verified with torch 2.7.1. The lesson runs to 30 slides.

Subject: Machine Learning · 61 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. Transformer Encoder-Decoder & Seq2Seq

Title

USAAIO · Lesson 86 · Phase 3

Architecture, teacher forcing, autoregressive inference, beam search, and label smoothing — the complete seq2seq pipeline from encoder to generated output.

2. By the end of this lesson you can

Objectives

  1. Describe the encoder-decoder split: encoder embeds source, decoder generates target token-by-token
  2. Implement teacher forcing in the training loop and explain why it speeds convergence
  3. Trace autoregressive inference step-by-step (causal mask, one token per forward pass)
  4. Execute a beam search (k=2) by hand — expand, score, and prune hypotheses
  5. Apply label smoothing (ε=0.1) and state the smoothed target distribution

3. What survived from Transformer Decoder Block?

Warm-up

Discussion prompt

Before we open Lesson 86: Transformer Encoder-Decoder & Seq2Seq: without looking back, what was the main idea of Transformer Decoder Block, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

the three sublayers of the transformer decoder block — causal masked self-attention (upper-triangular -inf mask), cross-attention (Q from decoder, K/V from encoder), and FFN — plus autoregressive generation.

4. Encoder-Decoder architecture

Section

Part 1 of 4

5. Two halves of seq2seq

Concept

The encoder-decoder Transformer (Vaswani 2017) splits the problem: the encoder converts the source sequence to a context representation; the decoder attends to that context while generating the target token by token.

halfinputoutputattention types
Encodersource tokens + PEmemory (same shape)self-attention only
Decodertarget tokens + PEhidden statesself-attn + cross-attn to memory
Projectiondecoder hiddenlogits over vocablinear + softmax

Encoder self-attention is bidirectional (every token attends to every other). Decoder self-attention uses a causal mask so position t cannot see positions t+1, t+2, … (Lesson 83).

6. Fill in: input for Two halves of seq2seq

Comparison

Comparison matrix

From Two halves of seq2seq: refill the input column from what you know. The rest of the table is as it appeared.

halfinputoutputattention types
Encodersource tokens + PEmemory (same shape)self-attention only
Decodertarget tokens + PEhidden statesself-attn + cross-attn to memory
Projectiondecoder hiddenlogits over vocablinear + softmax

7. Positional encoding in the encoder-decoder

Concept

Both encoder and decoder add sinusoidal positional encodings (Lesson 84) to their embeddings before the first layer. With d_model=16 and max_len=10 (verified):

\[ \text{PE}(\text{pos}, 2i) = \sin\!\left(\frac{\text{pos}}{10000^{2i/d}}\right),\quad \text{PE}(\text{pos}, 2i{+}1) = \cos\!\left(\frac{\text{pos}}{10000^{2i/d}}\right) \]

posPE[pos, 0]PE[pos, 1]PE[pos, 2]PE[pos, 3]
00.00001.00000.00001.0000
10.84150.54030.31100.9504
5-0.95890.28370.9999-0.0103

8. What each one costs: Positional encoding in the encoder-decoder

Trade off

Comparison matrix

From Positional encoding in the encoder-decoder: every row here is a choice with a cost. Fill the PE[pos, 2] column, then say which row you would actually pick and what you give up for it.

posPE[pos, 0]PE[pos, 1]PE[pos, 2]PE[pos, 3]
00.00001.00000.00001.0000
10.84150.54030.31100.9504
5-0.95890.28370.9999-0.0103

9. Guess the shape of the answer: Scaled dot-product attention (single head, 4…

Estimation

Predict first

Trace single-head attention on 4 source tokens with d_k=8. Q, K, V are random (seed 0). This is the same scaled dot-product from Lesson 83 — used in both self-attention and cross-attention.

Commit before you compute: what does Scaled dot-product attention (single head, 4 tokens) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Scores scale down by sqrt(d_k)=2.828 to prevent softmax saturation (Lesson 83)

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Without the sqrt(d_k) divisor, dot products grow with d_k, pushing softmax into near-zero-gradient regions and slowing training.

10. Scaled dot-product attention (single head, 4 tokens)

Worked example

Trace single-head attention on 4 source tokens with d_k=8. Q, K, V are random (seed 0). This is the same scaled dot-product from Lesson 83 — used in both self-attention and cross-attention.

import torch, torch.nn.functional as F
torch.manual_seed(0)
seq_len, d_k = 4, 8
Q = torch.randn(seq_len, d_k)
K = torch.randn(seq_len, d_k)
V = torch.randn(seq_len, d_k)
scores = Q @ K.T / (d_k ** 0.5)    # (4,4)
attn_w = F.softmax(scores, dim=-1)
out = attn_w @ V
print(attn_w.numpy().round(3))
print(out[0].numpy().round(4))

Scores scale down by sqrt(d_k)=2.828 to prevent softmax saturation (Lesson 83)

Why: Without the sqrt(d_k) divisor, dot products grow with d_k, pushing softmax into near-zero-gradient regions and slowing training.

queryattn to tok0attn to tok1attn to tok2attn to tok3
tok00.1210.7410.1060.032
tok10.2580.2650.1710.307
tok20.3360.0170.1870.460
tok30.2030.6230.1020.071

11. Watch it run: Scaled dot-product attention (single head, 4 tokens)

Pattern

Step through it

Step through Scaled dot-product attention (single head, 4 tokens) one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: query is tok0
  2. Step 2: query is tok1
  3. Step 3: query is tok2
  4. Step 4: query is tok3

12. Encoder and decoder shapes in PyTorch

Concept

A TransformerEncoderLayer leaves the shape unchanged: (batch, src_len, d_model) → (batch, src_len, d_model). The decoder takes both the target embedding AND the encoder memory.

tensorshapemeaning
encoder input (src)(2, 6, 16)batch=2, src_len=6, d_model=16
encoder output (memory)(2, 6, 16)same shape — each src token has a context vector
decoder input (tgt)(2, 5, 16)batch=2, tgt_len=5, d_model=16
decoder output(2, 5, 16)same shape — each tgt position's hidden state
projection logits(2, 5, V)one distribution over vocab per tgt position

The encoder runs once per source. At inference the memory is cached; the decoder runs once per generated token (autoregressive). This asymmetry is why inference is slow compared to encoding.

13. Which is which, by shape

Discrimination

Sort into buckets

Sort these by shape, from memory, without looking back at Encoder and decoder shapes in PyTorch. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

(2, 6, 16)
encoder input (src); encoder output (memory)
(2, 5, 16)
decoder input (tgt); decoder output
(2, 5, V)
projection logits
g1
shape is "(2, 6, 16)" for encoder input (src), encoder output (memory) — that is what the table on "Encoder and decoder shapes in PyTorch" records, and it is the single property separating this group from the rest.
g2
shape is "(2, 5, 16)" for decoder input (tgt), decoder output — that is what the table on "Encoder and decoder shapes in PyTorch" records, and it is the single property separating this group from the rest.
g3
shape is "(2, 5, V)" for projection logits — that is what the table on "Encoder and decoder shapes in PyTorch" records, and it is the single property separating this group from the rest.

14. Teacher forcing & training

Section

Part 2 of 4

15. Teacher forcing during training

Concept

During training, instead of feeding the model's own (potentially wrong) predicted tokens back into the decoder, we feed the ground-truth target tokens shifted by one. This is teacher forcing.

step tdecoder input tokendecoder target label
1[BOS]y_1
2y_1 (ground truth)y_2
3y_2 (ground truth)y_3
4y_3 (ground truth)[EOS]

Teacher forcing lets the whole target sequence be processed in one parallel forward pass (with the causal mask blocking future tokens). Without it, early wrong predictions would corrupt every subsequent step.

16. Break it if you can: Teacher forcing during training

Counterexample

Discussion prompt

During training, instead of feeding the model's own (potentially wrong) predicted tokens back into the decoder, we feed the ground-truth target tokens shifted by one. This is teacher forcing.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

17. What has to be given first: Teacher forcing in the training loop

Missing information

Discussion prompt

For a target sequence [BOS=10, 3, 3, EOS=11], the decoder input is target[:-1] and the labels are target[1:]. The causal mask ensures position t sees only positions 0..t.

What do you need to know — or decide — before the first line can be written? List everything the problem has to hand you.

Hint: Anything you would have to invent to get started is a thing the problem must supply.

Answer:

Adding -inf before softmax zeroes out those attention weights — the decoder cannot peek at future ground-truth tokens, so the training objective matches inference.

18. Teacher forcing in the training loop

Worked example

For a target sequence [BOS=10, 3, 3, EOS=11], the decoder input is target[:-1] and the labels are target[1:]. The causal mask ensures position t sees only positions 0..t.

import torch, torch.nn as nn
# target: [BOS=10, 3, 3, EOS=11]
target = torch.tensor([[10, 3, 3, 11]])  # (1, 4)
tgt_in  = target[:, :-1]                # (1, 3): [BOS, 3, 3]
tgt_out = target[:, 1:]                 # (1, 3): [3, 3, EOS]
# causal mask for tgt_len=3
tgt_len = tgt_in.size(1)
causal = torch.triu(torch.ones(tgt_len, tgt_len)*float('-inf'), diagonal=1)
print('tgt_in: ', tgt_in.tolist())
print('tgt_out:', tgt_out.tolist())
print('causal mask:\n', causal.numpy())

causal[i,j] = -inf for j>i: position 0 sees only position 0; position 2 sees 0,1,2

Why: Adding -inf before softmax zeroes out those attention weights — the decoder cannot peek at future ground-truth tokens, so the training objective matches inference.

decoder possees positionspredicts label
0 (BOS)0 only3
1 (3)0, 13
2 (3)0, 1, 2EOS=11

19. Watch it run: Teacher forcing in the training loop

Pattern

Step through it

Step through Teacher forcing in the training loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: decoder pos is 0 (BOS)
  2. Step 2: decoder pos is 1 (3)
  3. Step 3: decoder pos is 2 (3)

20. Cross-entropy loss with label smoothing

Concept

The training loss is cross-entropy over the decoder's output logits vs the target token IDs. Label smoothing (ε=0.1, Lesson 86 spec) regularizes by softening the one-hot target distribution.

\[ \tilde{y}_k = \begin{cases} 1 - \varepsilon + \varepsilon/V & k = \text{true class} \\ \varepsilon/V & \text{otherwise} \end{cases} \]

εVtrue-class probother-class probCE (top logit=3.0)
0.0 (none)51.00000.00000.2876
0.150.92000.02000.4916
interpretation—never 100% confidentnonzero everywherehigher loss → prevents overconfidence

21. By analogy: Cross-entropy loss with label smoothing

Analogy

Discussion prompt

Explain Cross-entropy loss with label smoothing by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

The training loss is cross-entropy over the decoder's output logits vs the target token IDs. Label smoothing (ε=0.1, Lesson 86 spec) regularizes by softening the one-hot target distribution.

22. Something is wrong here: using teacher forcing at inference time

Anomaly

Predict first

A student writes this, and it looks reasonable:

At inference, feed the ground-truth token at each step — the same approach as training. This keeps the decoder input correct and avoids compounding errors.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: At inference the target is unknown — that is the whole point.

At inference, feed the model's own previous output (autoregressive). Teacher forcing is training only.

Why: At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.

23. Trap: using teacher forcing at inference time

Trap

The trap

At inference, feed the ground-truth token at each step — the same approach as training. This keeps the decoder input correct and avoids compounding errors.

Feed target[t] (ground truth) at step t even during inference

Why: At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.

The fix

At inference, feed the model's own previous output (autoregressive). Teacher forcing is training only.

Step t: input = [BOS, y_hat_1, ..., y_hat_{t-1}]; take argmax of last position logit to get y_hat_t

Why: Autoregressive decoding extends the hypothesis one token at a time using the model's own predictions. This matches the deployment setting — no oracle tokens available.

24. Break it on purpose: using teacher forcing at inference time

Break the constraint

Discussion prompt

The rule this trap just fixed:

Autoregressive decoding extends the hypothesis one token at a time using the model's own predictions. This matches the deployment setting — no oracle tokens available.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

At inference the target is unknown — that is the whole point. There are no ground-truth tokens to feed. Using teacher forcing at inference time is data leakage: the model would be conditioning on the answer to predict the answer.

25. Autoregressive inference & beam search

Section

Part 3 of 4

26. Autoregressive inference

Concept

At inference, start with [BOS]. At each step, run the full decoder forward pass on the current hypothesis and take the argmax of the last position's logits to extend by one token. Repeat until [EOS] or max length.

stepdecoder input (length t)new token (greedy argmax)
1[BOS]y_hat_1
2[BOS, y_hat_1]y_hat_2
3[BOS, y_hat_1, y_hat_2]y_hat_3 or EOS
4...keep going or stop at EOS

Each step re-encodes the decoder input from scratch (naive) or uses a KV cache to avoid recomputing earlier positions (Lesson 87 topic). Cost without cache: O(T²) total — T steps, each with O(T) attention.

27. Fill in: new token (greedy argmax) for Autoregressive inference

Comparison

Comparison matrix

From Autoregressive inference: refill the new token (greedy argmax) column from what you know. The rest of the table is as it appeared.

stepdecoder input (length t)new token (greedy argmax)
1[BOS]y_hat_1
2[BOS, y_hat_1]y_hat_2
3[BOS, y_hat_1, y_hat_2]y_hat_3 or EOS
4...keep going or stop at EOS

28. Guess the shape of the answer: Autoregressive decode of a trained toy…

Estimation

Predict first

The toy seq2seq (digit → [digit, digit, EOS]) trains with teacher forcing for 300 epochs. After training, we decode digit=3 autoregressively. Verified loss trajectory below.

Commit before you compute: what does Autoregressive decode of a trained toy seq2seq come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Generated [BOS=10, 3, 3, EOS=11] — model learned to echo digit twice then stop

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Autoregressive generation: each step extends the hypothesis by 1.

29. Autoregressive decode of a trained toy seq2seq

Worked example

The toy seq2seq (digit → [digit, digit, EOS]) trains with teacher forcing for 300 epochs. After training, we decode digit=3 autoregressively. Verified loss trajectory below.

import torch, torch.nn as nn
torch.manual_seed(99)
# ToySeq2Seq: src vocab 12, tgt vocab 12, d=16, nhead=2
# BOS=10, EOS=11 — training code in milestone section
# After 300 epochs, greedy decode:
BOS, EOS = 10, 11
# (model already trained, shown below)
model2.eval()
with torch.no_grad():
    src_test = torch.tensor([[3]])    # digit 3
    generated = [BOS]
    for _ in range(4):
        tgt_in = torch.tensor([generated])
        tl = len(generated)
        cm = torch.triu(torch.ones(tl,tl)*float('-inf'), diagonal=1)
        logits = model2(src_test, tgt_in, tgt_mask=cm)
        next_t = logits[0,-1].argmax().item()
        if next_t == EOS: break
        generated.append(next_t)
print(generated)   # [10, 3, 3]

Generated [BOS=10, 3, 3, EOS=11] — model learned to echo digit twice then stop

Why: Autoregressive generation: each step extends the hypothesis by 1. The causal mask grows by one row/col per step. The model's own output at step t becomes the input at step t+1.

epochtraining loss
02.7717
500.8280
1000.4379
2000.3123
2990.0609

30. Watch it run: Autoregressive decode of a trained toy seq2seq

Pattern

Step through it

Step through Autoregressive decode of a trained toy seq2seq one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 50
  3. Step 3: epoch is 100
  4. Step 4: epoch is 200
  5. Step 5: epoch is 299

31. Greedy vs beam search

Concept

Greedy decoding takes the single highest-probability token at each step. It is fast but locally optimal — a strong early choice can crowd out a better overall sequence.

Beam search maintains the top-k hypotheses (beams) simultaneously. At each step it expands all k beams by every vocab token, scores the k×V candidates by cumulative log-probability, and retains the top k.

\[ \text{score}(y_{1:t}) = \sum_{i=1}^{t} \log P(y_i \mid y_{<i},\, x) \]

32. Where does each piece belong: Lesson 86: Transformer Encoder-Decoder &…

Sorting

Sort into buckets

These are the pieces of Lesson 86: Transformer Encoder-Decoder & Seq2Seq, out of order. Put each one back under the part of the lesson it belongs to.

Encoder-Decoder architecture
Two halves of seq2seq; Positional encoding in the encoder-decoder; Scaled dot-product attention (single head, 4 tokens)
Teacher forcing & training
Teacher forcing during training; Teacher forcing in the training loop; Cross-entropy loss with label smoothing
Autoregressive inference & beam search
Autoregressive inference; Autoregressive decode of a trained toy seq2seq; Greedy vs beam search
s1
Encoder-Decoder architecture is where Lesson 86: Transformer Encoder-Decoder & Seq2Seq puts Two halves of seq2seq, Positional encoding in the encoder-decoder, Scaled dot-product attention (single head, 4 tokens). Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Teacher forcing & training is where Lesson 86: Transformer Encoder-Decoder & Seq2Seq puts Teacher forcing during training, Teacher forcing in the training loop, Cross-entropy loss with label smoothing. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Autoregressive inference & beam search is where Lesson 86: Transformer Encoder-Decoder & Seq2Seq puts Autoregressive inference, Autoregressive decode of a trained toy seq2seq, Greedy vs beam search. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

33. Guess the shape of the answer: Beam search trace (k=2, vocab=5, 2 steps)

Estimation

Predict first

Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.

Commit before you compute: what does Beam search trace (k=2, vocab=5, 2 steps) come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.

Correct: Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates

Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Beam search scores sequences, not individual tokens.

34. Beam search trace (k=2, vocab=5, 2 steps)

Worked example

Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.

import torch
step1 = torch.log(torch.tensor([0.05,0.10,0.15,0.45,0.25]))
topk1 = torch.topk(step1, k=2)
print('Step-1 top-2:', topk1.indices.tolist(), topk1.values.numpy().round(3))
# Step 2: from beam1 (tok=3) and beam2 (tok=4)
b1_lp = torch.log(torch.tensor([0.02,0.03,0.10,0.15,0.70]))
b2_lp = torch.log(torch.tensor([0.05,0.05,0.20,0.40,0.30]))
b1_top2 = torch.topk(b1_lp, k=2)
b2_top2 = torch.topk(b2_lp, k=2)
print('Beam1 step2:', b1_top2.indices.tolist(), b1_top2.values.numpy().round(3))
print('Beam2 step2:', b2_top2.indices.tolist(), b2_top2.values.numpy().round(3))

Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates

Why: Beam search scores sequences, not individual tokens. A lower step-1 probability can be rescued by a very high step-2 probability — greedy would miss this.

sequencestep-1 log-pstep-2 log-pcum log-pkept?
[3, 4] (beam1 → tok4)-0.799-0.357-1.155yes (best)
[4, 3] (beam2 → tok3)-1.386-0.916-2.303yes (2nd)
[4, 4] (beam2 → tok4)-1.386-1.204-2.590dropped
[3, 3] (beam1 → tok3)-0.799-1.897-2.696dropped

35. Work backwards from the answer: Beam search trace (k=2, vocab=5, 2 steps)

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

Combine beam1_score + step2_logp for each extension; keep global top-2 across all 4 candidates

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

Vocab = {0,1,2,3,4}. Step 1 probs: [0.05, 0.10, 0.15, 0.45, 0.25]. Keep top-2 beams. Verified with torch.topk.

36. Something is wrong here: beam search always beats greedy on quality

Anomaly

Predict first

A student writes this, and it looks reasonable:

Beam search (k>1) is strictly better than greedy (k=1) — it explores more of the search space, so its output is always higher quality.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Beam search is heuristic, not exact.

Beam search is better than greedy in practice, but it is not exact and has failure modes — especially without length normalization.

Why: Beam search is heuristic, not exact. It can still miss the global optimum. Large k also penalizes short sequences: beams tend to keep longer hypotheses alive because cumulative log-probs keep decreasing. Without length normalization, beam search degenerates toward very long outputs.

37. Trap: beam search always beats greedy on quality

Trap

The trap

Beam search (k>1) is strictly better than greedy (k=1) — it explores more of the search space, so its output is always higher quality.

Prefer beam search with large k to maximize output quality

Why: Beam search is heuristic, not exact. It can still miss the global optimum. Large k also penalizes short sequences: beams tend to keep longer hypotheses alive because cumulative log-probs keep decreasing. Without length normalization, beam search degenerates toward very long outputs.

The fix

Beam search is better than greedy in practice, but it is not exact and has failure modes — especially without length normalization.

Use beam search with length normalization: divide cum log-prob by sequence length^alpha (alpha~0.6-0.7)

Why: Raw cumulative log-probs always decrease with length, so longer sequences score lower regardless of quality. Length normalization rescales scores so short and long sequences compete fairly.

38. Which of these survive contact with Lesson 86: Transformer Encoder-Decoder &…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Both encoder and decoder add sinusoidal positional encodings (Lesson 84) to their embeddings before the first layer. With d_model=16 and max_len=10 (verified):; Build rules: use batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.
Breaks
At inference, feed the ground-truth token at each step — the same approach as training. This keeps the decoder input correct and avoids compounding errors.; Beam search (k>1) is strictly better than greedy (k=1) — it explores more of the search space, so its output is always higher quality.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 86: Transformer Encoder-Decoder & Seq2Seq puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

39. Without one step: The encoder-decoder seq2seq recipe

Constraint

Discussion prompt

Run The encoder-decoder seq2seq recipe with this step confiscated:

Inference (greedy): start [BOS]; repeat: forward pass on current hypothesis, argmax last logit, append token; stop at [EOS] or max_len

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Architecture: TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear…
  2. Training: teacher forcing — feed target[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)
  3. Inference (greedy): start [BOS]; repeat: forward pass on current hypothesis, argmax last logit, append token; stop at [EOS] or max_len
  4. Inference (beam search k): keep top-k hypotheses; expand all k × V, score by cumulative log-prob, retain top-k; normalize by length
  5. Loss check: label smoothing ε=0.1 shifts target probs to 1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfident

40. The encoder-decoder seq2seq recipe

Pattern

  1. Architecture: TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear projection → logits
  2. Training: teacher forcing — feed target[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)
  3. Inference (greedy): start [BOS]; repeat: forward pass on current hypothesis, argmax last logit, append token; stop at [EOS] or max_len
  4. Inference (beam search k): keep top-k hypotheses; expand all k × V, score by cumulative log-prob, retain top-k; normalize by length
  5. Loss check: label smoothing ε=0.1 shifts target probs to 1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfident

41. Where does it stop working: The encoder-decoder seq2seq recipe

Edge cases

Discussion prompt

The encoder-decoder seq2seq recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Architecture: TransformerEncoder (bidirectional self-attn) → memory; TransformerDecoder (causal self-attn + cross-attn) → hidden states → linear…
  2. Training: teacher forcing — feed target[:-1] as decoder input, predict target[1:]; apply causal mask; CrossEntropyLoss(label_smoothing=0.1)
  3. Inference (greedy): start [BOS]; repeat: forward pass on current hypothesis, argmax last logit, append token; stop at [EOS] or max_len
  4. Inference (beam search k): keep top-k hypotheses; expand all k × V, score by cumulative log-prob, retain top-k; normalize by length
  5. Loss check: label smoothing ε=0.1 shifts target probs to 1−ε+ε/V (true) and ε/V (others); prevents the model from becoming overconfident

42. Rule out three: Check yourself — teacher forcing

Elimination

Eliminate the wrong options

During training with teacher forcing, what is fed as the decoder input at step t?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The ground-truth target token at position t−1
  • B. The model's own predicted token from step t−1
  • C. The encoder output (memory) at position t
  • D. A zero vector (the decoder starts cold at each step)

Survives elimination: A

Why: Teacher forcing feeds the true target token (shifted by one) at each decoder step. This allows the full target sequence to be processed in one parallel forward pass with a causal mask. The decoder input is target[:-1]; the labels are target[1:].

43. Check yourself — teacher forcing

Check

Work out the mechanism before clicking.

Check your understanding

During training with teacher forcing, what is fed as the decoder input at step t?

  • A. The ground-truth target token at position t−1 (correct)
  • B. The model's own predicted token from step t−1
  • C. The encoder output (memory) at position t
  • D. A zero vector (the decoder starts cold at each step)

Answer: A

Why: Teacher forcing feeds the true target token (shifted by one) at each decoder step. This allows the full target sequence to be processed in one parallel forward pass with a causal mask. The decoder input is target[:-1]; the labels are target[1:].

Why B tempts people
Feeding the model's own predictions is autoregressive decoding — that is inference, not training. During training this would slow convergence and compound early errors (exposure bias).
Why C tempts people
The encoder memory is attended to via cross-attention, not fed directly as the decoder's token input. The decoder's own embedding layer receives the target tokens.
Why D tempts people
A zero input at every step would destroy the sequential structure the causal mask is designed to expose — the decoder would have no token history to attend to.

44. Answer it before you see the options: Check yourself — beam search pruning

Prediction

Predict first

In the beam search trace (k=2), after step 2, which sequence is dropped first?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: [3, 3] with cum log-prob −2.696

Why: [3, 3] has cumulative log-prob −2.696, the lowest of the four candidates. After sorting descending by cum log-prob and keeping the top-2, [3, 3] and [4, 4] are dropped; [3, 4] (−1.155) and [4, 3] (−2.303) are retained.

45. Check yourself — beam search pruning

Check

Use the verified trace numbers.

Check your understanding

In the beam search trace (k=2), after step 2, which sequence is dropped first?

  • A. [4, 4] with cum log-prob −2.590
  • B. [3, 4] with cum log-prob −1.155
  • C. [4, 3] with cum log-prob −2.303
  • D. [3, 3] with cum log-prob −2.696 (correct)

Answer: D

Why: [3, 3] has cumulative log-prob −2.696, the lowest of the four candidates. After sorting descending by cum log-prob and keeping the top-2, [3, 3] and [4, 4] are dropped; [3, 4] (−1.155) and [4, 3] (−2.303) are retained.

Why A tempts people
[4, 4] at −2.590 is the third-ranked sequence. It is also dropped, but [3, 3] at −2.696 is ranked fourth — the first one eliminated.
Why B tempts people
[3, 4] at −1.155 is the best sequence (highest cum log-prob) and is kept, not dropped.
Why C tempts people
[4, 3] at −2.303 is the second-best and is retained in the top-2 beams.

46. Rule out three: Check yourself — label smoothing

Elimination

Eliminate the wrong options

Label smoothing ε=0.1 is applied to a problem with vocabulary size V=10. What is the smoothed probability assigned to the correct class?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. 0.91
  • B. 0.90
  • C. 0.10
  • D. 1.00

Survives elimination: A

Why: The true-class smoothed probability is 1 − ε + ε/V = 1 − 0.1 + 0.1/10 = 0.9 + 0.01 = 0.91. Each other class gets ε/V = 0.01, and 0.91 + 9×0.01 = 1.0.

47. Check yourself — label smoothing

Check

Apply the formula.

Check your understanding

Label smoothing ε=0.1 is applied to a problem with vocabulary size V=10. What is the smoothed probability assigned to the correct class?

  • A. 0.91 (correct)
  • B. 0.90
  • C. 0.10
  • D. 1.00

Answer: A

Why: The true-class smoothed probability is 1 − ε + ε/V = 1 − 0.1 + 0.1/10 = 0.9 + 0.01 = 0.91. Each other class gets ε/V = 0.01, and 0.91 + 9×0.01 = 1.0.

Why B tempts people
0.90 would be correct if ε/V were 0 — i.e., the ε/V redistribution term was ignored and only (1−ε) was used as the true-class prob.
Why C tempts people
0.10 is the value of ε itself (the total probability mass distributed to non-true classes), not the true-class probability.
Why D tempts people
1.00 is the un-smoothed one-hot probability. Label smoothing specifically moves this below 1.0 to prevent the model from becoming overconfident.

48. Your turn: seq2seq from scratch

Section

Project

49. Project: digit-echo seq2seq Transformer

Concept

Build a full encoder-decoder Transformer for a toy seq2seq task: given a source digit 0-9, generate [digit, digit, EOS]. Three milestones: training loop with teacher forcing → autoregressive inference → beam search (k=2).

#milestonekey tool
1Teacher-forcing training loop (300 epochs, loss → 0.06)nn.Transformer, CrossEntropyLoss
2Autoregressive greedy decode (src=3 → [3, 3, EOS])causal mask growing per step
3Beam search k=2 on step-1 probstorch.topk, cumulative log-probs

Build rules: use batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.

50. Break it if you can: Project: digit-echo seq2seq Transformer

Counterexample

Discussion prompt

Build rules: use batch_first=True throughout; always apply a causal mask in the decoder; feed target[:-1] as decoder input and target[1:] as labels.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

51. Milestone 1 — teacher-forcing training loop

Worked example

Your turn: set up ToySeq2Seq (src_vocab=12, tgt_vocab=12, d=16, nhead=2, ff=32). Train 300 epochs with Adam lr=5e-3. Predict the final loss.

Hint: decoder input = tgts_in (BOS prepended), labels = tgts_out (EOS appended). Causal mask shape = (3,3). loss = CrossEntropyLoss(logits.reshape(-1,VTGT), labels.reshape(-1)).

import torch, torch.nn as nn
torch.manual_seed(99)
BOS, EOS, VTGT = 10, 11, 12
srcs    = torch.arange(10).unsqueeze(1)     # (10,1)
tgts_in = torch.stack([torch.tensor([BOS,i,i]) for i in range(10)])
tgts_out= torch.stack([torch.tensor([i,i,EOS]) for i in range(10)])
causal  = torch.triu(torch.ones(3,3)*float('-inf'), diagonal=1)
model2  = ToySeq2Seq()   # defined above
opt     = torch.optim.Adam(model2.parameters(), lr=5e-3)
for ep in range(300):
    opt.zero_grad()
    logits = model2(srcs, tgts_in, tgt_mask=causal)
    loss = nn.CrossEntropyLoss()(logits.reshape(-1,VTGT), tgts_out.reshape(-1))
    loss.backward(); opt.step()
print(f'final loss: {loss.item():.4f}')
epochloss
02.7717
500.8280
1000.4379
2000.3123
2990.0609

52. Watch it run: Milestone 1 — teacher-forcing training loop

Pattern

Step through it

Step through Milestone 1 — teacher-forcing training loop one row at a time. What is driving the change, and what would the row after the last one be?

  1. Step 1: epoch is 0
  2. Step 2: epoch is 50
  3. Step 3: epoch is 100
  4. Step 4: epoch is 200
  5. Step 5: epoch is 299

53. Milestone 2 — autoregressive greedy decode

Worked example

Your turn: after training, decode src=3 step by step (no teacher forcing). Predict the output token at step 1.

Hint: start with generated=[BOS]; at each step extend the causal mask by 1 row+col; append logits[0,-1].argmax().item(); stop when EOS.

model2.eval()
with torch.no_grad():
    src_t = torch.tensor([[3]])
    gen   = [BOS]
    for _ in range(4):
        tl = len(gen)
        cm = torch.triu(torch.ones(tl,tl)*float('-inf'), diagonal=1)
        lg = model2(src_t, torch.tensor([gen]), tgt_mask=cm)
        nxt = lg[0,-1].argmax().item()
        gen.append(nxt)
        if nxt == EOS: break
print(gen)  # [10, 3, 3, 11]
stepdecoder inputnext token predicted
1[BOS=10]3
2[BOS, 3]3
3[BOS, 3, 3]EOS=11 → stop

54. What each one costs: Milestone 2 — autoregressive greedy decode

Trade off

Comparison matrix

From Milestone 2 — autoregressive greedy decode: every row here is a choice with a cost. Fill the next token predicted column, then say which row you would actually pick and what you give up for it.

stepdecoder inputnext token predicted
1[BOS=10]3
2[BOS, 3]3
3[BOS, 3, 3]EOS=11 → stop

55. Milestone 3 — beam search k=2

Worked example

Your turn: implement beam search for 2 steps on the given step-1 probs. Predict which sequences survive after step 2.

Hint: torch.topk(log_probs, k=2) at each step; accumulate log-probs by summing; after expanding all k beams, take global top-k.

import torch
step1 = torch.log(torch.tensor([0.05,0.10,0.15,0.45,0.25]))
tk1 = torch.topk(step1, k=2)   # ids=[3,4], lps=[-0.799,-1.386]
b1 = torch.log(torch.tensor([0.02,0.03,0.10,0.15,0.70]))
b2 = torch.log(torch.tensor([0.05,0.05,0.20,0.40,0.30]))
cands = []
for i,src_lp in enumerate([tk1.values[0], tk1.values[1]]):
    for tok,lp in zip(torch.topk(([b1,b2][i]),k=2).indices,
                      torch.topk(([b1,b2][i]),k=2).values):
        cands.append((tok.item(), src_lp.item()+lp.item()))
cands.sort(key=lambda x:-x[1])
print(cands[:2])  # kept beams
sequencecum log-probkept?
[3, 4]-1.155yes
[4, 3]-2.303yes
[4, 4]-2.590dropped
[3, 3]-2.696dropped

56. Which is which, by kept?

Discrimination

Sort into buckets

Sort these by kept?, from memory, without looking back at Milestone 3 — beam search k=2. Telling them apart on the spot is the skill; the table is only where the answer happens to be written down.

yes
[3, 4]; [4, 3]
dropped
[4, 4]; [3, 3]
g1
kept? is "yes" for [3, 4], [4, 3] — that is what the table on "Milestone 3 — beam search k=2" records, and it is the single property separating this group from the rest.
g2
kept? is "dropped" for [4, 4], [3, 3] — that is what the table on "Milestone 3 — beam search k=2" records, and it is the single property separating this group from the rest.

57. The full program

Concept

import torch, torch.nn as nn
torch.manual_seed(99)
BOS,EOS,V = 10,11,12
class ToySeq2Seq(nn.Module):
    def __init__(self,vsrc=V,vtgt=V,d=16,nh=2,ff=32,nl=1):
        super().__init__()
        self.se=nn.Embedding(vsrc,d); self.te=nn.Embedding(vtgt,d)
        self.tr=nn.Transformer(d,nh,nl,nl,ff,batch_first=True)
        self.pr=nn.Linear(d,vtgt)
    def forward(self,s,t,tgt_mask=None):
        return self.pr(self.tr(self.se(s),self.te(t),tgt_mask=tgt_mask))
srcs=torch.arange(10).unsqueeze(1)
ti=torch.stack([torch.tensor([BOS,i,i])for i in range(10)])
to=torch.stack([torch.tensor([i,i,EOS])for i in range(10)])
cm=torch.triu(torch.ones(3,3)*float('-inf'),diagonal=1)
m=ToySeq2Seq(); opt=torch.optim.Adam(m.parameters(),lr=5e-3)
for _ in range(300):
    opt.zero_grad()
    nn.CrossEntropyLoss()(m(srcs,ti,tgt_mask=cm).reshape(-1,V),to.reshape(-1)).backward()
    opt.step()
m.eval()
with torch.no_grad():
    g=[BOS]
    for _ in range(4):
        tl=len(g); cm2=torch.triu(torch.ones(tl,tl)*float('-inf'),diagonal=1)
        nxt=m(torch.tensor([[3]]),torch.tensor([g]),tgt_mask=cm2)[0,-1].argmax().item()
        g.append(nxt)
        if nxt==EOS: break
print(g)  # [10, 3, 3, 11]
componentchoicewhy
encodernn.TransformerEncoder (bidirectional)full context over source — no causal mask
decodernn.TransformerDecoder + causal maskcan't peek at future target tokens
trainingteacher forcing + CrossEntropyLossparallel over sequence, fast convergence
inferenceautoregressive (own output → next input)no ground truth available at deploy time
losslabel_smoothing=0.1prevents overconfident logits, improves generalization

Final training loss 0.0609; greedy decode of digit 3 → [BOS, 3, 3, EOS]. Same five-step training loop as Lesson 40 — only the model (Transformer) and the shifted target trick change.

58. Fill in: choice for The full program

Comparison

Comparison matrix

From The full program: refill the choice column from what you know. The rest of the table is as it appeared.

componentchoicewhy
encodernn.TransformerEncoder (bidirectional)full context over source — no causal mask
decodernn.TransformerDecoder + causal maskcan't peek at future target tokens
trainingteacher forcing + CrossEntropyLossparallel over sequence, fast convergence
inferenceautoregressive (own output → next input)no ground truth available at deploy time
losslabel_smoothing=0.1prevents overconfident logits, improves generalization

59. Show it off

Concept

Out loud, slides closed: (1) explain what the encoder produces and how the decoder uses it via cross-attention, (2) walk through one teacher-forcing training step showing tgt_in vs tgt_out, and (3) trace beam search for k=2 from [3, 4] at step 2.

Stretch (homework): add label_smoothing=0.1 to the training loss and compare final loss; implement length normalization for the beam search; extend the task to actual number-to-words translation (e.g., 42 → 'forty two'). Next up: Lesson 87 — KV cache and efficient inference.

60. Connect it up: Lesson 86: Transformer Encoder-Decoder & Seq2Seq

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Encoder-Decoder architecture · Teacher forcing & training · Autoregressive inference & beam search · Your turn: seq2seq from scratch. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

61. What you can do now

Recap

conceptthe one thing to remember
encoder-decoderencoder → memory (cached); decoder cross-attends every step
teacher forcingtraining only — feed ground truth, not model output
autoregressiveinference only — grow hypothesis one token at a time
beam searchkeep top-k cumulative log-probs; normalize by length
label smoothing1−ε+ε/V for true class; prevents overconfident logits

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 86 — Encoder-Decoder Transformer, Teacher Forcing, Beam Search — Barron · USAAIO Round 2 Preparation, 2026
  2. All attention scores, causal masks, label-smoothing values, beam search log-probs, and toy seq2seq training losses verified with torch 2.7.1+cpu and numpy 2.2.6, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108