Lesson 102: RLHF & DPO — Instruction Fine-Tuning

USAAIO Lesson 102, from Phase 3. It covers the full RLHF pipeline, running from supervised fine-tuning through a reward model to PPO with a KL penalty, along with the Bradley-Terry pairwise loss and the shaped reward r_env − β·KL. It then presents DPO as a closed-form alternative that eliminates the reward model entirely. Every loss value, KL divergence, and DPO margin was verified with torch 2.7.1+cpu at seed 42. The lesson runs to 26 slides.

Subject: Machine Learning · 51 slides · code lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. RLHF & DPO Instruction Fine-Tuning

Title

USAAIO · Lesson 102 · Phase 3

How raw pretrained LLMs become helpful assistants: supervised fine-tuning, reward modeling from human preferences, PPO with KL penalty, and DPO as a cleaner closed-form alternative. Every loss value derived and run.

2. By the end of this lesson you can

Objectives

  1. Draw the full 3-step RLHF pipeline (SFT → RM → RL) and state each stage's role and training signal
  2. Implement a reward model that scores completions using the Bradley-Terry pairwise loss on preference pairs
  3. Explain how PPO is used in RLHF and why the KL divergence penalty prevents policy collapse
  4. Derive the DPO objective and explain why it needs no separate reward model
  5. Compare RLHF and DPO on stability, compute, and data requirements — and state when each is preferred

3. What survived from LoRA — Low-Rank Adaptation?

Warm-up

Discussion prompt

Before we open Lesson 102: RLHF & DPO — Instruction Fine-Tuning: without looking back, what was the main idea of LoRA — Low-Rank Adaptation, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

LoRA (Low-Rank Adaptation) for parameter-efficient fine-tuning — low-rank decomposition delta_W=AB, B=0 initialization, alpha/r scaling, which layers to apply LoRA to, and full LoRALinear nn.Module implementation.

4. Why raw pretraining is not enough

Section

Part 1 of 4

5. Pretraining vs instruction following

Concept

A pretrained LLM minimizes next-token cross-entropy on internet text. It learns what language looks like, not what a helpful response looks like. Instruction fine-tuning bridges that gap.

stagedataobjectiveresult
pretrainingraw internet text (trillions of tokens)next-token cross-entropylanguage model p(x)
SFThuman-written instruction→response pairsnext-token cross-entropy on response onlyinstruction follower
RLHF RM+RLhuman pairwise preferencesmaximize reward - KLhelpful & aligned

SFT alone gives solid instruction following but cannot capture nuanced human preferences — it treats all demonstration tokens equally. RLHF uses human ranking signals to push the model further.

6. Fill in: data for Pretraining vs instruction following

Comparison

Comparison matrix

From Pretraining vs instruction following: refill the data column from what you know. The rest of the table is as it appeared.

stagedataobjectiveresult
pretrainingraw internet text (trillions of tokens)next-token cross-entropylanguage model p(x)
SFThuman-written instruction→response pairsnext-token cross-entropy on response onlyinstruction follower
RLHF RM+RLhuman pairwise preferencesmaximize reward - KLhelpful & aligned

7. The 3-step RLHF pipeline

Concept

  1. SFT (supervised fine-tuning): fine-tune the pretrained model on high-quality demonstration data (prompt → ideal response pairs). Produces the SFT model π_SFT.
  2. Reward model training: collect human preferences over pairs of completions. Train a scalar-output model r_φ(x, y) to rank them correctly.
  3. RL with PPO: treat the LLM as a policy π_θ. Use r_φ as environment reward; add a KL penalty against π_SFT to prevent runaway optimization.

Each stage feeds the next: SFT initializes the policy and the RM; the RM provides the reward signal; PPO then optimizes the policy while the SFT model acts as a frozen reference for the KL term.

8. Break it if you can: The 3-step RLHF pipeline

Counterexample

Discussion prompt

Each stage feeds the next: SFT initializes the policy and the RM; the RM provides the reward signal; PPO then optimizes the policy while the SFT model acts as a frozen reference for the KL term.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

9. Step 1: Supervised Fine-Tuning (SFT)

Section

Part 2 of 4

10. Predict the next row: SFT: cross-entropy on the response only

Pattern

Predict first

The table runs: 1 (init) | 2.7608 · 2 | 2.6412

In SFT: cross-entropy on the response only, given the rows so far: what is the next one — the row where step is 3?

Correct: 3 | 2.5241

steploss
1 (init)2.7608
22.6412
32.5241

Why: The relationship between the columns, not the individual numbers, is what generates the next row. We do NOT train on the prompt tokens; supervising on them teaches the model to copy prompts, not respond to them.

11. SFT: cross-entropy on the response only

Worked example

Define the loss — minimize negative log-likelihood over response tokens only

Why: We do NOT train on the prompt tokens; supervising on them teaches the model to copy prompts, not respond to them.

\[ \mathcal{L}_{\text{SFT}}(\theta) = -\sum_{t=1}^{T} \log \pi_\theta(y_t \mid x, y_{<t}) \]

Run 3 gradient steps on a tiny 8-token LM (vocab=8, d=8, Adam lr=0.01, seed=42)

Why: Each step computes cross-entropy on next-token targets [1,2,3,4] given input [0,1,2,3], backprops through embedding + linear head.

import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
embed = nn.Embedding(8, 8)
head  = nn.Linear(8, 8)
opt   = torch.optim.Adam(
    list(embed.parameters()) + list(head.parameters()), lr=0.01)
toks  = torch.tensor([[0, 1, 2, 3]])
targs = torch.tensor([[1, 2, 3, 4]])
for step in range(1, 4):
    opt.zero_grad()
    logits = head(embed(toks))          # (1, 4, 8)
    loss   = F.cross_entropy(logits.view(-1, 8), targs.view(-1))
    loss.backward()
    opt.step()
    print(f'step {step}: loss={loss.item():.4f}')
steploss
1 (init)2.7608
22.6412
32.5241

12. What each one costs: SFT: cross-entropy on the response only

Trade off

Comparison matrix

From SFT: cross-entropy on the response only: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.

steploss
1 (init)2.7608
22.6412
32.5241

13. Step 2: Reward Model from Human Preferences

Section

Part 3 of 4

14. Preference data and the Bradley-Terry model

Concept

Annotators see a prompt x and two completions y_w (winner) and y_l (loser) and mark which they prefer. The reward model learns to assign higher scalar reward to y_w.

\[ p(y_w \succ y_l \mid x) = \sigma\bigl(r_\phi(x, y_w) - r_\phi(x, y_l)\bigr) \]

\[ \mathcal{L}_{\text{RM}}(\phi) = -\mathbb{E}_{(x,y_w,y_l)}\bigl[\log \sigma\bigl(r_\phi(x,y_w) - r_\phi(x,y_l)\bigr)\bigr] \]

This is binary cross-entropy on the margin between the two reward scores — the Bradley-Terry ranking model. It never trains on absolute reward values, only on which completion is relatively better.

15. By analogy: Preference data and the Bradley-Terry model

Analogy

Discussion prompt

Explain Preference data and the Bradley-Terry model by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.

Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.

Answer:

Annotators see a prompt x and two completions y_w (winner) and y_l (loser) and mark which they prefer. The reward model learns to assign higher scalar reward to y_w.

16. Restore the missing line: Reward model: pairwise loss computation

Fill the middle

Fill in the blanks

From Reward model: pairwise loss computation — one line has had its right-hand side removed. Put it back.

class ToyRewardModel(nn.Module):
def __init__(self, vocab, d):
super().__init__()
self.embed = nn.Embedding(vocab, d)
self.head = nn.Linear(d, 1)
def forward(self, x):
return self.head(self.embed(x).mean(dim=1)).squeeze(-1)

rm = ToyRewardModel(vocab=8, d=16)
chosen = torch.tensor([[1, 3, 5, 6, 7]]) # higher quality
rejected = torch.tensor([[1, 2, 2, 2, 2]]) # lower quality
r_w = rm(chosen); r_l = rm(rejected)
loss_rm = -F.logsigmoid(r_w - r_l).mean()
print(f'r(chosen)=___ r(rejected)=___')
print(f'RM loss (Bradley-Terry)=___')

Why: chosen is what everything below it consumes, so the wrong expression here fails later and somewhere else. The RM architecture is the same as the SFT model but with a scalar head replacing the vocab projection; it can be initialized from the SFT checkpoint.

17. Reward model: pairwise loss computation

Worked example

Build a ToyRewardModel (embed → mean-pool → scalar) and forward both chosen and rejected completions

Why: The RM architecture is the same as the SFT model but with a scalar head replacing the vocab projection; it can be initialized from the SFT checkpoint.

class ToyRewardModel(nn.Module):
    def __init__(self, vocab, d):
        super().__init__()
        self.embed = nn.Embedding(vocab, d)
        self.head  = nn.Linear(d, 1)
    def forward(self, x):
        return self.head(self.embed(x).mean(dim=1)).squeeze(-1)

rm       = ToyRewardModel(vocab=8, d=16)
chosen   = torch.tensor([[1, 3, 5, 6, 7]])   # higher quality
rejected = torch.tensor([[1, 2, 2, 2, 2]])   # lower quality
r_w = rm(chosen);   r_l = rm(rejected)
loss_rm = -F.logsigmoid(r_w - r_l).mean()
print(f'r(chosen)={r_w.item():.4f}  r(rejected)={r_l.item():.4f}')
print(f'RM loss (Bradley-Terry)={loss_rm.item():.4f}')
quantityvaluemeaning
r(chosen)-0.6170scalar reward for y_w at random init
r(rejected)0.1790scalar reward for y_l at random init
margin r_w - r_l-0.7960negative: RM not yet trained
RM loss1.1683high — gradients push r_w > r_l

Interpret: loss = 1.1683 because r(chosen) < r(rejected) at random init

Why: After training, the RM loss minimizes when r(y_w) >> r(y_l) for all preference pairs — the model has learned the reward signal from human labels.

18. Inspect it line by line: Reward model: pairwise loss computation

Error analysis

Annotate

Walk the callouts on Reward model: pairwise loss computation. Each one is a place this is easy to get subtly wrong.

  • The RM architecture is the same as the SFT model but with a scalar head replacing the vocab projection; it can be initialized from the SFT checkpoint.
  • After training, the RM loss minimizes when r(y_w) >> r(y_l) for all preference pairs — the model has learned the reward signal from human labels.

19. Something is wrong here: treating RM training as regression on absolute scores

Anomaly

Predict first

A student writes this, and it looks reasonable:

Annotators assign numeric scores (1–10) to each completion separately. Train the RM with MSE: loss = MSE(r_phi(x,y), score).

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Seems natural — match the model's scalar to the annotator's rating.

Collect pairwise comparisons (A vs B, which is better?). Train the RM with Bradley-Terry loss on the margin: loss = -log_sigmoid(r_w - r_l).

Why: Seems natural — match the model's scalar to the annotator's rating.

20. Trap: treating RM training as regression on absolute scores

Trap

The trap

Annotators assign numeric scores (1–10) to each completion separately. Train the RM with MSE: loss = MSE(r_phi(x,y), score).

Compute loss = mean_squared_error(r_phi(completion), human_score)

Why: Seems natural — match the model's scalar to the annotator's rating.

Result: annotators are inconsistent and on different scales. The RM memorizes rater biases rather than learning which responses humans genuinely prefer.

The fix

Collect pairwise comparisons (A vs B, which is better?). Train the RM with Bradley-Terry loss on the margin: loss = -log_sigmoid(r_w - r_l).

Compute loss = -log sigma(r_phi(x, y_w) - r_phi(x, y_l)) for every (y_w, y_l) preference pair

Why: Relative comparison is far more consistent for humans than absolute scoring; the model learns what relatively better means, not a calibrated scale.

Pairwise preference data (A/B) is cheaper to collect and more reliable — this is why InstructGPT, Claude, and ChatGPT all use the pairwise paradigm.

21. Break it on purpose: treating RM training as regression on…

Break the constraint

Discussion prompt

The rule this trap just fixed:

Relative comparison is far more consistent for humans than absolute scoring; the model learns what relatively better means, not a calibrated scale.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Seems natural — match the model's scalar to the annotator's rating.

22. Step 3: PPO with KL Penalty

Section

Part 4a of 4

23. The policy optimization objective in RLHF

Concept

After SFT and RM training, the policy π_θ (initialized from π_SFT) is optimized to maximize expected reward while staying close to the SFT reference distribution.

\[ \max_{\pi_\theta} \mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi_\theta(\cdot|x)} \Bigl[ r_\phi(x, y) - \beta \cdot D_{\text{KL}}\bigl(\pi_\theta(\cdot|x) \| \pi_{\text{SFT}}(\cdot|x)\bigr) \Bigr] \]

Without the KL term the policy can exploit the reward model (reward hacking): generate grammatically broken or degenerate text that happens to fool the RM into high scores.

24. Break it if you can: The policy optimization objective in RLHF

Counterexample

Discussion prompt

After SFT and RM training, the policy π_θ (initialized from π_SFT) is optimized to maximize expected reward while staying close to the SFT reference distribution.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

Without the KL term the policy can exploit the reward model (reward hacking): generate grammatically broken or degenerate text that happens to fool the RM into high scores.

25. Restore the missing line: PPO shaped reward: r_env − β·KL

Fill the middle

Fill in the blanks

From PPO shaped reward: r_env − β·KL — one line has had its right-hand side removed. Put it back.

torch.manual_seed(7)
pol_logits = torch.randn(6, 8) # policy (6 tokens, vocab=8)
ref_logits = torch.randn(6, 8) # frozen SFT reference
pol_lp = F.log_softmax(pol_logits, dim=-1)
ref_lp = F.log_softmax(ref_logits, dim=-1)
# KL per token: sum over vocab
kl = (pol_lp.exp() * (pol_lp - ref_lp)).sum(dim=-1)
print('KL per token:', kl.detach().numpy().round(4))
print('Mean KL:', round(kl.mean().item(), 4))
r_env = 0.85 # reward from trained RM
beta = 0.1
shaped = r_env - beta * kl.mean().item()
print('shaped reward:', round(shaped, 4))

Why: ref_lp is what everything below it consumes, so the wrong expression here fails later and somewhere else. KL(pi || pi_ref) = sum_v pi(v|x) * (log pi(v|x) - log pi_ref(v|x)) measures how far the policy has drifted; summed/averaged over the response tokens.

26. PPO shaped reward: r_env − β·KL

Worked example

Compute per-token KL divergence between the current policy and the reference (SFT) distribution

Why: KL(pi || pi_ref) = sum_v pi(v|x) * (log pi(v|x) - log pi_ref(v|x)) measures how far the policy has drifted; summed/averaged over the response tokens.

torch.manual_seed(7)
pol_logits = torch.randn(6, 8)      # policy (6 tokens, vocab=8)
ref_logits = torch.randn(6, 8)      # frozen SFT reference
pol_lp = F.log_softmax(pol_logits, dim=-1)
ref_lp = F.log_softmax(ref_logits,  dim=-1)
# KL per token: sum over vocab
kl = (pol_lp.exp() * (pol_lp - ref_lp)).sum(dim=-1)
print('KL per token:', kl.detach().numpy().round(4))
print('Mean KL:', round(kl.mean().item(), 4))
r_env  = 0.85          # reward from trained RM
beta   = 0.1
shaped = r_env - beta * kl.mean().item()
print('shaped reward:', round(shaped, 4))
tokenKL(pi||pi_ref)
t=00.6565
t=10.2939
t=20.9556
t=30.6660
t=41.1052
t=50.5329
mean0.7017
shaped reward (r_env=0.85, beta=0.1)0.7798

Interpret: shaped_reward = 0.85 − 0.1 × 0.7017 = 0.7798

Why: The KL penalty reduces the effective reward; if the policy drifted more (higher KL), the penalty grows and discourages further divergence from the SFT baseline.

27. Fill in: KL(pi||pi_ref) for PPO shaped reward: r_env − β·KL

Comparison

Comparison matrix

From PPO shaped reward: r_env − β·KL: refill the KL(pi||pi_ref) column from what you know. The rest of the table is as it appeared.

tokenKL(pi||pi_ref)
t=00.6565
t=10.2939
t=20.9556
t=30.6660
t=41.1052
t=50.5329
mean0.7017
shaped reward (r_env=0.85, beta=0.1)0.7798

28. Why PPO — not vanilla policy gradient?

Concept

Vanilla REINFORCE has high variance and can take catastrophically large gradient steps. PPO clips the probability ratio to a trust region, preventing destructive updates.

\[ L^{\text{CLIP}}(\theta) = \mathbb{E}\Bigl[\min\bigl(r_t(\theta)\,\hat{A}_t,\;\text{clip}(r_t(\theta),\,1-\varepsilon,\,1+\varepsilon)\,\hat{A}_t\bigr)\Bigr] \]

termmeaning
r_t(theta) = pi_theta(a_t|s_t) / pi_old(a_t|s_t)probability ratio (how much policy changed)
A_hat_tadvantage estimate (reward − value baseline)
clip(..., 1-eps, 1+eps)limits ratio to [0.8, 1.2] when eps=0.2
min(unclipped, clipped)pessimistic: no credit for ratio outside trust region

29. Where does each piece belong: Lesson 102: RLHF & DPO — Instruction…

Sorting

Sort into buckets

These are the pieces of Lesson 102: RLHF & DPO — Instruction Fine-Tuning, out of order. Put each one back under the part of the lesson it belongs to.

Why raw pretraining is not enough
Pretraining vs instruction following; The 3-step RLHF pipeline
Step 2: Reward Model from Human Preferences
Preference data and the Bradley-Terry model; Reward model: pairwise loss computation
Step 3: PPO with KL Penalty
The policy optimization objective in RLHF; PPO shaped reward: r_env − β·KL; Why PPO — not vanilla policy gradient?
s1
Why raw pretraining is not enough is where Lesson 102: RLHF & DPO — Instruction Fine-Tuning puts Pretraining vs instruction following, The 3-step RLHF pipeline. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s2
Step 2: Reward Model from Human Preferences is where Lesson 102: RLHF & DPO — Instruction Fine-Tuning puts Preference data and the Bradley-Terry model, Reward model: pairwise loss computation. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.
s3
Step 3: PPO with KL Penalty is where Lesson 102: RLHF & DPO — Instruction Fine-Tuning puts The policy optimization objective in RLHF, PPO shaped reward: r_env − β·KL, Why PPO — not vanilla policy gradient?. Knowing which part of the lesson a problem belongs to is most of knowing which method to reach for.

30. DPO: Direct Preference Optimization

Section

Part 4b of 4

31. DPO: eliminating the reward model

Concept

Rafailov et al. (2023) showed the optimal RLHF policy implies a closed-form relationship between the policy, the reference, and the reward. DPO inverts this to directly optimize the policy on preference data — no RM needed.

\[ \mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x,y_w,y_l)} \Bigl[\log \sigma\Bigl(\beta \log\frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log\frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\Bigr)\Bigr] \]

The term beta * log(pi/pi_ref) on a sequence y is the implicit reward the optimal RL policy assigns to y. DPO trains the policy to rank the chosen completion's implicit reward above the rejected one — exactly what the RM would do, but without a separate reward model.

32. Predict the next row: DPO loss: computing implicit rewards and margin

Pattern

Predict first

The table runs: log pi(y_w|x) - log pi_ref(y_w|x) | 0.6000 | -4.2 - (-4.8) · log pi(y_l|x) - log pi_ref(y_l|x) | -0.2000 | -5.1 - (-4.9) · imp_w = beta * 0.6 | 0.0600 | implicit reward for chosen · imp_l = beta * (-0.2) | -0.0200 | implicit reward for rejected · margin = imp_w - imp_l | 0.0800 | beta * (ratio_w - ratio_l)

In DPO loss: computing implicit rewards and margin, given the rows so far: what is the next one — the row where quantity is DPO loss = -log sigma(0.08)?

Correct: DPO loss = -log sigma(0.08) | 0.6539 | positive: policy not yet trained

quantityvalueformula
log pi(y_w|x) - log pi_ref(y_w|x)0.6000-4.2 - (-4.8)
log pi(y_l|x) - log pi_ref(y_l|x)-0.2000-5.1 - (-4.9)
imp_w = beta * 0.60.0600implicit reward for chosen
imp_l = beta * (-0.2)-0.0200implicit reward for rejected
margin = imp_w - imp_l0.0800beta * (ratio_w - ratio_l)
DPO loss = -log sigma(0.08)0.6539positive: policy not yet trained

Why: The relationship between the columns, not the individual numbers, is what generates the next row. Sum the per-token log-probs under both the policy and the reference, then take the difference.

33. DPO loss: computing implicit rewards and margin

Worked example

Compute sequence-level log-prob ratios log(pi_theta/pi_ref) for chosen and rejected completions

Why: Sum the per-token log-probs under both the policy and the reference, then take the difference. This is the implicit reward for each completion.

# Toy sequence-level log-probs (pretend these are summed over tokens)
pi_chosen_lp  = -4.2;  ref_chosen_lp  = -4.8   # log p for y_w
pi_rej_lp     = -5.1;  ref_rej_lp     = -4.9   # log p for y_l
beta = 0.1
# implicit rewards
imp_w = beta * (pi_chosen_lp - ref_chosen_lp)   # 0.1 * 0.6
imp_l = beta * (pi_rej_lp   - ref_rej_lp)       # 0.1 * (-0.2)
margin = imp_w - imp_l
loss_dpo = -float(torch.tensor(margin).sigmoid().log())
print(f'imp_w={imp_w:.4f}  imp_l={imp_l:.4f}')
print(f'margin={margin:.4f}  DPO_loss={loss_dpo:.4f}')
quantityvalueformula
log pi(y_w|x) - log pi_ref(y_w|x)0.6000-4.2 - (-4.8)
log pi(y_l|x) - log pi_ref(y_l|x)-0.2000-5.1 - (-4.9)
imp_w = beta * 0.60.0600implicit reward for chosen
imp_l = beta * (-0.2)-0.0200implicit reward for rejected
margin = imp_w - imp_l0.0800beta * (ratio_w - ratio_l)
DPO loss = -log sigma(0.08)0.6539positive: policy not yet trained

Interpret: DPO loss decreases as imp_w grows and imp_l shrinks — i.e. the policy up-weights chosen and down-weights rejected relative to the reference

Why: The gradient of DPO loss pushes the policy to increase log prob of y_w and decrease log prob of y_l, while the pi_ref denominator anchors the magnitude — β controls the trade-off strength.

34. What each one costs: DPO loss: computing implicit rewards and margin

Trade off

Comparison matrix

From DPO loss: computing implicit rewards and margin: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.

quantityvalueformula
log pi(y_w|x) - log pi_ref(y_w|x)0.6000-4.2 - (-4.8)
log pi(y_l|x) - log pi_ref(y_l|x)-0.2000-5.1 - (-4.9)
imp_w = beta * 0.60.0600implicit reward for chosen
imp_l = beta * (-0.2)-0.0200implicit reward for rejected
margin = imp_w - imp_l0.0800beta * (ratio_w - ratio_l)
DPO loss = -log sigma(0.08)0.6539positive: policy not yet trained

35. Something is wrong here: DPO does not need the reference model at inference time

Anomaly

Predict first

A student writes this, and it looks reasonable:

After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: The DPO objective is framed around the ratio, so the reference must stay active.

The reference model pi_ref is needed only during DPO training to compute the implicit reward ratios. At inference, just sample from pi_theta normally.

Why: The DPO objective is framed around the ratio, so the reference must stay active.

36. Trap: DPO does not need the reference model at inference time

Trap

The trap

After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.

At serving time, compute beta * (log pi_theta(y|x) - log pi_ref(y|x)) for each candidate response before returning the best one

Why: The DPO objective is framed around the ratio, so the reference must stay active.

The fix

The reference model pi_ref is needed only during DPO training to compute the implicit reward ratios. At inference, just sample from pi_theta normally.

At serving time, use pi_theta.generate(prompt) with no reference model

Why: DPO bakes the preference signal into the policy weights during training. The trained pi_theta already reflects the human preferences — pi_ref is discarded after training is complete.

37. Which of these survive contact with Lesson 102: RLHF & DPO — Instruction…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Annotators see a prompt x and two completions y_w (winner) and y_l (loser) and mark which they prefer. The reward model learns to assign higher scalar reward to y_w.; After SFT and RM training, the policy π_θ (initialized from π_SFT) is optimized to maximize expected reward while staying close to the SFT reference distribution.; Vanilla REINFORCE has high variance and can take catastrophically large gradient steps. PPO clips the probability ratio to a trust region, preventing destructive updates.
Breaks
Annotators assign numeric scores (1–10) to each completion separately. Train the RM with MSE: loss = MSE(r_phi(x,y), score).; After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.
sound
These are stated as this lesson states them — each one survives the edge cases Lesson 102: RLHF & DPO — Instruction Fine-Tuning puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

38. RLHF vs DPO: trade-offs

Concept

dimensionRLHF (PPO)DPO
componentsSFT + RM + policy + value netSFT + policy (pi_ref frozen)
training stabilitycomplex: RM updates, PPO clip, KLsimple: single SL loss
online vs offlineonline (samples during training)offline (preference dataset fixed)
computehigh — runs 4 models simultaneouslylow — equivalent to SFT
reward hackingpossible (RM can be gamed)implicit RM cannot diverge past pi_ref
use todayGPT-4, Gemini Ultramany open-source models (Zephyr, Tulu-3)

DPO is preferred when compute is limited or training stability is a concern. RLHF with PPO allows online exploration — the policy can discover new behaviors during training — which matters at frontier scale.

39. Fill in: RLHF (PPO) for RLHF vs DPO: trade-offs

Comparison

Comparison matrix

From RLHF vs DPO: trade-offs: refill the RLHF (PPO) column from what you know. The rest of the table is as it appeared.

dimensionRLHF (PPO)DPO
componentsSFT + RM + policy + value netSFT + policy (pi_ref frozen)
training stabilitycomplex: RM updates, PPO clip, KLsimple: single SL loss
online vs offlineonline (samples during training)offline (preference dataset fixed)
computehigh — runs 4 models simultaneouslylow — equivalent to SFT
reward hackingpossible (RM can be gamed)implicit RM cannot diverge past pi_ref
use todayGPT-4, Gemini Ultramany open-source models (Zephyr, Tulu-3)

40. Without one step: Recipe: RLHF or DPO on a new task

Constraint

Discussion prompt

Run Recipe: RLHF or DPO on a new task with this step confiscated:

RLHF path: Train RM with Bradley-Terry loss: L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value network as the PPO critic.

Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.

Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.

Answer:

  1. Collect demonstrations (prompt → ideal response pairs). Fine-tune the pretrained model with cross-entropy on response tokens only. This is π_SFT.
  2. Collect preference pairs (prompt → (y_w, y_l)): for each prompt, generate or write two completions; have annotators rank them.
  3. RLHF path: Train RM with Bradley-Terry loss: L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value…
  4. DPO path (simpler): Directly fine-tune π_SFT with L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.
  5. Evaluate: human evals on helpfulness, harmlessness, honesty; automated win-rate vs SFT baseline; safety benchmarks. Monitor KL vs SFT to detect drift.

41. Recipe: RLHF or DPO on a new task

Pattern

  1. Collect demonstrations (prompt → ideal response pairs). Fine-tune the pretrained model with cross-entropy on response tokens only. This is π_SFT.
  2. Collect preference pairs (prompt → (y_w, y_l)): for each prompt, generate or write two completions; have annotators rank them.
  3. RLHF path: Train RM with Bradley-Terry loss: L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value network as the PPO critic.
  4. DPO path (simpler): Directly fine-tune π_SFT with L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.
  5. Evaluate: human evals on helpfulness, harmlessness, honesty; automated win-rate vs SFT baseline; safety benchmarks. Monitor KL vs SFT to detect drift.

42. Where does it stop working: Recipe: RLHF or DPO on a new task

Edge cases

Discussion prompt

Recipe: RLHF or DPO on a new task works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

  1. Collect demonstrations (prompt → ideal response pairs). Fine-tune the pretrained model with cross-entropy on response tokens only. This is π_SFT.
  2. Collect preference pairs (prompt → (y_w, y_l)): for each prompt, generate or write two completions; have annotators rank them.
  3. RLHF path: Train RM with Bradley-Terry loss: L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value…
  4. DPO path (simpler): Directly fine-tune π_SFT with L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.
  5. Evaluate: human evals on helpfulness, harmlessness, honesty; automated win-rate vs SFT baseline; safety benchmarks. Monitor KL vs SFT to detect drift.

43. Rule out three: Check 1: RLHF pipeline stage order

Elimination

Eliminate the wrong options

In the standard 3-step RLHF pipeline, which model is used as the frozen reference during the PPO stage?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. The reward model r_phi
  • B. The original pretrained LM before any fine-tuning
  • C. The SFT model pi_SFT
  • D. The value network (PPO critic)

Survives elimination: C

Why: The SFT model pi_SFT is frozen and used to compute the KL penalty KL(pi_theta || pi_SFT) during PPO. It anchors the policy near the instruction-following baseline, preventing reward hacking.

44. Check 1: RLHF pipeline stage order

Check

Answer from the pipeline structure — no calculation needed.

Check your understanding

In the standard 3-step RLHF pipeline, which model is used as the frozen reference during the PPO stage?

  • A. The reward model r_phi
  • B. The original pretrained LM before any fine-tuning
  • C. The SFT model pi_SFT (correct)
  • D. The value network (PPO critic)

Answer: C

Why: The SFT model pi_SFT is frozen and used to compute the KL penalty KL(pi_theta || pi_SFT) during PPO. It anchors the policy near the instruction-following baseline, preventing reward hacking.

Why A tempts people
The reward model r_phi provides the environment reward signal, but it is not the KL reference — it is a separate frozen model queried for scalar scores.
Why B tempts people
The raw pretrained LM is not used after SFT. Using it as the reference would allow the policy to drift far from instruction-following behavior before the KL kicks in.
Why D tempts people
The value network (critic) estimates expected future reward for PPO's advantage computation — it is not the reference distribution for the KL divergence term.

45. Answer it before you see the options: Check 2: Bradley-Terry reward model loss

Prediction

Predict first

A reward model outputs r(x, y_w) = 1.2 and r(x, y_l) = 0.7 for a preference pair. What is the Bradley-Terry loss for this example (to 3 decimal places)?

Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.

Correct: -log σ(0.5) ≈ 0.474

Why: margin = r(y_w) - r(y_l) = 1.2 - 0.7 = 0.5. Loss = -log σ(0.5) = -log(1/(1+e^{-0.5})) ≈ 0.474. The model is partially correct — loss is positive and would decrease as the margin grows.

46. Check 2: Bradley-Terry reward model loss

Check

Work through the loss calculation before selecting an answer.

Check your understanding

A reward model outputs r(x, y_w) = 1.2 and r(x, y_l) = 0.7 for a preference pair. What is the Bradley-Terry loss for this example (to 3 decimal places)?

  • A. -log σ(0.5) ≈ 0.474 (correct)
  • B. -log σ(1.9) ≈ 0.135
  • C. log σ(0.5) ≈ -0.474
  • D. -log σ(-0.5) ≈ 0.974

Answer: A

Why: margin = r(y_w) - r(y_l) = 1.2 - 0.7 = 0.5. Loss = -log σ(0.5) = -log(1/(1+e^{-0.5})) ≈ 0.474. The model is partially correct — loss is positive and would decrease as the margin grows.

Why B tempts people
The margin 1.9 would arise if both scores were summed (1.2 + 0.7) rather than subtracted. The Bradley-Terry model uses the difference, not the sum.
Why C tempts people
The sign is wrong — the RM loss is the negative log-sigmoid of the margin, making it a positive quantity that gradient descent minimizes.
Why D tempts people
A margin of -0.5 would mean r(y_l) > r(y_w), which is not the case here. This would correspond to the RM ranking the completions backwards.

47. Rule out three: Check 3: DPO implicit reward

Elimination

Eliminate the wrong options

With beta=0.1, a trained policy has log pi_theta(y_w|x) - log pi_ref(y_w|x) = 3.0 and log pi_theta(y_l|x) - log pi_ref(y_l|x) = -1.0. Which statement is correct?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. DPO loss margin = beta*(3.0 - (-1.0)) = 0.4; the policy correctly ranks y_w above y_l
  • B. DPO loss margin = 3.0 - (-1.0) = 4.0; beta is applied only to the chosen completion
  • C. DPO loss margin = beta*(3.0 + (-1.0)) = 0.2; the two log-ratios are summed
  • D. DPO loss margin = beta*(3.0 - (-1.0)) = 0.4; but the policy still minimizes loss when y_l is ranked higher

Survives elimination: A

Why: margin = beta(ratio_w - ratio_l) = 0.1(3.0 - (-1.0)) = 0.4. A positive margin means sigma(0.4) > 0.5, so -log sigma(0.4) < log 2 ≈ 0.693 — the policy is correctly placing y_w above y_l and the loss is below the random-init baseline.

48. Check 3: DPO implicit reward

Check

Use the DPO implicit reward formula. beta = 0.1.

Check your understanding

With beta=0.1, a trained policy has log pi_theta(y_w|x) - log pi_ref(y_w|x) = 3.0 and log pi_theta(y_l|x) - log pi_ref(y_l|x) = -1.0. Which statement is correct?

  • A. DPO loss margin = beta*(3.0 - (-1.0)) = 0.4; the policy correctly ranks y_w above y_l (correct)
  • B. DPO loss margin = 3.0 - (-1.0) = 4.0; beta is applied only to the chosen completion
  • C. DPO loss margin = beta*(3.0 + (-1.0)) = 0.2; the two log-ratios are summed
  • D. DPO loss margin = beta*(3.0 - (-1.0)) = 0.4; but the policy still minimizes loss when y_l is ranked higher

Answer: A

Why: margin = beta(ratio_w - ratio_l) = 0.1(3.0 - (-1.0)) = 0.4. A positive margin means sigma(0.4) > 0.5, so -log sigma(0.4) < log 2 ≈ 0.693 — the policy is correctly placing y_w above y_l and the loss is below the random-init baseline.

Why B tempts people
beta multiplies the difference of both log-ratios, not just the chosen one. Dropping beta inflates the margin by 10x and gives the wrong scale.
Why C tempts people
The DPO formula subtracts the rejected ratio from the chosen ratio: (ratio_w - ratio_l). Summing them (3.0 + (-1.0) = 2.0) confuses the sign of the rejected term.
Why D tempts people
A positive margin means the policy already ranks y_w above y_l (imp_w = 0.3 > imp_l = -0.1). Loss is minimized when the margin is large and positive, not when y_l is ranked higher.

49. Your turn: build a toy DPO training step

Worked example

Task: implement one DPO gradient step on a single preference pair. Use the tiny model from the SFT slide (vocab=8, d=8, seed=42). One prompt x = [0,1,2], chosen y_w = [3,4,5], rejected y_l = [2,2,2]. beta=0.1.

  1. Milestone 1: Initialize pi_theta (embed + head) and freeze a copy as pi_ref.
  2. Milestone 2: Compute sequence-level log-probs — F.log_softmax then index the target tokens.
  3. Milestone 3: Compute ratio_w = sum(log pi_theta(y_w)) - sum(log pi_ref(y_w)) and ratio_l similarly.
  4. Milestone 4: loss = -F.logsigmoid(beta * (ratio_w - ratio_l)); backprop and step.
import torch, torch.nn as nn, torch.nn.functional as F, copy
torch.manual_seed(42)
# tiny model
embed = nn.Embedding(8, 8); head = nn.Linear(8, 8)
pi_theta = lambda x: head(embed(x))
# freeze reference
embed_r = copy.deepcopy(embed); head_r = copy.deepcopy(head)
for p in list(embed_r.parameters()) + list(head_r.parameters()):
    p.requires_grad_(False)
pi_ref = lambda x: head_r(embed_r(x))
# data
x   = torch.tensor([[0, 1, 2]])       # prompt (unused here — concat model)
y_w = torch.tensor([[3, 4, 5]])
y_l = torch.tensor([[2, 2, 2]])
def seq_logprob(policy, inp, tgt):
    logits = policy(inp)               # (1, T, V)
    lp = F.log_softmax(logits, dim=-1)
    return lp.gather(2, tgt.unsqueeze(2)).squeeze(2).sum(dim=1)  # (1,)
rw_theta = seq_logprob(pi_theta, y_w, y_w)
rw_ref   = seq_logprob(pi_ref,   y_w, y_w)
rl_theta = seq_logprob(pi_theta, y_l, y_l)
rl_ref   = seq_logprob(pi_ref,   y_l, y_l)
beta = 0.1
loss = -F.logsigmoid(beta * ((rw_theta-rw_ref) - (rl_theta-rl_ref))).mean()
print(f'DPO loss (step 0): {loss.item():.4f}')
loss.backward()
print('Gradient computed successfully')
quantityexpected output
DPO loss (step 0)positive value < log(2) ≈ 0.693 or > 0.693 depending on random init
Gradient computedsuccessfully (no error = pi_ref frozen correctly)
After many stepsratio_w grows, ratio_l shrinks, loss approaches 0

Show it off: print ratio_w - ratio_l before and after 50 Adam steps. Confirm the margin grows and the loss falls, while pi_ref (frozen) outputs the same log-probs throughout.

50. Connect it up: Lesson 102: RLHF & DPO — Instruction Fine-Tuning

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — Why raw pretraining is not enough · Step 1: Supervised Fine-Tuning (SFT) · Step 2: Reward Model from Human Preferences · Step 3: PPO with KL Penalty · DPO: Direct Preference Optimization. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

51. Lesson 102 recap: RLHF & DPO

Recap

Sources

  1. USAAIO Year-Long Master Lesson Plan, Lesson 102 — Instruction Fine-Tuning & RLHF — Barron · USAAIO Round 2 Preparation, 2026
  2. Ouyang et al. 'Training language models to follow instructions with human feedback' (InstructGPT, NeurIPS 2022) — arXiv:2203.02155
  3. Rafailov et al. 'Direct Preference Optimization: Your Language Model is Secretly a Reward Model' (NeurIPS 2023) — arXiv:2305.18290
  4. SFT loss=2.7608/2.6412/2.5241, RM Bradley-Terry loss=1.1683, KL_mean=0.7017, shaped_reward=0.7798, DPO loss=0.6539 all verified with torch 2.7.1+cpu seed=42, June 2026 — Real execution, verified

Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108