USAAIO Lesson 102, from Phase 3. It covers the full RLHF pipeline, running from supervised fine-tuning through a reward model to PPO with a KL penalty, along with the Bradley-Terry pairwise loss and the shaped reward r_env − β·KL. It then presents DPO as a closed-form alternative that eliminates the reward model entirely. Every loss value, KL divergence, and DPO margin was verified with torch 2.7.1+cpu at seed 42. The lesson runs to 26 slides.
Subject: Machine Learning · 51 slides · code lesson
Open the interactive version of this deck · Homework for this lesson
Title
USAAIO · Lesson 102 · Phase 3
How raw pretrained LLMs become helpful assistants: supervised fine-tuning, reward modeling from human preferences, PPO with KL penalty, and DPO as a cleaner closed-form alternative. Every loss value derived and run.
Objectives
Warm-up
Discussion prompt
Before we open Lesson 102: RLHF & DPO — Instruction Fine-Tuning: without looking back, what was the main idea of LoRA — Low-Rank Adaptation, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
LoRA (Low-Rank Adaptation) for parameter-efficient fine-tuning — low-rank decomposition delta_W=AB, B=0 initialization, alpha/r scaling, which layers to apply LoRA to, and full LoRALinear nn.Module implementation.
Section
Part 1 of 4
Concept
A pretrained LLM minimizes next-token cross-entropy on internet text. It learns what language looks like, not what a helpful response looks like. Instruction fine-tuning bridges that gap.
| stage | data | objective | result |
|---|---|---|---|
| pretraining | raw internet text (trillions of tokens) | next-token cross-entropy | language model p(x) |
| SFT | human-written instruction→response pairs | next-token cross-entropy on response only | instruction follower |
| RLHF RM+RL | human pairwise preferences | maximize reward - KL | helpful & aligned |
SFT alone gives solid instruction following but cannot capture nuanced human preferences — it treats all demonstration tokens equally. RLHF uses human ranking signals to push the model further.
Comparison
Comparison matrix
From Pretraining vs instruction following: refill the data column from what you know. The rest of the table is as it appeared.
| stage | data | objective | result |
|---|---|---|---|
| pretraining | raw internet text (trillions of tokens) | next-token cross-entropy | language model p(x) |
| SFT | human-written instruction→response pairs | next-token cross-entropy on response only | instruction follower |
| RLHF RM+RL | human pairwise preferences | maximize reward - KL | helpful & aligned |
Concept
Each stage feeds the next: SFT initializes the policy and the RM; the RM provides the reward signal; PPO then optimizes the policy while the SFT model acts as a frozen reference for the KL term.
Counterexample
Discussion prompt
Each stage feeds the next: SFT initializes the policy and the RM; the RM provides the reward signal; PPO then optimizes the policy while the SFT model acts as a frozen reference for the KL term.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Section
Part 2 of 4
Pattern
Predict first
The table runs: 1 (init) | 2.7608 · 2 | 2.6412
In SFT: cross-entropy on the response only, given the rows so far: what is the next one — the row where step is 3?
Correct: 3 | 2.5241
| step | loss |
|---|---|
| 1 (init) | 2.7608 |
| 2 | 2.6412 |
| 3 | 2.5241 |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. We do NOT train on the prompt tokens; supervising on them teaches the model to copy prompts, not respond to them.
Worked example
Define the loss — minimize negative log-likelihood over response tokens only
Why: We do NOT train on the prompt tokens; supervising on them teaches the model to copy prompts, not respond to them.
\[ \mathcal{L}_{\text{SFT}}(\theta) = -\sum_{t=1}^{T} \log \pi_\theta(y_t \mid x, y_{<t}) \]
Run 3 gradient steps on a tiny 8-token LM (vocab=8, d=8, Adam lr=0.01, seed=42)
Why: Each step computes cross-entropy on next-token targets [1,2,3,4] given input [0,1,2,3], backprops through embedding + linear head.
import torch, torch.nn as nn, torch.nn.functional as F
torch.manual_seed(42)
embed = nn.Embedding(8, 8)
head = nn.Linear(8, 8)
opt = torch.optim.Adam(
list(embed.parameters()) + list(head.parameters()), lr=0.01)
toks = torch.tensor([[0, 1, 2, 3]])
targs = torch.tensor([[1, 2, 3, 4]])
for step in range(1, 4):
opt.zero_grad()
logits = head(embed(toks)) # (1, 4, 8)
loss = F.cross_entropy(logits.view(-1, 8), targs.view(-1))
loss.backward()
opt.step()
print(f'step {step}: loss={loss.item():.4f}')| step | loss |
|---|---|
| 1 (init) | 2.7608 |
| 2 | 2.6412 |
| 3 | 2.5241 |
Trade off
Comparison matrix
From SFT: cross-entropy on the response only: every row here is a choice with a cost. Fill the loss column, then say which row you would actually pick and what you give up for it.
| step | loss |
|---|---|
| 1 (init) | 2.7608 |
| 2 | 2.6412 |
| 3 | 2.5241 |
Section
Part 3 of 4
Concept
Annotators see a prompt x and two completions y_w (winner) and y_l (loser) and mark which they prefer. The reward model learns to assign higher scalar reward to y_w.
\[ p(y_w \succ y_l \mid x) = \sigma\bigl(r_\phi(x, y_w) - r_\phi(x, y_l)\bigr) \]
\[ \mathcal{L}_{\text{RM}}(\phi) = -\mathbb{E}_{(x,y_w,y_l)}\bigl[\log \sigma\bigl(r_\phi(x,y_w) - r_\phi(x,y_l)\bigr)\bigr] \]
This is binary cross-entropy on the margin between the two reward scores — the Bradley-Terry ranking model. It never trains on absolute reward values, only on which completion is relatively better.
Analogy
Discussion prompt
Explain Preference data and the Bradley-Terry model by analogy to something with no Machine Learning in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
Annotators see a prompt x and two completions y_w (winner) and y_l (loser) and mark which they prefer. The reward model learns to assign higher scalar reward to y_w.
Fill the middle
Fill in the blanks
From Reward model: pairwise loss computation — one line has had its right-hand side removed. Put it back.
class ToyRewardModel(nn.Module):
def __init__(self, vocab, d):
super().__init__()
self.embed = nn.Embedding(vocab, d)
self.head = nn.Linear(d, 1)
def forward(self, x):
return self.head(self.embed(x).mean(dim=1)).squeeze(-1)
rm = ToyRewardModel(vocab=8, d=16)
chosen = torch.tensor([[1, 3, 5, 6, 7]]) # higher quality
rejected = torch.tensor([[1, 2, 2, 2, 2]]) # lower quality
r_w = rm(chosen); r_l = rm(rejected)
loss_rm = -F.logsigmoid(r_w - r_l).mean()
print(f'r(chosen)=___ r(rejected)=___')
print(f'RM loss (Bradley-Terry)=___')
Why: chosen is what everything below it consumes, so the wrong expression here fails later and somewhere else. The RM architecture is the same as the SFT model but with a scalar head replacing the vocab projection; it can be initialized from the SFT checkpoint.
Worked example
Build a ToyRewardModel (embed → mean-pool → scalar) and forward both chosen and rejected completions
Why: The RM architecture is the same as the SFT model but with a scalar head replacing the vocab projection; it can be initialized from the SFT checkpoint.
class ToyRewardModel(nn.Module):
def __init__(self, vocab, d):
super().__init__()
self.embed = nn.Embedding(vocab, d)
self.head = nn.Linear(d, 1)
def forward(self, x):
return self.head(self.embed(x).mean(dim=1)).squeeze(-1)
rm = ToyRewardModel(vocab=8, d=16)
chosen = torch.tensor([[1, 3, 5, 6, 7]]) # higher quality
rejected = torch.tensor([[1, 2, 2, 2, 2]]) # lower quality
r_w = rm(chosen); r_l = rm(rejected)
loss_rm = -F.logsigmoid(r_w - r_l).mean()
print(f'r(chosen)={r_w.item():.4f} r(rejected)={r_l.item():.4f}')
print(f'RM loss (Bradley-Terry)={loss_rm.item():.4f}')| quantity | value | meaning |
|---|---|---|
| r(chosen) | -0.6170 | scalar reward for y_w at random init |
| r(rejected) | 0.1790 | scalar reward for y_l at random init |
| margin r_w - r_l | -0.7960 | negative: RM not yet trained |
| RM loss | 1.1683 | high — gradients push r_w > r_l |
Interpret: loss = 1.1683 because r(chosen) < r(rejected) at random init
Why: After training, the RM loss minimizes when r(y_w) >> r(y_l) for all preference pairs — the model has learned the reward signal from human labels.
Error analysis
Annotate
Walk the callouts on Reward model: pairwise loss computation. Each one is a place this is easy to get subtly wrong.
Anomaly
Predict first
A student writes this, and it looks reasonable:
Annotators assign numeric scores (1–10) to each completion separately. Train the RM with MSE: loss = MSE(r_phi(x,y), score).
It is wrong. Say what breaks — and say it before you turn the page.
Correct: Seems natural — match the model's scalar to the annotator's rating.
Collect pairwise comparisons (A vs B, which is better?). Train the RM with Bradley-Terry loss on the margin: loss = -log_sigmoid(r_w - r_l).
Why: Seems natural — match the model's scalar to the annotator's rating.
Trap
Annotators assign numeric scores (1–10) to each completion separately. Train the RM with MSE: loss = MSE(r_phi(x,y), score).
Compute loss = mean_squared_error(r_phi(completion), human_score)
Why: Seems natural — match the model's scalar to the annotator's rating.
Result: annotators are inconsistent and on different scales. The RM memorizes rater biases rather than learning which responses humans genuinely prefer.
Collect pairwise comparisons (A vs B, which is better?). Train the RM with Bradley-Terry loss on the margin: loss = -log_sigmoid(r_w - r_l).
Compute loss = -log sigma(r_phi(x, y_w) - r_phi(x, y_l)) for every (y_w, y_l) preference pair
Why: Relative comparison is far more consistent for humans than absolute scoring; the model learns what relatively better means, not a calibrated scale.
Pairwise preference data (A/B) is cheaper to collect and more reliable — this is why InstructGPT, Claude, and ChatGPT all use the pairwise paradigm.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Relative comparison is far more consistent for humans than absolute scoring; the model learns what relatively better means, not a calibrated scale.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Answer:
Seems natural — match the model's scalar to the annotator's rating.
Section
Part 4a of 4
Concept
After SFT and RM training, the policy π_θ (initialized from π_SFT) is optimized to maximize expected reward while staying close to the SFT reference distribution.
\[ \max_{\pi_\theta} \mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi_\theta(\cdot|x)} \Bigl[ r_\phi(x, y) - \beta \cdot D_{\text{KL}}\bigl(\pi_\theta(\cdot|x) \| \pi_{\text{SFT}}(\cdot|x)\bigr) \Bigr] \]
Without the KL term the policy can exploit the reward model (reward hacking): generate grammatically broken or degenerate text that happens to fool the RM into high scores.
Counterexample
Discussion prompt
After SFT and RM training, the policy π_θ (initialized from π_SFT) is optimized to maximize expected reward while staying close to the SFT reference distribution.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Answer:
Without the KL term the policy can exploit the reward model (reward hacking): generate grammatically broken or degenerate text that happens to fool the RM into high scores.
Fill the middle
Fill in the blanks
From PPO shaped reward: r_env − β·KL — one line has had its right-hand side removed. Put it back.
torch.manual_seed(7)
pol_logits = torch.randn(6, 8) # policy (6 tokens, vocab=8)
ref_logits = torch.randn(6, 8) # frozen SFT reference
pol_lp = F.log_softmax(pol_logits, dim=-1)
ref_lp = F.log_softmax(ref_logits, dim=-1)
# KL per token: sum over vocab
kl = (pol_lp.exp() * (pol_lp - ref_lp)).sum(dim=-1)
print('KL per token:', kl.detach().numpy().round(4))
print('Mean KL:', round(kl.mean().item(), 4))
r_env = 0.85 # reward from trained RM
beta = 0.1
shaped = r_env - beta * kl.mean().item()
print('shaped reward:', round(shaped, 4))
Why: ref_lp is what everything below it consumes, so the wrong expression here fails later and somewhere else. KL(pi || pi_ref) = sum_v pi(v|x) * (log pi(v|x) - log pi_ref(v|x)) measures how far the policy has drifted; summed/averaged over the response tokens.
Worked example
Compute per-token KL divergence between the current policy and the reference (SFT) distribution
Why: KL(pi || pi_ref) = sum_v pi(v|x) * (log pi(v|x) - log pi_ref(v|x)) measures how far the policy has drifted; summed/averaged over the response tokens.
torch.manual_seed(7)
pol_logits = torch.randn(6, 8) # policy (6 tokens, vocab=8)
ref_logits = torch.randn(6, 8) # frozen SFT reference
pol_lp = F.log_softmax(pol_logits, dim=-1)
ref_lp = F.log_softmax(ref_logits, dim=-1)
# KL per token: sum over vocab
kl = (pol_lp.exp() * (pol_lp - ref_lp)).sum(dim=-1)
print('KL per token:', kl.detach().numpy().round(4))
print('Mean KL:', round(kl.mean().item(), 4))
r_env = 0.85 # reward from trained RM
beta = 0.1
shaped = r_env - beta * kl.mean().item()
print('shaped reward:', round(shaped, 4))| token | KL(pi||pi_ref) |
|---|---|
| t=0 | 0.6565 |
| t=1 | 0.2939 |
| t=2 | 0.9556 |
| t=3 | 0.6660 |
| t=4 | 1.1052 |
| t=5 | 0.5329 |
| mean | 0.7017 |
| shaped reward (r_env=0.85, beta=0.1) | 0.7798 |
Interpret: shaped_reward = 0.85 − 0.1 × 0.7017 = 0.7798
Why: The KL penalty reduces the effective reward; if the policy drifted more (higher KL), the penalty grows and discourages further divergence from the SFT baseline.
Comparison
Comparison matrix
From PPO shaped reward: r_env − β·KL: refill the KL(pi||pi_ref) column from what you know. The rest of the table is as it appeared.
| token | KL(pi||pi_ref) |
|---|---|
| t=0 | 0.6565 |
| t=1 | 0.2939 |
| t=2 | 0.9556 |
| t=3 | 0.6660 |
| t=4 | 1.1052 |
| t=5 | 0.5329 |
| mean | 0.7017 |
| shaped reward (r_env=0.85, beta=0.1) | 0.7798 |
Concept
Vanilla REINFORCE has high variance and can take catastrophically large gradient steps. PPO clips the probability ratio to a trust region, preventing destructive updates.
\[ L^{\text{CLIP}}(\theta) = \mathbb{E}\Bigl[\min\bigl(r_t(\theta)\,\hat{A}_t,\;\text{clip}(r_t(\theta),\,1-\varepsilon,\,1+\varepsilon)\,\hat{A}_t\bigr)\Bigr] \]
| term | meaning |
|---|---|
| r_t(theta) = pi_theta(a_t|s_t) / pi_old(a_t|s_t) | probability ratio (how much policy changed) |
| A_hat_t | advantage estimate (reward − value baseline) |
| clip(..., 1-eps, 1+eps) | limits ratio to [0.8, 1.2] when eps=0.2 |
| min(unclipped, clipped) | pessimistic: no credit for ratio outside trust region |
Sorting
Sort into buckets
These are the pieces of Lesson 102: RLHF & DPO — Instruction Fine-Tuning, out of order. Put each one back under the part of the lesson it belongs to.
Section
Part 4b of 4
Concept
Rafailov et al. (2023) showed the optimal RLHF policy implies a closed-form relationship between the policy, the reference, and the reward. DPO inverts this to directly optimize the policy on preference data — no RM needed.
\[ \mathcal{L}_{\text{DPO}}(\theta) = -\mathbb{E}_{(x,y_w,y_l)} \Bigl[\log \sigma\Bigl(\beta \log\frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log\frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}\Bigr)\Bigr] \]
The term beta * log(pi/pi_ref) on a sequence y is the implicit reward the optimal RL policy assigns to y. DPO trains the policy to rank the chosen completion's implicit reward above the rejected one — exactly what the RM would do, but without a separate reward model.
Pattern
Predict first
The table runs: log pi(y_w|x) - log pi_ref(y_w|x) | 0.6000 | -4.2 - (-4.8) · log pi(y_l|x) - log pi_ref(y_l|x) | -0.2000 | -5.1 - (-4.9) · imp_w = beta * 0.6 | 0.0600 | implicit reward for chosen · imp_l = beta * (-0.2) | -0.0200 | implicit reward for rejected · margin = imp_w - imp_l | 0.0800 | beta * (ratio_w - ratio_l)
In DPO loss: computing implicit rewards and margin, given the rows so far: what is the next one — the row where quantity is DPO loss = -log sigma(0.08)?
Correct: DPO loss = -log sigma(0.08) | 0.6539 | positive: policy not yet trained
| quantity | value | formula |
|---|---|---|
| log pi(y_w|x) - log pi_ref(y_w|x) | 0.6000 | -4.2 - (-4.8) |
| log pi(y_l|x) - log pi_ref(y_l|x) | -0.2000 | -5.1 - (-4.9) |
| imp_w = beta * 0.6 | 0.0600 | implicit reward for chosen |
| imp_l = beta * (-0.2) | -0.0200 | implicit reward for rejected |
| margin = imp_w - imp_l | 0.0800 | beta * (ratio_w - ratio_l) |
| DPO loss = -log sigma(0.08) | 0.6539 | positive: policy not yet trained |
Why: The relationship between the columns, not the individual numbers, is what generates the next row. Sum the per-token log-probs under both the policy and the reference, then take the difference.
Worked example
Compute sequence-level log-prob ratios log(pi_theta/pi_ref) for chosen and rejected completions
Why: Sum the per-token log-probs under both the policy and the reference, then take the difference. This is the implicit reward for each completion.
# Toy sequence-level log-probs (pretend these are summed over tokens)
pi_chosen_lp = -4.2; ref_chosen_lp = -4.8 # log p for y_w
pi_rej_lp = -5.1; ref_rej_lp = -4.9 # log p for y_l
beta = 0.1
# implicit rewards
imp_w = beta * (pi_chosen_lp - ref_chosen_lp) # 0.1 * 0.6
imp_l = beta * (pi_rej_lp - ref_rej_lp) # 0.1 * (-0.2)
margin = imp_w - imp_l
loss_dpo = -float(torch.tensor(margin).sigmoid().log())
print(f'imp_w={imp_w:.4f} imp_l={imp_l:.4f}')
print(f'margin={margin:.4f} DPO_loss={loss_dpo:.4f}')| quantity | value | formula |
|---|---|---|
| log pi(y_w|x) - log pi_ref(y_w|x) | 0.6000 | -4.2 - (-4.8) |
| log pi(y_l|x) - log pi_ref(y_l|x) | -0.2000 | -5.1 - (-4.9) |
| imp_w = beta * 0.6 | 0.0600 | implicit reward for chosen |
| imp_l = beta * (-0.2) | -0.0200 | implicit reward for rejected |
| margin = imp_w - imp_l | 0.0800 | beta * (ratio_w - ratio_l) |
| DPO loss = -log sigma(0.08) | 0.6539 | positive: policy not yet trained |
Interpret: DPO loss decreases as imp_w grows and imp_l shrinks — i.e. the policy up-weights chosen and down-weights rejected relative to the reference
Why: The gradient of DPO loss pushes the policy to increase log prob of y_w and decrease log prob of y_l, while the pi_ref denominator anchors the magnitude — β controls the trade-off strength.
Trade off
Comparison matrix
From DPO loss: computing implicit rewards and margin: every row here is a choice with a cost. Fill the value column, then say which row you would actually pick and what you give up for it.
| quantity | value | formula |
|---|---|---|
| log pi(y_w|x) - log pi_ref(y_w|x) | 0.6000 | -4.2 - (-4.8) |
| log pi(y_l|x) - log pi_ref(y_l|x) | -0.2000 | -5.1 - (-4.9) |
| imp_w = beta * 0.6 | 0.0600 | implicit reward for chosen |
| imp_l = beta * (-0.2) | -0.0200 | implicit reward for rejected |
| margin = imp_w - imp_l | 0.0800 | beta * (ratio_w - ratio_l) |
| DPO loss = -log sigma(0.08) | 0.6539 | positive: policy not yet trained |
Anomaly
Predict first
A student writes this, and it looks reasonable:
After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.
It is wrong. Say what breaks — and say it before you turn the page.
Correct: The DPO objective is framed around the ratio, so the reference must stay active.
The reference model pi_ref is needed only during DPO training to compute the implicit reward ratios. At inference, just sample from pi_theta normally.
Why: The DPO objective is framed around the ratio, so the reference must stay active.
Trap
After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.
At serving time, compute beta * (log pi_theta(y|x) - log pi_ref(y|x)) for each candidate response before returning the best one
Why: The DPO objective is framed around the ratio, so the reference must stay active.
The reference model pi_ref is needed only during DPO training to compute the implicit reward ratios. At inference, just sample from pi_theta normally.
At serving time, use pi_theta.generate(prompt) with no reference model
Why: DPO bakes the preference signal into the policy weights during training. The trained pi_theta already reflects the human preferences — pi_ref is discarded after training is complete.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
loss = MSE(r_phi(x,y), score).; After training with DPO, keep both pi_theta and pi_ref loaded at inference time to compute the implicit reward before each generation.Concept
| dimension | RLHF (PPO) | DPO |
|---|---|---|
| components | SFT + RM + policy + value net | SFT + policy (pi_ref frozen) |
| training stability | complex: RM updates, PPO clip, KL | simple: single SL loss |
| online vs offline | online (samples during training) | offline (preference dataset fixed) |
| compute | high — runs 4 models simultaneously | low — equivalent to SFT |
| reward hacking | possible (RM can be gamed) | implicit RM cannot diverge past pi_ref |
| use today | GPT-4, Gemini Ultra | many open-source models (Zephyr, Tulu-3) |
DPO is preferred when compute is limited or training stability is a concern. RLHF with PPO allows online exploration — the policy can discover new behaviors during training — which matters at frontier scale.
Comparison
Comparison matrix
From RLHF vs DPO: trade-offs: refill the RLHF (PPO) column from what you know. The rest of the table is as it appeared.
| dimension | RLHF (PPO) | DPO |
|---|---|---|
| components | SFT + RM + policy + value net | SFT + policy (pi_ref frozen) |
| training stability | complex: RM updates, PPO clip, KL | simple: single SL loss |
| online vs offline | online (samples during training) | offline (preference dataset fixed) |
| compute | high — runs 4 models simultaneously | low — equivalent to SFT |
| reward hacking | possible (RM can be gamed) | implicit RM cannot diverge past pi_ref |
| use today | GPT-4, Gemini Ultra | many open-source models (Zephyr, Tulu-3) |
Constraint
Discussion prompt
Run Recipe: RLHF or DPO on a new task with this step confiscated:
RLHF path: Train RM with Bradley-Terry loss: L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value network as the PPO critic.
Is it still possible? If it is, say what takes its place and what it costs you. If it is not, say exactly what that step was providing that nothing else does.
Hint: A step you can drop for free was never load-bearing. If you cannot drop it, name the thing that goes wrong the moment it is gone.
Answer:
L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value…L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.Pattern
L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value network as the PPO critic.L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.Edge cases
Discussion prompt
Recipe: RLHF or DPO on a new task works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".
Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.
Answer:
L = -log σ(r(x,y_w) − r(x,y_l)). Then run PPO maximizing E[r(x,y) − β·KL(π||π_SFT)], loading a value…L_DPO = -log σ(β(log π/π_ref on y_w) − β(log π/π_ref on y_l)). No separate RM or value network.Elimination
Eliminate the wrong options
In the standard 3-step RLHF pipeline, which model is used as the frozen reference during the PPO stage?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: C
Why: The SFT model pi_SFT is frozen and used to compute the KL penalty KL(pi_theta || pi_SFT) during PPO. It anchors the policy near the instruction-following baseline, preventing reward hacking.
Check
Answer from the pipeline structure — no calculation needed.
Check your understanding
In the standard 3-step RLHF pipeline, which model is used as the frozen reference during the PPO stage?
Answer: C
Why: The SFT model pi_SFT is frozen and used to compute the KL penalty KL(pi_theta || pi_SFT) during PPO. It anchors the policy near the instruction-following baseline, preventing reward hacking.
Prediction
Predict first
A reward model outputs r(x, y_w) = 1.2 and r(x, y_l) = 0.7 for a preference pair. What is the Bradley-Terry loss for this example (to 3 decimal places)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: -log σ(0.5) ≈ 0.474
Why: margin = r(y_w) - r(y_l) = 1.2 - 0.7 = 0.5. Loss = -log σ(0.5) = -log(1/(1+e^{-0.5})) ≈ 0.474. The model is partially correct — loss is positive and would decrease as the margin grows.
Check
Work through the loss calculation before selecting an answer.
Check your understanding
A reward model outputs r(x, y_w) = 1.2 and r(x, y_l) = 0.7 for a preference pair. What is the Bradley-Terry loss for this example (to 3 decimal places)?
Answer: A
Why: margin = r(y_w) - r(y_l) = 1.2 - 0.7 = 0.5. Loss = -log σ(0.5) = -log(1/(1+e^{-0.5})) ≈ 0.474. The model is partially correct — loss is positive and would decrease as the margin grows.
Elimination
Eliminate the wrong options
With beta=0.1, a trained policy has log pi_theta(y_w|x) - log pi_ref(y_w|x) = 3.0 and log pi_theta(y_l|x) - log pi_ref(y_l|x) = -1.0. Which statement is correct?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: margin = beta(ratio_w - ratio_l) = 0.1(3.0 - (-1.0)) = 0.4. A positive margin means sigma(0.4) > 0.5, so -log sigma(0.4) < log 2 ≈ 0.693 — the policy is correctly placing y_w above y_l and the loss is below the random-init baseline.
Check
Use the DPO implicit reward formula. beta = 0.1.
Check your understanding
With beta=0.1, a trained policy has log pi_theta(y_w|x) - log pi_ref(y_w|x) = 3.0 and log pi_theta(y_l|x) - log pi_ref(y_l|x) = -1.0. Which statement is correct?
Answer: A
Why: margin = beta(ratio_w - ratio_l) = 0.1(3.0 - (-1.0)) = 0.4. A positive margin means sigma(0.4) > 0.5, so -log sigma(0.4) < log 2 ≈ 0.693 — the policy is correctly placing y_w above y_l and the loss is below the random-init baseline.
Worked example
Task: implement one DPO gradient step on a single preference pair. Use the tiny model from the SFT slide (vocab=8, d=8, seed=42). One prompt x = [0,1,2], chosen y_w = [3,4,5], rejected y_l = [2,2,2]. beta=0.1.
F.log_softmax then index the target tokens.loss = -F.logsigmoid(beta * (ratio_w - ratio_l)); backprop and step.import torch, torch.nn as nn, torch.nn.functional as F, copy
torch.manual_seed(42)
# tiny model
embed = nn.Embedding(8, 8); head = nn.Linear(8, 8)
pi_theta = lambda x: head(embed(x))
# freeze reference
embed_r = copy.deepcopy(embed); head_r = copy.deepcopy(head)
for p in list(embed_r.parameters()) + list(head_r.parameters()):
p.requires_grad_(False)
pi_ref = lambda x: head_r(embed_r(x))
# data
x = torch.tensor([[0, 1, 2]]) # prompt (unused here — concat model)
y_w = torch.tensor([[3, 4, 5]])
y_l = torch.tensor([[2, 2, 2]])
def seq_logprob(policy, inp, tgt):
logits = policy(inp) # (1, T, V)
lp = F.log_softmax(logits, dim=-1)
return lp.gather(2, tgt.unsqueeze(2)).squeeze(2).sum(dim=1) # (1,)
rw_theta = seq_logprob(pi_theta, y_w, y_w)
rw_ref = seq_logprob(pi_ref, y_w, y_w)
rl_theta = seq_logprob(pi_theta, y_l, y_l)
rl_ref = seq_logprob(pi_ref, y_l, y_l)
beta = 0.1
loss = -F.logsigmoid(beta * ((rw_theta-rw_ref) - (rl_theta-rl_ref))).mean()
print(f'DPO loss (step 0): {loss.item():.4f}')
loss.backward()
print('Gradient computed successfully')| quantity | expected output |
|---|---|
| DPO loss (step 0) | positive value < log(2) ≈ 0.693 or > 0.693 depending on random init |
| Gradient computed | successfully (no error = pi_ref frozen correctly) |
| After many steps | ratio_w grows, ratio_l shrinks, loss approaches 0 |
Show it off: print ratio_w - ratio_l before and after 50 Adam steps. Confirm the margin grows and the loss falls, while pi_ref (frozen) outputs the same log-probs throughout.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Why raw pretraining is not enough · Step 1: Supervised Fine-Tuning (SFT) · Step 2: Reward Model from Human Preferences · Step 3: PPO with KL Penalty · DPO: Direct Preference Optimization. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
−log σ(r(x,y_w) − r(x,y_l)) — trains on relative rankings, not absolute scores.−log σ(β·(log π/π_ref on y_w) − β·(log π/π_ref on y_l)) — implicit reward replaces the explicit RM; pi_ref discarded at inference.Want this taught 1-on-1? Alexander tutors Machine Learning — $55/session, free consultation.