All articles LLM Alignment

Direct Preference Optimization (DPO) & Alignment: Beyond RLHF with PPO

A rigorous mathematical breakdown of LLM alignment techniques: Bradley-Terry preference modeling, PPO loss dynamics, DPO implicit reward derivation, KTO, and ORPO.

Aligning Large Language Models (LLMs) with human intent historically relied on Reinforcement Learning from Human Feedback (RLHF) using Proximal Policy Optimization (PPO). However, PPO pipeline setups are notoriously unstable — requiring four separate models in GPU memory simultaneously (Policy, Reference, Reward, and Value Critic). Direct Preference Optimization (DPO) simplifies preference learning by mathematically reparameterizing the reward function directly in terms of the policy, eliminating the need for a separate reward model or RL critic loop.

RLHF with PPO optimizes a policy against an explicit scalar reward model; DPO proves that the optimal policy can be derived directly from preference pairs via an implicit reward function.

The classical RLHF / PPO objective

In standard RLHF, given a prompt x and model response y, we train a Reward Model r_theta(x, y) on pairwise human preferences (y_w > y_l) using the Bradley-Terry preference model:

P(y_w > y_l | x) = sigmoid( r_theta(x, y_w) - r_theta(x, y_l) )

Then, PPO optimizes the policy pi_theta(y|x) against r_theta with a KL-divergence penalty relative to the reference SFT model pi_ref:

max_pi E [ r_theta(x, y) - beta * KL( pi_theta(y|x) || pi_ref(y|x) ) ]

Mathematical derivation of DPO

Rafailov et al. (2023) demonstrated that the exact closed-form solution for the optimal policy pi_* in the RLHF objective can be expressed as:

pi_*(y|x) = (1 / Z(x)) * pi_ref(y|x) * exp( (1 / beta) * r(x, y) )

Rearranging this equation yields the exact relationship between the implicit ground-truth reward r(x, y) and the policy ratios:

r(x, y) = beta * log( pi_theta(y|x) / pi_ref(y|x) ) + beta * log( Z(x) )

Substituting this implicit reward expression directly into the Bradley-Terry preference probability cancels out the intractable partition function Z(x)!

The DPO loss function explained

The resulting DPO Loss Function evaluates directly over preferred responses y_w and dispreferred responses y_l:

L_DPO(pi_theta; pi_ref) = - E_(x, y_w, y_l) [ log sigmoid( beta * log( pi_theta(y_w|x) / pi_ref(y_w|x) ) - beta * log( pi_theta(y_l|x) / pi_ref(y_l|x) ) ) ]

Notice how DPO acts as an adaptive gradient force:

  • If the reference model already assigns high probability to y_w relative to y_l, the gradient update is small.
  • If the policy model degrades on y_w or increases probability on dispreferred response y_l, the loss penalty spikes heavily.

Beyond DPO: KTO, ORPO & SimPO

  • KTO (Kahneman-Tversky Optimization): Does not require paired preference data (y_w, y_l). Learns from single binary signals (thumbs up / thumbs down) based on prospect theory.
  • ORPO (Odds Ratio Preference Optimization): Combines SFT loss and preference alignment into a single training step, eliminating the need for a reference model pi_ref altogether.
  • SimPO (Simple Preference Optimization): Replaces the reference model log ratio with an explicit length-normalized reward margin.

Python implementation with PyTorch & TRL

import torch
import torch.nn.functional as F

def compute_dpo_loss(
    policy_chosen_logps: torch.FloatTensor,
    policy_rejected_logps: torch.FloatTensor,
    ref_chosen_logps: torch.FloatTensor,
    ref_rejected_logps: torch.FloatTensor,
    beta: float = 0.1
) -> torch.FloatTensor:
    """
    Computes exact Direct Preference Optimization (DPO) Loss.
    """
    # Compute log ratios between policy and reference model
    pi_logratios = policy_chosen_logps - policy_rejected_logps
    ref_logratios = ref_chosen_logps - ref_rejected_logps
    
    # Calculate implicit reward delta
    logits = pi_logratios - ref_logratios
    
    # Negative log-sigmoid DPO loss
    losses = -F.logsigmoid(beta * logits)
    return losses.mean()

# Example log probability calculation
policy_chosen = torch.tensor([-1.2, -0.8])
policy_rejected = torch.tensor([-3.4, -4.1])
ref_chosen = torch.tensor([-1.5, -1.0])
ref_rejected = torch.tensor([-2.8, -3.0])

loss = compute_dpo_loss(policy_chosen, policy_rejected, ref_chosen, ref_rejected, beta=0.1)
print(f"DPO Loss: {loss.item():.4f}")

Engineering best practices for preference tuning

  1. Beta Hyperparameter Tuning: Keep beta between 0.01 and 0.1. Larger values strictly enforce reference model constraint; smaller values risk policy collapse.
  2. Data Cleanliness Over Volume: Filter out tie pairs where chosen and rejected completions are almost identical in quality.
  3. Use Length Normalization: Normalize log probabilities by token length to prevent DPO from exploiting response length bias.
← Back to all articles