Traditional Reinforcement Learning from Human Feedback (RLHF) relies on a neural Reward Model trained on human preference ratings. However, preference reward models suffer from reward hacking, length bias, and noise. Reinforcement Learning with Verifiable Rewards (RLVR) — popularized by DeepSeek R1 and OpenAI o1 — bypasses learned reward models entirely for domain problems with objective ground truth (such as competitive mathematics, formal logic, and software code compilation).
RLHF learns from human opinions; RLVR learns from execution truth: python compilers, math solvers, and unit tests.
Why subjective reward models fail on reasoning tasks
When a language model outputs a 2,000-token proof for a mathematical theorem, human annotators and neural reward models struggle to detect subtle algebraic errors at step 14. As a result, RLHF pushes the model to write authoritative-sounding, verbose answers rather than mathematically correct ones.
In contrast, RLVR evaluates outputs against a deterministic binary or scalar verification oracle:
- Code Execution: Does the code pass 100% of unit test cases in a sandboxed Docker execution environment? (+1 or 0)
- Mathematical Equality: Does the parsed SymPy expression evaluate to exact identity with the target answer? (+1 or 0)
- Form Format Verification: Does the output properly wrap internal reasoning in
<think> ... </think>tags? (+0.1 structural reward)
Group Relative Policy Optimization (GRPO) explained
Standard PPO requires a Critic model of equal size to the Actor to estimate state values V(s). GRPO (Shao et al., 2024 / DeepSeek Math) eliminates the Critic model entirely by sampling a group of G candidate outputs {y_1, y_2, ..., y_G} for a single prompt x and computing relative advantages across the group:
Advantage for completion i:
A_i = ( R_i - mean(R_1..R_G) ) / std(R_1..R_G)
The GRPO objective maximizes policy likelihood scaled by group-relative advantages, constrained by KL divergence against the initial reference model:
L_GRPO = (1 / G) * sum_i [ min( (pi_theta / pi_old) * A_i, clip(pi_theta / pi_old, 1-eps, 1+eps) * A_i ) - beta * KL(pi_theta || pi_ref) ]
Python implementation of GRPO Advantage Normalization
import torch
def compute_grpo_advantages(rewards: torch.Tensor, eps: float = 1e-8) -> torch.Tensor:
"""
Computes Group Relative Policy Optimization (GRPO) relative advantages.
rewards shape: (batch_size, group_size)
"""
group_mean = rewards.mean(dim=-1, keepdim=True)
group_std = rewards.std(dim=-1, keepdim=True)
# Normalize rewards across the group
advantages = (rewards - group_mean) / (group_std + eps)
return advantages
# Example: Prompt with 4 sampled completions
# Rewards evaluated via Python code execution test suite (1.0 = All pass, 0.0 = Fail)
sample_rewards = torch.tensor([[1.0, 0.0, 1.0, 0.0],
[0.0, 0.0, 1.0, 0.0]])
advantages = compute_grpo_advantages(sample_rewards)
print("GRPO Group Advantages:\n", advantages)
Key insights from RLVR research
- Self-Correction Emergence: Given pure RLVR compute without human prompts, models autonomously learn to pause, re-read earlier assumptions, and write self-correction phrases like "Wait, let me double check this step...".
- Zero Critic Memory Overhead: GRPO reduces GPU memory requirements by 50% compared to PPO by calculating baseline advantages directly from group sampling.
- Deterministic Feedback Loop: Grounding reward signals in verifiable execution prevents reward hacking and guarantees asymptotic convergence to true correctness.