All articles Reinforcement Learning

RLVR & Reasoning Alignment: Reinforcement Learning with Verifiable Rewards

How RLVR replaces subjective reward models with objective code compilers and math solvers: Group Relative Policy Optimization (GRPO) and self-play.

Traditional Reinforcement Learning from Human Feedback (RLHF) relies on a neural Reward Model trained on human preference ratings. However, preference reward models suffer from reward hacking, length bias, and noise. Reinforcement Learning with Verifiable Rewards (RLVR) — popularized by DeepSeek R1 and OpenAI o1 — bypasses learned reward models entirely for domain problems with objective ground truth (such as competitive mathematics, formal logic, and software code compilation).

RLHF learns from human opinions; RLVR learns from execution truth: python compilers, math solvers, and unit tests.

Why subjective reward models fail on reasoning tasks

When a language model outputs a 2,000-token proof for a mathematical theorem, human annotators and neural reward models struggle to detect subtle algebraic errors at step 14. As a result, RLHF pushes the model to write authoritative-sounding, verbose answers rather than mathematically correct ones.

In contrast, RLVR evaluates outputs against a deterministic binary or scalar verification oracle:

  • Code Execution: Does the code pass 100% of unit test cases in a sandboxed Docker execution environment? (+1 or 0)
  • Mathematical Equality: Does the parsed SymPy expression evaluate to exact identity with the target answer? (+1 or 0)
  • Form Format Verification: Does the output properly wrap internal reasoning in <think> ... </think> tags? (+0.1 structural reward)

Group Relative Policy Optimization (GRPO) explained

Standard PPO requires a Critic model of equal size to the Actor to estimate state values V(s). GRPO (Shao et al., 2024 / DeepSeek Math) eliminates the Critic model entirely by sampling a group of G candidate outputs {y_1, y_2, ..., y_G} for a single prompt x and computing relative advantages across the group:

Advantage for completion i:
A_i = ( R_i - mean(R_1..R_G) ) / std(R_1..R_G)

The GRPO objective maximizes policy likelihood scaled by group-relative advantages, constrained by KL divergence against the initial reference model:

L_GRPO = (1 / G) * sum_i [ min( (pi_theta / pi_old) * A_i, clip(pi_theta / pi_old, 1-eps, 1+eps) * A_i ) - beta * KL(pi_theta || pi_ref) ]

Python implementation of GRPO Advantage Normalization

import torch

def compute_grpo_advantages(rewards: torch.Tensor, eps: float = 1e-8) -> torch.Tensor:
    """
    Computes Group Relative Policy Optimization (GRPO) relative advantages.
    rewards shape: (batch_size, group_size)
    """
    group_mean = rewards.mean(dim=-1, keepdim=True)
    group_std = rewards.std(dim=-1, keepdim=True)
    
    # Normalize rewards across the group
    advantages = (rewards - group_mean) / (group_std + eps)
    return advantages

# Example: Prompt with 4 sampled completions
# Rewards evaluated via Python code execution test suite (1.0 = All pass, 0.0 = Fail)
sample_rewards = torch.tensor([[1.0, 0.0, 1.0, 0.0],
                               [0.0, 0.0, 1.0, 0.0]])

advantages = compute_grpo_advantages(sample_rewards)
print("GRPO Group Advantages:\n", advantages)

Key insights from RLVR research

  1. Self-Correction Emergence: Given pure RLVR compute without human prompts, models autonomously learn to pause, re-read earlier assumptions, and write self-correction phrases like "Wait, let me double check this step...".
  2. Zero Critic Memory Overhead: GRPO reduces GPU memory requirements by 50% compared to PPO by calculating baseline advantages directly from group sampling.
  3. Deterministic Feedback Loop: Grounding reward signals in verifiable execution prevents reward hacking and guarantees asymptotic convergence to true correctness.
← Back to all articles