Generative modeling has undergone a paradigm shift from Autoregressive models and Generative Adversarial Networks (GANs) toward Continuous-Time Diffusion Models and Rectified Flow Matching (powering architectures like Flux.1 and SD3). By formulating image synthesis as learning a vector field that transports a simple Gaussian noise distribution to a complex data distribution, Flow Matching achieves superior sample quality with straight, deterministic ODE trajectories.
DDPM models learn reverse Gaussian noise steps; Flow Matching learns optimal straight trajectories connecting noise to data.
From DDPM to Stochastic Differential Equations (SDEs)
In classical Denoising Diffusion Probabilistic Models (DDPM), a forward process gradually adds Gaussian noise to data x_0 over discrete timesteps t = 1..T. Song et al. (2020) generalized this to continuous time via Probability Flow SDEs:
dx = f(x, t) dt + g(t) dw (Forward SDE)
dx = [ f(x, t) - g(t)^2 * grad_x log p_t(x) ] dt + g(t) dw_bar (Reverse SDE)
The reverse process requires learning the Score Function grad_x log p_t(x) using a U-Net or Diffusion Transformer (DiT).
Rectified Flow Matching mathematics
While diffusion paths curve unpredictably through high-dimensional space, Flow Matching (Lipman et al., 2022) constructs a linear interpolation path between random Gaussian noise x_0 ~ N(0, I) and target image data x_1 ~ p_data:
x_t = (1 - t) * x_0 + t * x_1 for t in [0, 1]
The target velocity vector field u_t(x_t | x_0, x_1) that drives x_0 straight to x_1 is simply the constant derivative:
u_t = d(x_t) / dt = x_1 - x_0
Velocity field regression loss
The neural network v_theta(x_t, t) (typically a Diffusion Transformer) is trained to predict this velocity vector field using a simple Mean Squared Error loss:
L_FlowMatching = E_(t, x_0, x_1) [ || v_theta(x_t, t) - (x_1 - x_0) ||^2 ]
Because the conditional vector paths are straight lines, the learned Probability Flow Ordinary Differential Equation (ODE) can be integrated with far fewer function evaluations (e.g. 4 to 16 Euler steps instead of 50 to 1,000 DDPM steps)!
PyTorch implementation of Flow Matching vector field training
import torch
import torch.nn as nn
def train_flow_matching_step(
model: nn.Module,
x_1: torch.Tensor, # Real images: (batch_size, channels, height, width)
optimizer: torch.optim.Optimizer
) -> float:
"""
Executes a single Rectified Flow Matching training iteration.
"""
batch_size = x_1.shape[0]
# 1. Sample random Gaussian noise x_0
x_0 = torch.randn_like(x_1)
# 2. Sample uniform timesteps t in [0, 1]
t = torch.rand(batch_size, 1, 1, 1, device=x_1.device)
# 3. Compute linear interpolation x_t
x_t = (1 - t) * x_0 + t * x_1
# 4. Target velocity vector field
target_velocity = x_1 - x_0
# 5. Model predicts velocity field v_theta(x_t, t)
pred_velocity = model(x_t, t.squeeze())
# 6. Compute MSE Loss
loss = torch.mean((pred_velocity - target_velocity) ** 2)
optimizer.zero_grad()
loss.backward()
optimizer.step()
return loss.item()
Flow Matching vs DDPM architectural comparison
- Straight Trajectories: Flow Matching constructs straight probability trajectories, enabling fast Euler solver sampling in 10-20 steps.
- Diffusion Transformers (DiT): Replaces classical convolutional U-Nets with Patchified Transformer backbones for enhanced compute scaling.
- Simplified Loss Function: Direct velocity field regression
v_theta(x_t, t) = x_1 - x_0simplifies training stability compared to noise-predictioneps_thetaepsilon loss.