All articles Generative Modeling

Diffusion Models & Flow Matching: Continuous-Time Generative Architectures

The mathematical formulation of Score-based Generative Models, SDE/ODE formulations, Rectified Flow Matching, and Classifier-Free Guidance (CFG).

Generative modeling has undergone a paradigm shift from Autoregressive models and Generative Adversarial Networks (GANs) toward Continuous-Time Diffusion Models and Rectified Flow Matching (powering architectures like Flux.1 and SD3). By formulating image synthesis as learning a vector field that transports a simple Gaussian noise distribution to a complex data distribution, Flow Matching achieves superior sample quality with straight, deterministic ODE trajectories.

DDPM models learn reverse Gaussian noise steps; Flow Matching learns optimal straight trajectories connecting noise to data.

From DDPM to Stochastic Differential Equations (SDEs)

In classical Denoising Diffusion Probabilistic Models (DDPM), a forward process gradually adds Gaussian noise to data x_0 over discrete timesteps t = 1..T. Song et al. (2020) generalized this to continuous time via Probability Flow SDEs:

dx = f(x, t) dt + g(t) dw  (Forward SDE)
dx = [ f(x, t) - g(t)^2 * grad_x log p_t(x) ] dt + g(t) dw_bar  (Reverse SDE)

The reverse process requires learning the Score Function grad_x log p_t(x) using a U-Net or Diffusion Transformer (DiT).

Rectified Flow Matching mathematics

While diffusion paths curve unpredictably through high-dimensional space, Flow Matching (Lipman et al., 2022) constructs a linear interpolation path between random Gaussian noise x_0 ~ N(0, I) and target image data x_1 ~ p_data:

x_t = (1 - t) * x_0 + t * x_1   for t in [0, 1]

The target velocity vector field u_t(x_t | x_0, x_1) that drives x_0 straight to x_1 is simply the constant derivative:

u_t = d(x_t) / dt = x_1 - x_0

Velocity field regression loss

The neural network v_theta(x_t, t) (typically a Diffusion Transformer) is trained to predict this velocity vector field using a simple Mean Squared Error loss:

L_FlowMatching = E_(t, x_0, x_1) [ || v_theta(x_t, t) - (x_1 - x_0) ||^2 ]

Because the conditional vector paths are straight lines, the learned Probability Flow Ordinary Differential Equation (ODE) can be integrated with far fewer function evaluations (e.g. 4 to 16 Euler steps instead of 50 to 1,000 DDPM steps)!

PyTorch implementation of Flow Matching vector field training

import torch
import torch.nn as nn

def train_flow_matching_step(
    model: nn.Module,
    x_1: torch.Tensor, # Real images: (batch_size, channels, height, width)
    optimizer: torch.optim.Optimizer
) -> float:
    """
    Executes a single Rectified Flow Matching training iteration.
    """
    batch_size = x_1.shape[0]
    
    # 1. Sample random Gaussian noise x_0
    x_0 = torch.randn_like(x_1)
    
    # 2. Sample uniform timesteps t in [0, 1]
    t = torch.rand(batch_size, 1, 1, 1, device=x_1.device)
    
    # 3. Compute linear interpolation x_t
    x_t = (1 - t) * x_0 + t * x_1
    
    # 4. Target velocity vector field
    target_velocity = x_1 - x_0
    
    # 5. Model predicts velocity field v_theta(x_t, t)
    pred_velocity = model(x_t, t.squeeze())
    
    # 6. Compute MSE Loss
    loss = torch.mean((pred_velocity - target_velocity) ** 2)
    
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    
    return loss.item()

Flow Matching vs DDPM architectural comparison

  1. Straight Trajectories: Flow Matching constructs straight probability trajectories, enabling fast Euler solver sampling in 10-20 steps.
  2. Diffusion Transformers (DiT): Replaces classical convolutional U-Nets with Patchified Transformer backbones for enhanced compute scaling.
  3. Simplified Loss Function: Direct velocity field regression v_theta(x_t, t) = x_1 - x_0 simplifies training stability compared to noise-prediction eps_theta epsilon loss.
← Back to all articles