Modern Vision-Language Models (VLMs) — such as LLaVA, GPT-4o, and Gemini 1.5 Pro — treat visual inputs not as raw pixel arrays, but as discrete visual tokens projected into the LLM's text embedding space. This architectural breakdown explores the mechanics of patch tokenization, contrastive pre-training with SigLIP (Sigmoid Loss for Language-Image Pre-training), and multi-modal projection layers.
Pixels are to Vision-Language Models what subword tokens are to Text LLMs.
Vision Transformers & patch tokenization
To process an image of dimension (H x W x C) through a Transformer, the image is divided into non-overlapping spatial patches of size (P x P) (typically 14x14 pixels). The total number of visual tokens N is:
N = (H / P) * (W / P)
For a 336x336 image with 14x14 patches: N = (336/14) * (336/14) = 24 * 24 = 576 visual tokens
Each 14x14 pixel patch (14 x 14 x 3 = 588 values) is flattened into a vector and passed through a trainable Linear Projection matrix to map it into embedding dimension D_vision.
SigLIP vs CLIP InfoNCE contrastive loss
Standard CLIP uses the InfoNCE loss, which requires softmax normalization over all pair combinations across GPUs in a batch:
L_InfoNCE = - log ( exp(sim(i, t_pos) / tau) / sum_j exp(sim(i, t_j) / tau) )
SigLIP (Zhai et al., 2023) replaces global softmax normalization with a pairwise sigmoid loss:
L_SigLIP = - sum_ij log sigmoid( z_ij * (sim(img_i, txt_j) * t + b) )
where z_ij = +1 if i == j (positive pair) and -1 if i != j (negative pair)
By operating independently on individual image-text pairs without global softmax normalization, SigLIP allows batch sizes to scale past 1,000,000 without inter-GPU memory synchronization bottlenecks.
The LLaVA vision-language connector
Once the Vision Transformer (e.g., SigLIP-So400M) extracts 576 patch embeddings of shape (576, D_vision), a multimodal projection layer maps these vectors directly into the LLM's hidden dimension D_text:
+-----------------------+
| Input Image (336x336)|
+-----------+-----------+
| (14x14 Patching)
+-----------v-----------+
| SigLIP ViT Encoder | -> Output Shape: (576, 1152)
+-----------+-----------+
| (MLP Projection Connector: 1152 -> 4096)
+-----------v-----------+
| LLM Token Stream | -> [ <image_patch_0>, ..., <image_patch_575>, "Describe this image..." ]
+-----------------------+
PyTorch visual patch projection implementation
import torch
import torch.nn as nn
class VisionLanguageConnector(nn.Module):
"""
LLaVA-style 2-layer MLP projection connector mapping
vision transformer embeddings to LLM text hidden dimension.
"""
def __init__(self, vision_dim: int = 1152, llm_dim: int = 4096):
super().__init__()
self.proj = nn.Sequential(
nn.Linear(vision_dim, llm_dim),
nn.GELU(),
nn.Linear(llm_dim, llm_dim)
)
def forward(self, image_features: torch.Tensor) -> torch.Tensor:
# image_features: (batch_size, num_patches, vision_dim)
return self.proj(image_features)
# Example execution
connector = VisionLanguageConnector(vision_dim=1152, llm_dim=4096)
raw_vit_features = torch.randn(2, 576, 1152) # 2 images, 576 patches each
projected_tokens = connector(raw_vit_features)
print("Projected Visual Token Shape:", projected_tokens.shape)
# Output: torch.Size([2, 576, 4096]) - Ready to concatenate with LLM text tokens!
Key multimodal design patterns
- SigLIP over CLIP: Use SigLIP pre-trained encoders for superior zero-shot classification and higher batch training efficiency.
- 2-Layer MLP Connectors: Replace linear projection layers with 2-layer MLP GELU connectors for non-linear feature mapping.
- Dynamic High-Resolution Grid: Crop high-res images into multi-tile sub-grids (e.g. 2x2 grids) to preserve small text and fine UI details.