All articles Deep Learning

Transformers: The Architecture That Changed AI Forever

A deep, modern dive into Transformer architecture — attention math, encoder/decoder design, the 2025-2026 stack (RoPE, RMSNorm, GQA, FlashAttention, MoE, long context), efficiency techniques, evaluation, and where the field is heading.

Transformers revolutionized AI by replacing recurrence with attention, enabling scalability, parallelism, and long-term dependency handling. What began as a machine-translation architecture in 2017 is now the substrate for essentially every frontier model in language, vision, audio, and code. This post walks from the original design through the components that define the modern 2025-2026 Transformer stack.

Attention turned sequence modeling from a step-by-step recurrence into a fully parallel, differentiable lookup — and that single shift unlocked the scaling era.

Introduction

Introduced in 2017's "Attention is All You Need" by Vaswani et al. at Google, the Transformer replaced the recurrent (RNN, LSTM, GRU) and convolutional backbones that dominated sequence modeling with a design built almost entirely around self-attention. The payoff was profound: because attention processes every position in parallel rather than one timestep at a time, Transformers train efficiently on modern accelerators, capture long-range dependencies without gradient decay, and — crucially — keep getting better as you add parameters, data, and compute.

That last property, later formalized as scaling laws (Kaplan et al., 2020; the Chinchilla compute-optimal analysis, Hoffmann et al., 2022), is why the architecture became the foundation of large language models. The same core block underlies BERT, the GPT family, T5, Llama, Mistral, Gemini, Claude, Qwen, and DeepSeek. Understanding the Transformer is therefore the single highest-leverage thing you can learn to reason about modern AI systems.

Why Attention Replaced Recurrence

To appreciate why the architecture mattered, it helps to compare it directly with what came before.

PropertyRNN / LSTMCNN (seq)Transformer
Parallelism across sequenceNo (sequential in time)PartialFull
Path length between two tokensO(n)O(log n) with dilationO(1)
Long-range dependency handlingWeak (vanishing gradients)Limited by receptive fieldStrong (direct attention)
Compute per layerO(n · d²)O(k · n · d²)O(n² · d)
Hardware fit (GPU/TPU)PoorGoodExcellent

The key insight: the Transformer trades a favorable-looking linear compute cost for a quadratic one, but buys back constant path length and full parallelism. On accelerators that reward large dense matrix multiplications, that trade is overwhelmingly worth it — and most of the last several years of systems research has been about taming the quadratic term without giving up the parallelism.

Core Components

  • Embeddings & Positional Encoding: Convert discrete tokens into dense vectors and inject information about sequence order (attention itself is permutation-invariant, so position must be added explicitly).
  • Self-Attention: Computes a relevance-weighted mixture of all tokens using queries, keys, and values.
  • Multi-Head Attention: Runs several attention computations in parallel subspaces so the model can attend to different relationships (syntax, coreference, position) simultaneously.
  • Position-wise Feedforward Networks (FFN/MLP): Two linear layers with a non-linearity applied independently at each position; this is where most parameters and much of the stored "knowledge" live.
  • Residual Connections & Normalization: Skip connections plus layer normalization stabilize training and keep gradients flowing through deep stacks.

The Attention Mechanism, Step by Step

Scaled dot-product attention is the heart of the model. Given an input matrix of token embeddings, three learned projections produce queries (Q), keys (K), and values (V):

Q = X W_Q      # what each token is "looking for"
K = X W_K      # what each token "offers" as a match key
V = X W_V      # the content each token contributes

scores  = (Q @ K.T) / sqrt(d_k)   # similarity, scaled to control variance
weights = softmax(scores)         # row-wise probabilities over positions
output  = weights @ V             # relevance-weighted mixture of values

The division by sqrt(d_k) keeps the dot products from growing with dimension, which would otherwise push softmax into saturated, low-gradient regions. In multi-head attention, this is done h times in parallel with smaller per-head dimensions, then concatenated and projected:

head_i = Attention(X W_Q^i, X W_K^i, X W_V^i)
MHA(X) = Concat(head_1, ..., head_h) W_O

Three attention flavors appear in a full Transformer, distinguished only by where Q, K, and V come from and by masking:

TypeQ fromK, V fromMaskingUsed in
Encoder self-attentionInputInputNone (bidirectional)BERT, encoders
Masked (causal) self-attentionOutput so farOutput so farFuture positions hiddenGPT, decoders
Cross-attentionDecoder stateEncoder outputNoneTranslation, T5

Encoder-Decoder Structure

The encoder maps the input sequence into a stack of contextual representations using bidirectional self-attention — every token can see every other token. The decoder generates output autoregressively: masked self-attention prevents a position from peeking at future tokens, while cross-attention lets the decoder query the encoder's representations. This split is ideal for sequence-transduction tasks such as machine translation and summarization, where a full source is available but the target is produced left to right.

Modern generative LLMs mostly drop the encoder and use a decoder-only stack, folding "understanding" and "generation" into a single causal model. Encoder-only and encoder-decoder designs remain strong for embeddings, retrieval, and constrained transduction.

Transformer Architecture Diagram

To better understand how Transformers process input and generate output, here is a simplified flow through the original encoder-decoder design:

flowchart TD A[Input Text] --> B[Tokenization] B --> C[Embeddings + Positional Encoding] C --> D[Encoder Stack] D --> E[Contextual Representations] E --> F[Decoder Stack] F --> G[Linear Projection to Vocabulary] G --> H[Softmax] H --> I[Predicted Tokens] subgraph Encoder D1[Self-Attention] --> D2[Feedforward Network] D2 --> D3[Residual + LayerNorm] end subgraph Decoder F1[Masked Self-Attention] --> F2[Cross-Attention with Encoder Outputs] F2 --> F3[Feedforward Network] F3 --> F4[Residual + LayerNorm] end

A single decoder-only block — the workhorse of contemporary LLMs — looks like this internally, with pre-normalization and residual streams:

flowchart TD X[Hidden state x] --> N1[RMSNorm] N1 --> ATT[Causal Self-Attention with RoPE] ATT --> R1[Add residual] X --> R1 R1 --> N2[RMSNorm] N2 --> FFN[Gated MLP / MoE experts] FFN --> R2[Add residual] R1 --> R2 R2 --> Y[Output to next layer]

The Modern Transformer Stack (2025-2026)

The 2017 block still runs, but the components that ship in current frontier models differ in several well-established ways. These changes improve stability, throughput, and context length without altering the core attention idea.

ComponentOriginal (2017)Modern defaultWhy it changed
Normalization placementPost-LayerNormPre-normalizationStabler gradients in very deep stacks
Norm typeLayerNormRMSNormCheaper, no mean subtraction, comparable quality
Positional infoAbsolute sinusoidalRotary embeddings (RoPE), ALiBiRelative positions, better length extrapolation
Activation / FFNReLUSwiGLU / GeGLU gated MLPHigher quality per parameter
Attention KV sharingMulti-head (MHA)Grouped-query (GQA) / multi-query (MQA)Smaller KV cache, faster decoding
FFN capacityDenseMixture-of-Experts (MoE)More parameters at fixed compute per token
Attention kernelNaive softmaxFlashAttention (IO-aware)Memory-linear, much faster on GPU

Rotary Position Embeddings (RoPE)

Instead of adding a position vector to the input, RoPE rotates the query and key vectors by an angle proportional to their position. Because the attention score then depends on the difference of positions, the model naturally encodes relative distance and extrapolates better to sequences longer than those seen in training — especially when combined with interpolation or NTK-aware scaling for context extension.

Grouped-Query Attention (GQA)

During generation, the dominant memory cost is the KV cache that stores keys and values for every past token. MQA shares a single K/V head across all query heads; GQA is the middle ground, sharing K/V across small groups. This shrinks the cache and memory bandwidth dramatically with negligible quality loss, and is now standard in most open-weight models.

Mixture-of-Experts (MoE)

MoE replaces the dense FFN with many expert MLPs and a router that activates only a few per token. This decouples total parameter count from per-token compute, letting a model hold far more knowledge while keeping inference affordable. Sparse MoE routing (with load-balancing losses and, in some designs, shared always-on experts) is a defining feature of several 2024-2025 frontier and open-weight systems.

flowchart TD T[Token hidden state] --> R{Router / gating} R -->|top-k| E1[Expert 1] R -->|top-k| E3[Expert 3] R -.->|skipped| E2[Expert 2] R -.->|skipped| E4[Expert 4] E1 --> S[Weighted sum] E3 --> S S --> O[Output]

Variants of Transformers

  • BERT: Encoder-only, bidirectional, trained with masked language modeling. Strong for classification, retrieval embeddings, and extractive QA.
  • GPT family: Decoder-only, autoregressive; the template for modern generative LLMs.
  • T5 / FLAN-T5: Encoder-decoder that casts every task as text-to-text.
  • Transformer-XL: Adds segment-level recurrence for longer effective context.
  • Longformer / BigBird / Reformer / Performer: Sparse or linearized attention for very long sequences.
  • Vision Transformer (ViT): Treats image patches as tokens, bringing the architecture to computer vision.
  • Modern decoder-only families: Llama, Mistral/Mixtral, Qwen, Gemma, and DeepSeek combine the modern-stack components above at scale.

Training Objectives and the Post-Training Stack

Pretraining objectives still define the variant:

  • Masked Language Modeling (MLM): Predict masked tokens from bidirectional context (BERT).
  • Autoregressive / Causal LM (ALM): Predict the next token (GPT and virtually all generative LLMs).
  • Span corruption / seq2seq: Reconstruct corrupted spans or translate/summarize (T5).

What changed most in recent years is the post-training pipeline that turns a raw pretrained model into a useful assistant:

  1. Supervised fine-tuning (SFT): Instruction and conversation data teach the model to follow requests.
  2. Preference optimization: RLHF (reward model + PPO) or the simpler, reward-model-free DPO align outputs with human preferences.
  3. Reasoning-focused RL: Reinforcement learning with verifiable rewards trains long chain-of-thought "reasoning" models that spend more inference-time compute on hard problems — a defining theme of 2024-2025 releases.
flowchart TD A[Web-scale pretraining corpus] --> B[Base model: next-token prediction] B --> C[Supervised fine-tuning on instructions] C --> D[Preference optimization: RLHF / DPO] D --> E[Reasoning RL with verifiable rewards] E --> F[Aligned assistant model]

Inference, KV Cache, and Decoding

Generation is autoregressive: each new token is produced from all previous ones. Two practical concerns dominate serving.

The KV cache

Recomputing attention over the whole prefix at every step would be wasteful, so keys and values are cached. The cache grows linearly with context length and is the main driver of memory use during long-context inference — which is exactly why GQA/MQA and paged attention (efficient cache memory management) matter so much.

Decoding strategies

  • Greedy / beam search: Deterministic; good for translation, weak for open-ended text.
  • Sampling with temperature, top-k, top-p (nucleus): Controls diversity vs. coherence.
  • Speculative decoding: A small draft model proposes several tokens that the large model verifies in one pass, cutting latency without changing the output distribution.

Challenges

  • Computational cost: Self-attention scales quadratically with sequence length in compute (and the KV cache linearly in memory).
  • Bias: Models inherit and can amplify biases present in training data.
  • Hallucination: Fluent output is not grounded output; models can state falsehoods confidently.
  • Interpretability: Attributing a prediction to specific internal mechanisms is hard, though mechanistic interpretability is advancing.
  • Data provenance and privacy: Web-scale corpora raise copyright, consent, and leakage concerns.

Efficiency Techniques

Reducing cost and improving scalability spans training, serving, and architecture:

  • Pruning: Remove redundant weights, heads, or layers.
  • Quantization: Store and compute in lower precision (INT8, INT4, FP8); post-training and quantization-aware variants preserve most quality.
  • Knowledge distillation: Train a compact student to mimic a larger teacher.
  • Parameter-efficient fine-tuning: LoRA and QLoRA update small adapter matrices instead of all weights.
  • FlashAttention: An IO-aware, fused attention kernel that avoids materializing the full score matrix, giving memory-linear, hardware-friendly attention.
  • Sparse / linear attention: Reformer, Longformer, Performer, and related methods reduce the quadratic term.

Scaling Context and Beyond Attention

Two parallel research threads aim at the quadratic bottleneck.

Longer context

Context windows have grown from a few thousand tokens to hundreds of thousands and, in some systems, into the million-token range, enabled by RoPE-based length extension, FlashAttention, and better memory management. The open problem is effective use of context — models can accept long inputs but may still under-utilize the middle of them.

State-space models and hybrids

Structured state-space models — most prominently the Mamba family — process sequences in near-linear time with a recurrent-style scan, offering strong long-sequence throughput. In practice, hybrid architectures that interleave a few full-attention layers with many SSM or linear-attention layers have emerged as a pragmatic sweet spot, keeping attention's in-context precision while reducing cost. This is an active, well-established 2024-2025 direction rather than a wholesale replacement of Transformers.

ApproachTime complexityStrengthTrade-off
Full attentionO(n²)Exact in-context recallExpensive at long n
Sparse / linear attention~O(n)Cheaper long sequencesApproximate, may lose detail
State-space (Mamba)~O(n)Fast long-range throughputWeaker exact recall
Hybrid (attention + SSM)MixedBalance of bothMore design complexity

Evaluation Metrics

No single number captures a modern model; evaluation is layered:

  • Perplexity: Intrinsic language-modeling quality.
  • Accuracy / F1: Classification and extractive tasks.
  • BLEU / ROUGE / chrF: Translation and summarization overlap metrics.
  • Task benchmarks: MMLU (knowledge), GSM8K and MATH (math reasoning), HumanEval and MBPP (code), and agentic/tool-use suites.
  • Human and LLM-as-judge preference: Coherence, helpfulness, factuality; arena-style pairwise comparisons.

A recurring caution: public benchmarks leak into training data over time, so benchmark scores must be read alongside contamination checks and held-out evaluations.

Applications

  • Machine translation and multilingual understanding.
  • Conversational assistants and text generation (ChatGPT, Claude, Gemini, Copilot).
  • Code generation, review, and agentic software workflows.
  • Summarization of news, legal, and scientific documents.
  • Retrieval and question answering, typically via embeddings plus RAG.
  • Multimodal systems spanning text, image, audio, and video (CLIP-style encoders and native multimodal LLMs).

Future Directions

  • Efficient long context: Million-token windows that are also reliably used end to end.
  • Memory-augmented and retrieval-native models: External and learned memory (including RAG and test-time memory frameworks) to extend effective context and freshness.
  • Reasoning and agents: Inference-time compute scaling, tool use, and multi-step planning.
  • Architectural diversity: MoE at larger scale and Transformer/SSM hybrids as defaults rather than experiments.
  • Alignment and governance: Bias mitigation, interpretability, provenance, and fairness auditing built into the pipeline.

Conclusion

Transformers are not just a model architecture — they are the foundation of modern AI. By replacing recurrence with attention, they unlocked parallel training, long-range reasoning, and predictable scaling across language, vision, and multimodal data. The core block from 2017 endures, but the working architecture of 2025-2026 is meaningfully different: pre-norm RMSNorm, rotary positions, gated MLPs, grouped-query attention, sparse MoE capacity, FlashAttention kernels, and an elaborate post-training stack. As long-context methods, state-space hybrids, and reasoning-focused training continue to mature, attention-based models will keep shaping the future of intelligent systems.



Transformers revolutionized AI by replacing recurrence with attention, enabling scalability, parallelism, and long-term dependency handling.

← Back to all articles