All articles Research

Titans Architecture and the MIRAS Framework

A deep dive into Google Research's Titans architecture and the MIRAS framework: how test-time memorization, surprise-driven updates, and attentional bias push models past the quadratic-attention wall.

Understanding how Titans and MIRAS redefine efficiency and memory in modern AI systems. Both come out of Google Research and share a single provocative idea: memory is not a fixed table you read from, it is a small network you keep training while the model is running.

Titans is the architecture that learns to memorize at test time; MIRAS is the framework that explains why that works and how far it generalizes.

Introduction: the long-context problem

Traditional Transformers face challenges when scaling to very long sequences due to quadratic attention costs. Every token attends to every other token, so doubling the context roughly quadruples the compute and blows up the key-value (KV) cache that must live in memory during generation. Titans architecture and the MIRAS framework were introduced to overcome these limitations by combining efficiency, scalability, and continuous learning.

The deeper problem is not just cost, it is expressivity. Softmax attention gives you a perfect, lossless lookup over the context window, but nothing beyond it. Recurrent models give you unbounded context for a fixed cost, but they cram everything into one fixed-size state and forget the details. Titans and MIRAS ask a different question: what if the model carried a small neural network as its memory, and kept training that network on the fly as new tokens arrive?

Background: why attention hits a wall

It helps to name the three families of sequence models before we place Titans among them.

  • Softmax Transformers. Exact pairwise attention. Quadratic time and memory in sequence length, but extremely accurate in-context recall.
  • Recurrent / linear models. RNNs, and modern state-space models like Mamba, plus linear-attention variants. They maintain a fixed-size state updated token by token, giving linear cost and constant memory, at the price of lossy compression of history.
  • Test-time memory models. The newer family Titans belongs to, where the recurrent "state" is itself a learnable module whose parameters are updated by an inner optimization loop during inference.

The key insight the 2024-2025 literature converged on is that many linear-attention and state-space updates can be read as online learning: the state update is a gradient step on some inner objective. Titans and MIRAS take that observation seriously and turn it into a design space.

flowchart TD A[Long input sequence] --> B{How is history stored?} B -->|All tokens, pairwise| C[Softmax Transformer] B -->|Fixed-size vector state| D[RNN / SSM / Linear attention] B -->|Small network trained at test time| E[Titans neural memory] C --> C1[Lossless recall / quadratic cost] D --> D1[Constant cost / lossy compression] E --> E1[Large capacity / linear cost / adaptive]

Titans Architecture

Titans is a new AI architecture designed to address the context length problem in Transformers. It merges the speed of recurrent neural networks with the accuracy of Transformers. It was introduced by Google Research in "Titans: Learning to Memorize at Test Time" (Ali Behrouz, Peilin Zhong, Vahab Mirrokni), and its central contribution is a deep neural long-term memory module.

  • Neural long-term memory module: Helps the model memorize historical context while focusing on current input. Instead of a lookup table, memory is a small multi-layer network whose weights are the stored knowledge.
  • Efficient scaling: Handles extremely long contexts such as full documents, genome sequences, or extended video streams, reportedly scaling to context windows beyond two million tokens.
  • Real-time adaptability: Updates memory dynamically during inference, enabling continuous adaptation.

Titans organizes memory into three cooperating components, which is easiest to see as a hierarchy:

flowchart TD IN[Incoming tokens] --> STM[Short-term memory: attention over a local window] IN --> LTM[Long-term memory: deep neural memory updated at test time] PM[Persistent memory: fixed, task-level learnable tokens] --> STM STM --> OUT[Output representation] LTM --> OUT LTM -->|surprise-driven writes| LTM
ComponentTime scaleWhat it holdsUpdated when?
Persistent memoryFixedTask-general knowledge as learnable tokensOnly at training time
Short-term memory (attention)ImmediateExact local context in the current windowEvery step (via attention)
Long-term neural memoryAcross the whole sequenceCompressed, salient historyEvery step, at test time

Learning to Memorize at Test Time

The interesting question is what the long-term memory chooses to store. Titans answers with a biologically flavored heuristic: store what is surprising. Formally, surprise is measured by the gradient of an associative-memory loss with respect to the memory's parameters. A token that the memory already predicts well produces a small gradient and is largely ignored; a token that violates expectations produces a large gradient and is written strongly.

Two refinements make this practical:

  • Momentary vs. momentum surprise. Using only the instantaneous gradient makes the model forget that it was surprised a moment ago and stop tracking an ongoing surprising event. Titans adds a momentum term so surprise decays smoothly rather than resetting each step, much like momentum in SGD.
  • Adaptive forgetting. Unbounded writing eventually saturates any finite memory. A data-dependent weight-decay term acts as a learnable forgetting gate, letting the module flush stale information when the context calls for it.

Conceptually the inner update looks like a gradient-descent-with-momentum step on the memory network, run once per incoming token, with the "loss" being how badly the memory reconstructs the current key-value association:

# Pseudocode for the inner (test-time) memory update
# M      : the neural memory network (its weights are the "state")
# k, v   : key/value projected from the current token
# theta  : learning rate (data-dependent gate)
# alpha  : forgetting / weight-decay gate (data-dependent)
# eta    : momentum coefficient

loss     = associative_loss(M(k), v)      # how surprised are we?
grad     = d loss / d M.weights           # surprise signal
S        = eta * S_prev - theta * grad    # momentum surprise
M.weights = (1 - alpha) * M.weights + S   # forget a bit, then write

Because the update is a well-defined recurrence, Titans can be reformulated to run in parallel over chunks of the sequence during training, so it keeps the hardware efficiency of modern linear-attention kernels rather than being stuck with a slow token-by-token loop.

flowchart TD T[New token] --> P[Project to key/value] P --> L[Compute memory prediction loss] L --> G[Gradient = surprise] G --> M{Momentum: combine with past surprise} M --> W[Decay old memory, then apply update] W --> R[Updated long-term memory] R --> T2[Next token]

Three Ways to Wire in Memory: MAC, MAG, MAL

Titans does not prescribe a single topology. The paper proposes three ways to combine the neural memory with attention, trading off recall precision against efficiency.

VariantFull nameHow memory and attention combineCharacter
MACMemory as ContextRetrieved long-term memory is concatenated with the current window and fed to attentionStrongest recall; attention sees history as extra context
MAGMemory as GateAttention branch and memory branch are combined through a learned gateBalanced; blends local precision with global memory
MALMemory as LayerMemory is stacked as a dedicated layer before/after attentionSimplest to implement; memory acts as a pre-processing state
flowchart TD subgraph MAC[Memory as Context] A1[Long-term memory read] --> A2[Concatenate with window] --> A3[Attention] end subgraph MAG[Memory as Gate] B1[Attention branch] --> B3[Learned gate] B2[Memory branch] --> B3 end subgraph MAL[Memory as Layer] C1[Memory layer] --> C2[Attention layer] end

The MIRAS Framework

"MIRAS provides the foundation for continuously learning AI models with functional long-term memory."

The MIRAS framework complements Titans by enabling continuous learning and efficient handling of massive contexts. In the author's reading it stands for a Memory-Informed Recurrent Attention System: a lens for treating attention and recurrence as two faces of the same associative-memory machinery. In the formal Google Research paper it is introduced under the title "It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization" (Behrouz et al., 2025), where the same acronym names a unifying design space rather than a single model.

  • Continuous learning: Models can keep learning during use, not just during training. The memory's parameters are still moving while the model serves requests.
  • Efficient memory updates: Allows AI to update its core memory dynamically while processing data streams, using cheap online-optimization steps rather than full retraining.
  • Balance between compression and detail: Retains fine-grained information across long sequences, unlike RNNs that compress context into a single hidden state.

The core claim of MIRAS is unifying: attention, linear attention, state-space models, and Titans-style memory are all instances of one template — an inner learner that minimizes an "attentional bias" objective under a retention constraint. Once you see the template, you can mix and match its pieces to invent new architectures on purpose instead of by accident.

The Four Design Axes of MIRAS

MIRAS decomposes any test-time memory model into four independent choices. Classical attention, Mamba-style SSMs, and Titans each correspond to particular settings of these knobs.

AxisQuestion it answersExample choices
Memory architectureWhat structure stores the state?Vector, matrix, or a deep neural network (as in Titans)
Attentional biasWhat objective does the inner learner minimize?L2 (dot-product) loss, robust p-norm losses, Huber-style losses
Retention gateHow is old information forgotten or preserved?Weight decay, elastic-net / regularized retention, data-dependent gates
Learning algorithmHow is the inner objective optimized per step?Gradient descent, momentum, adaptive optimizers

The important conceptual move is around attentional bias. Standard attention implicitly minimizes an L2 reconstruction objective, which is sensitive to outliers — a single noisy token can dominate the update. MIRAS argues that swapping in more robust objectives yields memory that is less distracted by noise and better at retaining structure over very long streams. Similarly, generalizing the forgetting term from plain weight decay to richer regularizers gives finer control over the compression/detail trade-off.

flowchart TD START[Sequence model as test-time learner] --> A[Axis 1: memory architecture] A --> B[Axis 2: attentional bias / objective] B --> C[Axis 3: retention gate] C --> D[Axis 4: optimization algorithm] D --> OUT[A concrete architecture instance] OUT -.example.-> T[Titans = deep memory + L2 + decay + momentum GD]

Moneta, Yaad, and Memora

To show the framework is generative and not just descriptive, the MIRAS paper instantiates several new models by turning the knobs above. They are worth naming because they make the abstract axes concrete:

  • Moneta — explores generalized p-norm attentional biases, decoupling the objective used to write memory from the standard dot-product form.
  • Yaad — uses a robust, Huber-like objective so that surprising-but-noisy tokens do not overwrite well-established memory.
  • Memora — pairs a bounded/normalized objective with a retention scheme aimed at stable long-horizon behavior.

Each is just a different coordinate in the four-axis space, which is exactly the point: Titans becomes one recipe among many that the framework can express.

How It Compares

Placing these ideas next to the architectures practitioners already know clarifies the trade-offs.

PropertySoftmax TransformerSSM / linear attention (e.g. Mamba)Titans + MIRAS
Cost vs. lengthQuadraticLinearLinear (with chunked parallel training)
Memory stateFull KV cacheFixed-size vector/matrixA small neural network
Recall fidelityExact, in-windowLossyAdaptive, surprise-weighted
Adapts at inference?NoState updates, weights frozenYes — memory weights keep learning
Practical context lengthTens to hundreds of KVery longReported past 2M tokens

A useful mental model: the Transformer is a perfect but expensive scratchpad, the SSM is a cheap but forgetful summary, and Titans is a student who keeps a compact notebook and rewrites it whenever something surprising happens.

Why They Matter

Traditional Transformers are powerful but limited by quadratic scaling with sequence length. Titans and MIRAS offer a path toward scalable, memory-efficient, and continuously learning AI. They land in a 2025-2026 moment where the field is actively hunting for post-Transformer or hybrid designs — state-space models, linear attention, and hybrid stacks that interleave a few full-attention layers with many efficient ones. Titans fits this trend but adds the distinctive twist of a memory that is itself trained during inference.

Applications include:

  • Full-document understanding — reasoning over entire books, codebases, or legal filings without chunking hacks.
  • Genomic analysis — sequences where meaningful dependencies span millions of base pairs.
  • Long video or multimodal sequence processing — hours of frames where salient events are sparse and must be remembered.
  • Streaming and agentic settings — where the model runs continuously and benefits from adapting its memory to the ongoing session.

Caveats and Open Questions

Honest engineering means naming the sharp edges too:

  • Test-time compute. Running an inner optimization step per token adds work at inference; the chunked-parallel formulation mitigates but does not eliminate this.
  • Stability of continuous learning. A memory that keeps updating can drift or be steered by adversarial inputs, so retention and forgetting gates are safety-relevant, not just accuracy knobs.
  • Evaluation. Long-context benchmarks are still maturing; strong needle-in-a-haystack scores do not always translate to genuine long-range reasoning.
  • Ecosystem maturity. Softmax attention has years of kernel optimization behind it; newer memory modules are still catching up on tooling and hardware support.

Conclusion

Titans is the architecture that implements efficient long-term memory, while MIRAS is the framework that formalizes how models can continuously update and use that memory in practice. Titans supplies a concrete, surprise-driven neural memory with three integration patterns (MAC, MAG, MAL); MIRAS steps back and shows that this memory is one point in a four-axis design space spanning architecture, attentional bias, retention, and optimization. Together, they represent a significant step forward in building AI systems that are scalable, adaptable, and capable of handling long-term dependencies.



As AI evolves, architectures like Titans and frameworks like MIRAS will shape the future of intelligent systems — not by making attention obsolete, but by giving models a memory that keeps learning long after training ends.

← Back to all articles