Understanding how Titans and MIRAS redefine efficiency and memory in modern AI systems. Both come out of Google Research and share a single provocative idea: memory is not a fixed table you read from, it is a small network you keep training while the model is running.
Titans is the architecture that learns to memorize at test time; MIRAS is the framework that explains why that works and how far it generalizes.
Introduction: the long-context problem
Traditional Transformers face challenges when scaling to very long sequences due to quadratic attention costs. Every token attends to every other token, so doubling the context roughly quadruples the compute and blows up the key-value (KV) cache that must live in memory during generation. Titans architecture and the MIRAS framework were introduced to overcome these limitations by combining efficiency, scalability, and continuous learning.
The deeper problem is not just cost, it is expressivity. Softmax attention gives you a perfect, lossless lookup over the context window, but nothing beyond it. Recurrent models give you unbounded context for a fixed cost, but they cram everything into one fixed-size state and forget the details. Titans and MIRAS ask a different question: what if the model carried a small neural network as its memory, and kept training that network on the fly as new tokens arrive?
Background: why attention hits a wall
It helps to name the three families of sequence models before we place Titans among them.
- Softmax Transformers. Exact pairwise attention. Quadratic time and memory in sequence length, but extremely accurate in-context recall.
- Recurrent / linear models. RNNs, and modern state-space models like Mamba, plus linear-attention variants. They maintain a fixed-size state updated token by token, giving linear cost and constant memory, at the price of lossy compression of history.
- Test-time memory models. The newer family Titans belongs to, where the recurrent "state" is itself a learnable module whose parameters are updated by an inner optimization loop during inference.
The key insight the 2024-2025 literature converged on is that many linear-attention and state-space updates can be read as online learning: the state update is a gradient step on some inner objective. Titans and MIRAS take that observation seriously and turn it into a design space.
Titans Architecture
Titans is a new AI architecture designed to address the context length problem in Transformers. It merges the speed of recurrent neural networks with the accuracy of Transformers. It was introduced by Google Research in "Titans: Learning to Memorize at Test Time" (Ali Behrouz, Peilin Zhong, Vahab Mirrokni), and its central contribution is a deep neural long-term memory module.
- Neural long-term memory module: Helps the model memorize historical context while focusing on current input. Instead of a lookup table, memory is a small multi-layer network whose weights are the stored knowledge.
- Efficient scaling: Handles extremely long contexts such as full documents, genome sequences, or extended video streams, reportedly scaling to context windows beyond two million tokens.
- Real-time adaptability: Updates memory dynamically during inference, enabling continuous adaptation.
Titans organizes memory into three cooperating components, which is easiest to see as a hierarchy:
| Component | Time scale | What it holds | Updated when? |
|---|---|---|---|
| Persistent memory | Fixed | Task-general knowledge as learnable tokens | Only at training time |
| Short-term memory (attention) | Immediate | Exact local context in the current window | Every step (via attention) |
| Long-term neural memory | Across the whole sequence | Compressed, salient history | Every step, at test time |
Learning to Memorize at Test Time
The interesting question is what the long-term memory chooses to store. Titans answers with a biologically flavored heuristic: store what is surprising. Formally, surprise is measured by the gradient of an associative-memory loss with respect to the memory's parameters. A token that the memory already predicts well produces a small gradient and is largely ignored; a token that violates expectations produces a large gradient and is written strongly.
Two refinements make this practical:
- Momentary vs. momentum surprise. Using only the instantaneous gradient makes the model forget that it was surprised a moment ago and stop tracking an ongoing surprising event. Titans adds a momentum term so surprise decays smoothly rather than resetting each step, much like momentum in SGD.
- Adaptive forgetting. Unbounded writing eventually saturates any finite memory. A data-dependent weight-decay term acts as a learnable forgetting gate, letting the module flush stale information when the context calls for it.
Conceptually the inner update looks like a gradient-descent-with-momentum step on the memory network, run once per incoming token, with the "loss" being how badly the memory reconstructs the current key-value association:
# Pseudocode for the inner (test-time) memory update
# M : the neural memory network (its weights are the "state")
# k, v : key/value projected from the current token
# theta : learning rate (data-dependent gate)
# alpha : forgetting / weight-decay gate (data-dependent)
# eta : momentum coefficient
loss = associative_loss(M(k), v) # how surprised are we?
grad = d loss / d M.weights # surprise signal
S = eta * S_prev - theta * grad # momentum surprise
M.weights = (1 - alpha) * M.weights + S # forget a bit, then write
Because the update is a well-defined recurrence, Titans can be reformulated to run in parallel over chunks of the sequence during training, so it keeps the hardware efficiency of modern linear-attention kernels rather than being stuck with a slow token-by-token loop.
Three Ways to Wire in Memory: MAC, MAG, MAL
Titans does not prescribe a single topology. The paper proposes three ways to combine the neural memory with attention, trading off recall precision against efficiency.
| Variant | Full name | How memory and attention combine | Character |
|---|---|---|---|
| MAC | Memory as Context | Retrieved long-term memory is concatenated with the current window and fed to attention | Strongest recall; attention sees history as extra context |
| MAG | Memory as Gate | Attention branch and memory branch are combined through a learned gate | Balanced; blends local precision with global memory |
| MAL | Memory as Layer | Memory is stacked as a dedicated layer before/after attention | Simplest to implement; memory acts as a pre-processing state |
The MIRAS Framework
"MIRAS provides the foundation for continuously learning AI models with functional long-term memory."
The MIRAS framework complements Titans by enabling continuous learning and efficient handling of massive contexts. In the author's reading it stands for a Memory-Informed Recurrent Attention System: a lens for treating attention and recurrence as two faces of the same associative-memory machinery. In the formal Google Research paper it is introduced under the title "It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization" (Behrouz et al., 2025), where the same acronym names a unifying design space rather than a single model.
- Continuous learning: Models can keep learning during use, not just during training. The memory's parameters are still moving while the model serves requests.
- Efficient memory updates: Allows AI to update its core memory dynamically while processing data streams, using cheap online-optimization steps rather than full retraining.
- Balance between compression and detail: Retains fine-grained information across long sequences, unlike RNNs that compress context into a single hidden state.
The core claim of MIRAS is unifying: attention, linear attention, state-space models, and Titans-style memory are all instances of one template — an inner learner that minimizes an "attentional bias" objective under a retention constraint. Once you see the template, you can mix and match its pieces to invent new architectures on purpose instead of by accident.
The Four Design Axes of MIRAS
MIRAS decomposes any test-time memory model into four independent choices. Classical attention, Mamba-style SSMs, and Titans each correspond to particular settings of these knobs.
| Axis | Question it answers | Example choices |
|---|---|---|
| Memory architecture | What structure stores the state? | Vector, matrix, or a deep neural network (as in Titans) |
| Attentional bias | What objective does the inner learner minimize? | L2 (dot-product) loss, robust p-norm losses, Huber-style losses |
| Retention gate | How is old information forgotten or preserved? | Weight decay, elastic-net / regularized retention, data-dependent gates |
| Learning algorithm | How is the inner objective optimized per step? | Gradient descent, momentum, adaptive optimizers |
The important conceptual move is around attentional bias. Standard attention implicitly minimizes an L2 reconstruction objective, which is sensitive to outliers — a single noisy token can dominate the update. MIRAS argues that swapping in more robust objectives yields memory that is less distracted by noise and better at retaining structure over very long streams. Similarly, generalizing the forgetting term from plain weight decay to richer regularizers gives finer control over the compression/detail trade-off.
Moneta, Yaad, and Memora
To show the framework is generative and not just descriptive, the MIRAS paper instantiates several new models by turning the knobs above. They are worth naming because they make the abstract axes concrete:
- Moneta — explores generalized p-norm attentional biases, decoupling the objective used to write memory from the standard dot-product form.
- Yaad — uses a robust, Huber-like objective so that surprising-but-noisy tokens do not overwrite well-established memory.
- Memora — pairs a bounded/normalized objective with a retention scheme aimed at stable long-horizon behavior.
Each is just a different coordinate in the four-axis space, which is exactly the point: Titans becomes one recipe among many that the framework can express.
How It Compares
Placing these ideas next to the architectures practitioners already know clarifies the trade-offs.
| Property | Softmax Transformer | SSM / linear attention (e.g. Mamba) | Titans + MIRAS |
|---|---|---|---|
| Cost vs. length | Quadratic | Linear | Linear (with chunked parallel training) |
| Memory state | Full KV cache | Fixed-size vector/matrix | A small neural network |
| Recall fidelity | Exact, in-window | Lossy | Adaptive, surprise-weighted |
| Adapts at inference? | No | State updates, weights frozen | Yes — memory weights keep learning |
| Practical context length | Tens to hundreds of K | Very long | Reported past 2M tokens |
A useful mental model: the Transformer is a perfect but expensive scratchpad, the SSM is a cheap but forgetful summary, and Titans is a student who keeps a compact notebook and rewrites it whenever something surprising happens.
Why They Matter
Traditional Transformers are powerful but limited by quadratic scaling with sequence length. Titans and MIRAS offer a path toward scalable, memory-efficient, and continuously learning AI. They land in a 2025-2026 moment where the field is actively hunting for post-Transformer or hybrid designs — state-space models, linear attention, and hybrid stacks that interleave a few full-attention layers with many efficient ones. Titans fits this trend but adds the distinctive twist of a memory that is itself trained during inference.
Applications include:
- Full-document understanding — reasoning over entire books, codebases, or legal filings without chunking hacks.
- Genomic analysis — sequences where meaningful dependencies span millions of base pairs.
- Long video or multimodal sequence processing — hours of frames where salient events are sparse and must be remembered.
- Streaming and agentic settings — where the model runs continuously and benefits from adapting its memory to the ongoing session.
Caveats and Open Questions
Honest engineering means naming the sharp edges too:
- Test-time compute. Running an inner optimization step per token adds work at inference; the chunked-parallel formulation mitigates but does not eliminate this.
- Stability of continuous learning. A memory that keeps updating can drift or be steered by adversarial inputs, so retention and forgetting gates are safety-relevant, not just accuracy knobs.
- Evaluation. Long-context benchmarks are still maturing; strong needle-in-a-haystack scores do not always translate to genuine long-range reasoning.
- Ecosystem maturity. Softmax attention has years of kernel optimization behind it; newer memory modules are still catching up on tooling and hardware support.
Conclusion
Titans is the architecture that implements efficient long-term memory, while MIRAS is the framework that formalizes how models can continuously update and use that memory in practice. Titans supplies a concrete, surprise-driven neural memory with three integration patterns (MAC, MAG, MAL); MIRAS steps back and shows that this memory is one point in a four-axis design space spanning architecture, attentional bias, retention, and optimization. Together, they represent a significant step forward in building AI systems that are scalable, adaptable, and capable of handling long-term dependencies.
As AI evolves, architectures like Titans and frameworks like MIRAS will shape the future of intelligent systems — not by making attention obsolete, but by giving models a memory that keeps learning long after training ends.