Anamaria-Roberta Hartl, Levente Zólyomi, David Stap, Pieter-Jan Hoedt, Niklas Schmidinger, Lukas Hauzenberger, Sebastian Böck, Günter Klambauer, Sepp Hochreiter
5 min
Abstract
Transformers dominate modern sequence modeling, but their quadratic attention incurs substantial computational cost. Subquadratic architectures offer a scalable alternative. However, it remains unclear which designs yield the most effective sequence models. We compare three leading approaches: xLSTM, Mamba-2, and Gated DeltaNet. We evaluate these models on tasks with complex dependencies: (1) code-model pre-training, (2) distillation of code models from large language models, and (3) pre-training of time-series foundation models. Across these settings, xLSTM delivers the strongest overall performance. To explain xLSTM's advantage, we present a unified formulation and analyze the underlying architectural mechanisms, focusing on state tracking and memory dynamics. Our results show that xLSTM enables more flexible and stable memory correction via its gating scheme. We corroborate these findings on controlled synthetic length-generalization tasks. Overall, our findings indicate that xLSTM's gains on complex tasks stem from robust state tracking and accumulation.
Sam: That is what the evidence suggests. When the researchers tested these models on synthetic tasks—think of them as carefully designed logic puzzles meant to isolate specific abilities—xLSTM was the only architecture that could reliably count and track states across very long sequences. The others would eventually lose the thread.
Alex: And how did they move from those controlled puzzles to something more real-world?
Sam: They used a process called "linearization," which is a form of knowledge distillation. The idea is this: you start with a large, powerful model that already knows a lot—the "teacher." Then you train a much smaller, more efficient model—the "student"—to reproduce the teacher's outputs. By swapping the standard processing blocks in the student model for these new, faster subquadratic operators, they found that xLSTM-based students consistently performed better at code generation than the alternatives.
Alex: So the architectural advantage isn't just theoretical. It shows up in practice when you're actually trying to build something useful.
Sam: It does. Though it's worth being clear about the limits. Most of this research was done on models with around 400 million parameters. The models powering the most advanced coding assistants today can be hundreds of times larger than that. We have a clear signal at smaller scales, but whether these advantages hold at the very top end of the scale is still an open question.
Alex: So the "inductive bias"—that structural advantage—is a promising starting point, but it isn't a guaranteed win at every level of complexity.
Sam: Exactly right. And there's another nuance worth mentioning. In practice, researchers often combine these new efficient operators with a smaller number of standard attention layers. So it isn't a complete replacement of the old approach—it's more of a partnership, where the efficient parts handle the heavy lifting and the standard layers provide a safety net for the most complex reasoning.
Alex: What does the near-term path look like for this, then?
Sam: One direction is toward more capable devices that don't rely on massive servers. Because these models use a fixed amount of memory regardless of how long the sequence gets, they can handle very long contexts without running out of resources—which is a real constraint for standard models. The other direction is time-series forecasting. The same principles of precise memory management that help with code also help these models detect patterns in data that unfolds over time, like energy grid signals, where older architectures might miss subtle long-range structure.
Alex: So the broader shift here is away from just throwing more computing power at the problem, and toward building architectures that are actually suited to the logic of the task.
Sam: That is a good way to put it. It's a meaningful step toward understanding what structural properties actually allow a model to reason carefully—rather than simply relying on scale. And that kind of understanding is likely to matter as these systems are asked to do more precise, consequential work.
Alex: Thanks for listening to ResearchPod.