Research ยท August 14, 2026
What if attention isn't all you need?
I still remember reading "Attention Is All You Need." At the time, I saw it as the paper that changed sequence modeling. I didn't realize how long one question from it would stay with me: what if attention isn't all you need?
That question feels even more interesting today. A lot of recent open models are experimenting with hybrid architectures, using Mamba, state-space models, recurrent mechanisms, or other forms of persistent state alongside attention.
Nemotron 3 Nano 30B-A3B has 52 layers: 23 Mamba-2 layers, 6 attention layers, and 23 MoE layers. Bamba-9B uses 29 Mamba-2 layers and only 3 full-attention layers across 32 layers. Falcon-H1 takes another route, combining Mamba-2 and attention within the same hybrid blocks.
The interesting shift isn't "attention is dead." It's that we're asking a more precise question: how much of sequence modeling actually needs attention?
For open models, we can inspect the architecture. For the strongest closed models, we simply don't know how much of the same experimentation is happening internally. That uncertainty is part of what pulled me in.
So I started experimenting with HGDM, Hierarchical Gated Delta Memory. It is a byte-level, attention-free architecture built around hierarchical recurrent memory, content-aware temporal decimation, cross-scale state highways, and a custom Triton fused scan.
I also pushed a 1.006B parameter version to 1.46B training tokens. And yes, I know the problem. Under the classic Chinchilla compute-optimal rule of roughly 20 tokens per parameter, a 1B model would need around 20B tokens, not 1.46B.
So this is not a claim that a 1B HGDM is fully trained or compute-optimal. It's an early experiment. A way to see what the architecture learns before spending orders of magnitude more compute.
Still small. Still unfinished. Still building. HGDM on GitHub.
Written by Thanniru Sai Teja. More writing.