✳flâneur — a map of the web's best reading
Tensor-Transformer Variants are Surprisingly Performant — LessWrong
lesswrong.com · 4,577 words · saved by 3 readers
I've been researching tensor networks as a more interpretable architecture, but whenever I tell people this, they always ask "But is it any good?" …
x Tensor-Transformer Variants are Surprisingly Performant — LessWrong AI Frontpage 87 Tensor-Transformer Variants are Surprisingly Performant by Logan Riggs 12th Jan 2026 5 min read 16 87 I've been researching tensor networks as a more interpretable architecture, but whenever I tell people this, they always ask "But is it any good?" So I trained multiple 500M parameter LLMs on fineweb, showing the tensor variant needed ~4% more batches of data to match CE-loss. There's a few caveats, so my personal estimate is around 15% worst to 10% better. Details below. The Architecture Replacing MLP w/ a B
Explore this link on the map →saved by
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- The Annotated Transformernlp.seas.harvard.edu
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- 1706.03762arxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Annotated Transformernlp.seas.harvard.edu
- [2512.19428] Attention Is Not What You Needarxiv.org
- Softmax Linear Unitstransformer-circuits.pub
- The Transformer Family | Hacker Newsnews.ycombinator.com