Fetching the paper…
Reading the bibliography…
The attention mechanism in Transformers is an important primitive for accurate and scalable sequence modeling.
Learning fast algorithms for linear transforms using butterfly factorizations, 2020
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 1903
Earlier work this paper cites.
Adaptive switching circuits
Bernard Widrow, Marcian E Hoff, et al · 1960
Earlier work this paper cites.
The WY representation for products of householder matrices
Christian H. Bischof and Charles Van Loan · 1985
Earlier work this paper cites.
Prefix sums and their applications
Guy E Blelloch · 1990
Earlier work this paper cites.
A fast on-line adaptive code
B. Ya. Ryabko · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
A new data structure for cumulative frequency tables
Peter M. Fenwick · 1994
Earlier work this paper cites.
Hierarchical matrices based on a weak admissibility criterion
Wolfgang Hackbusch, Boris N Khoromskij, and Ronald Kriemann · 2004
Earlier work this paper cites.
Accumulating householder transformations, revisited
Thierry Joffrain, Tze Meng Low, Enrique S. Quintana-Ortí, Robert A. van de Geijn, and Field G. Van Zee · 2006
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu, et al · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Low-rank updates and a divide-and-conquer method for linear matrix equations
Daniel Kressner, Stefano Massei, and Leonardo Robol · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, Hsiang-Tsung Kung, and David Cox · 2019
Earlier work this paper cites.
Bp-transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang · 2019
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
hm-toolbox: Matlab software for hodlr and hss matrices
Stefano Massei, Leonardo Robol, and Daniel Kressner · 2020
Earlier work this paper cites.
Finetuning pretrained transformers into rnns
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, and Noah A Smith · 2021
Earlier work this paper cites.
Fmmformer: Efficient and flexible transformer via decomposed near-field and far-field attention
Tan Nguyen, Vai Suliafu, Stanley Osher, Long Chen, and Bao Wang · 2021
Earlier work this paper cites.
In Proceedings of ICLR , 2021
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong · 2021
Earlier work this paper cites.
Linear Transformers Are Secretly Fast Weight Programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Earlier work this paper cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Cited alongside, same era.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2021
Cited alongside, same era.
H-transformer-1d: Fast one-dimensional hierarchical attention for sequences, 2021
Zhenhai Zhu and Radu Soricut · 2021
Cited alongside, same era.
Monarch: Expressive structured matrices for efficient and accurate training
Tri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
xLSTM: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael K Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Later among the works it cites.
Reparameterized multi-resolution convolutions for long sequence modelling
Harry Jake Cunningham, Giorgio Giannone, Mingtian Zhang, and Marc Peter Deisenroth · 2024
Later among the works it cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Later among the works it cites.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Tri Dao and Albert Gu · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Later among the works it cites.
Ruler: What’s the real context size of your long-context language models?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le · 2022
Cited alongside, same era.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Cited alongside, same era.
Multi resolution analysis (mra) for approximate self-attention, 2022
Zhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn M Fung, and Vikas Singh · 2022
Cited alongside, same era.
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re · 2023
Cited alongside, same era.
Polysketchformer: Fast transformers via sketching polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Cited alongside, same era.
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg · 2024
Later among the works it cites.
Fast multipole attention: A divide-and-conquer attention mechanism for long sequences, 2024
Yanming Kang, Giang Tran, and Hans De Sterck · 2024
Later among the works it cites.
Ring attention with blockwise transformers for near-infinite context
Hao Liu, Matei Zaharia, and Pieter Abbeel · 2024
Later among the works it cites.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Later among the works it cites.
Leave no context behind: Efficient infinite context transformers with infini-attention, 2024
Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal · 2024
Later among the works it cites.
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou, Przemysław Kazienko, et al · 2024
Later among the works it cites.
Compute better spent: Replacing dense layers with structured matrices
Shikai Qiu, Andres Potapczynski, Marc Finzi, Micah Goldblum, and Andrew Gordon Wilson · 2024
Later among the works it cites.
FlashAttention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Later among the works it cites.
Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024
Songlin Yang and Yu Zhang · 2024
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff, 2025
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré · 2025
Closest in time.
Unlocking state-tracking in linear RNNs through negative eigenvalues
Riccardo Grazzi, Julien Siems, Jörg K.H. Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil · 2025
Closest in time.
Forgetting transformer: Softmax attention with a forget gate
Zhixuan Lin, Evgenii Nikishin, Xu He, and Aaron Courville · 2025
Closest in time.
Flash inference: Near linear time inference for long convolution sequence models and beyond
Costin-Andrei Oncescu, Sanket Purandare, Stratos Idreos, and Sham M. Kakade · 2025
Closest in time.
Rwkv-7" goose" with expressive dynamic state evolution
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, et al · 2025
Closest in time.
Deltaproduct: Improving state-tracking in linear rnns via householder products
Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi · 2025
Closest in time.
Mesanet: Sequence modeling by locally optimal test-time training
Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, et al · 2025
Closest in time.
Sequential-parallel duality in prefix scannable models
Morris Yau, Sharut Gupta, Valerie Engelmayer, Kazuki Irie, Stefanie Jegelka, and Jacob Andreas · 2025
Closest in time.
Flame: Flash language modeling made easy, January 2025
Yu Zhang and Songlin Yang · 2025
Closest in time.