Fetching the paper…
Reading the bibliography…
Transformers with linear attention (i.e., linear transformers) and state-space models have recently been suggested as a viable linear-time alternative to transformers with softmax attention.
Adaptive switching circuits
B. Widrow, M. E. Hoff, et al · 1960
Earlier work this paper cites.
The WY representation for products of householder matrices
C. H. Bischof and C. V. Loan · 1985
Earlier work this paper cites.
The capacity of the hopfield associative memory
R. J. McEliece, E. C. Posner, E. R. Rodemich, and S. S. Venkatesh · 1987
Earlier work this paper cites.
The space of interactions in neural network models
E. Gardner · 1988
Earlier work this paper cites.
Neural network capacity using delta rule
D. Prados and S. Kak · 1989
Earlier work this paper cites.
Prefix sums and their applications
G. E. Blelloch · 1990
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Adaptive filter theory
S. Rogers · 1996
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2002
Earlier work this paper cites.
Accumulating householder transformations, revisited
T. Joffrain, T. M. Low, E. S. Quintana-Ortí, R. A. van de Geijn, and F. G. V. Zee · 2006
Earlier work this paper cites.
What if neural networks had svds?
A. Mathiasen, F. Hvilshoj, J. R. Jørgensen, A. Nasery, and D. Mottin · 2009
Earlier work this paper cites.
Delta learning rule for the active sites model
K. C. Lingashetty · 2010
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D. Clevert, T. Unterthiner, and S. Hochreiter · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Improving Variational Auto-Encoders using Householder Flow, 2016
J. M. Tomczak and M. Welling · 2016
Earlier work this paper cites.
On a model of associative memory with huge storage capacity
M. Demircigil, J. Heusel, M. Löwe, S. Upgang, and F. Vermet · 2017
Earlier work this paper cites.
S. Elfwing, E. Uchibe, and K. Doya · 2017
Earlier work this paper cites.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Z. Mhammedi, A. D. Hellicar, A. Rahman, and J. Bailey · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
On orthogonality and learning recurrent networks with long term dependencies
E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Fast blocking of householder reflectors on graphics processors
A. E. T. Dominguez and E. S. Q. Orti · 2018
Earlier work this paper cites.
Orthogonal recurrent neural networks with scaled cayley transform
K. Helfrich, D. Willmott, and Q. Ye · 2018
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
E. Martin and C. Cundy · 2018
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
E. Martin and C. Cundy · 2018
Earlier work this paper cites.
Rational recurrences
H. Peng, R. Schwartz, S. Thomson, and N. A. Smith · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Sylvester normalizing flows for variational inference
R. van den Berg, L. Hasenclever, J. M. Tomczak, and M. Welling · 2018
Earlier work this paper cites.
Stabilizing gradients for deep neural networks via efficient SVD parameterization
J. Zhang, Q. Lei, and I. S. Dhillon · 2018
Earlier work this paper cites.
Gated Orthogonal Recurrent Units: On Learning to Forget
L. Jing, C. Gulcehre, J. Peurifoy, Y. Shen, M. Tegmark, M. Soljacic, and Y. Bengio · 2019
Earlier work this paper cites.
OpenCeres: When open information extraction meets the semi-structured web
C. Lockard, P. Shiralkar, and X. L. Dong · 2019
Earlier work this paper cites.
Metalearned neural memory
T. Munkhdalai, A. Sordoni, T. Wang, and A. Trischler · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
P. Tillet, H. Kung, and D. D. Cox · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. LeBras, J. Gao, and Y. Choi · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
P. S. H. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontañón, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation, 2021
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2021
Cited alongside, same era.
Going beyond linear transformers with recurrent fast weight programmers
K. Irie, I. Schlag, R. Csordás, and J. Schmidhuber · 2021
Cited alongside, same era.
Going beyond linear transformers with recurrent fast weight programmers
K. Irie, I. Schlag, R. Csordás, and J. Schmidhuber · 2021
Cited alongside, same era.
Finetuning pretrained transformers into RNNs
Simplified state space layers for sequence modeling
J. T. H. Smith, A. Warrington, and S. W. Linderman · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. H. Smith, A. Warrington, and S. W. Linderman · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers
J. von Oswald, E. Niklasson, M. Schlegel, S. Kobayashi, N. Zucchet, N. Scherrer, N. Miller, M. Sandler, B. A. y Arcas, M. Vladymyrov, R. Pascanu, and J. Sacramento · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kasai, H. Peng, Y. Zhang, D. Yogatama, G. Ilharco, N. Pappas, Y. Mao, W. Chen, and N. A. Smith · 2021
Cited alongside, same era.
Large associative memory problem in neurobiology and machine learning
D. Krotov and J. J. Hopfield · 2021
Cited alongside, same era.
When attention meets fast recurrence: Training language models with reduced compute
T. Lei · 2021
Cited alongside, same era.
Fmmformer: Efficient and flexible transformer via decomposed near-field and far-field attention
T. M. Nguyen, V. Suliafu, S. J. Osher, L. Chen, and B. Wang · 2021
Cited alongside, same era.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. A. Smith, and L. Kong · 2021
Cited alongside, same era.
Hopfield networks is all you need
H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, L. Gruber, M. Holzleitner, T. Adler, D. P. Kreil, M. K. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2021
Cited alongside, same era.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2021
Cited alongside, same era.
Later among the works it cites.
Pretraining without attention
J. Wang, J. N. Yan, A. Gu, and A. Rush · 2023
Later among the works it cites.
Diffusion models without attention
J. N. Yan, J. Gu, and A. M. Rush · 2023
Later among the works it cites.
FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism, 2024
S. Yang and Y. Zhang · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
S. Yang, B. Wang, Y. Shen, R. Panda, and Y. Kim · 2023
Later among the works it cites.
Efficient long-range transformers: You need to attend more, but not necessarily at every layer
Q. Zhang, D. Ram, C. Hawkins, S. Zha, and T. Zhao · 2023
Later among the works it cites.
Linear Transformers with Learnable Kernel Functions are Better In-Context Models, 2024
Y. Aksenov, N. Balagansky, S. M. L. C. Vaina, B. Shaposhnikov, A. Gorbatovski, and D. Gavrilov · 2024
Closest in time.
In-context language learning: Arhitectures and algorithms
E. Akyürek, B. Wang, Y. Kim, and J. Andreas · 2024
Closest in time.
The hidden attention of mamba models, 2024
A. Ali, I. Zimerman, and L. Wolf · 2024
Closest in time.
xlstm: Extended long short-term memory
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2024
Closest in time.
Titans: Learning to memorize at test time, 2024
A. Behrouz, P. Zhong, and V. Mirrokni · 2024
Closest in time.
MetaLA: Unified optimal linear approximation to softmax attention map
Y. Chou, M. Yao, K. Wang, Y. Pan, R.-J. Zhu, J. Wu, Y. Zhong, Y. Qiao, B. XU, and G. Li · 2024
Closest in time.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
T. Dao and A. Gu · 2024
Closest in time.
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models, 2024
S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, G. Desjardins, A. Doucet, D. Budden, Y. W. Teh, R. Pascanu, N. De Freitas, and C. Gulcehre · 2024
Closest in time.
Recurrentgemma: Moving past transformers for efficient open language models
R. Griffin and G. Teams · 2024
Closest in time.
Vig: Linear-complexity visual sequence learning with gated linear attention
B. Liao, X. Wang, L. Zhu, Q. Zhang, and C. Huang · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
O. Lieber, B. Lenz, H. Bata, G. Cohen, J. Osin, I. Dalmedigos, E. Safahi, S. Meirom, Y. Belinkov, S. Shalev-Shwartz, et al · 2024
Closest in time.
Longhorn: State space models are amortized online learners
B. Liu, R. Wang, L. Wu, Y. Feng, P. Stone, and Q. Liu · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length
X. Ma, X. Yang, W. Xiong, B. Chen, L. Yu, H. Zhang, J. May, L. Zettlemoyer, O. Levy, and C. Zhou · 2024
Closest in time.
Linearizing large language models
J. Mercat, I. Vasiljevic, S. Keh, K. Arora, A. Dave, A. Gaidon, and T. Kollar · 2024
Closest in time.
The Illusion of State in State-Space Models, 2024
W. Merrill, J. Petty, and A. Sabharwal · 2024
Closest in time.
Linear Attention as Iterated Hopfield Networks
B. Millidge · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention
T. Munkhdalai, M. Faruqui, and S. Gopal · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks
J. Park, J. Park, Z. Xiong, N. Lee, J. Cho, S. Oymak, K. Lee, and D. Papailiopoulos · 2024
Closest in time.
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, 2024
B. Peng, D. Goldstein, Q. Anthony, A. Albalak, E. Alcaide, S. Biderman, E. Cheah, X. Du, T. Ferdinan, H. Hou, P. Kazienko, K. K. GV, J. Kocoń, B. Koptyra, S. Krishna, R. McClelland Jr., N. Muennighoff, F. Obeid, A. Saito, G. Song, H. Tu, S. Woźniak, R. Zhang, B. Zhao, Q. Zhao, P. Zhou, J. Zhu, and R.-J. Zhu · 2024
Closest in time.
Mechanistic Design and Scaling of Hybrid Architectures, 2024
M. Poli, A. W. Thomas, E. Nguyen, P. Ponnusamy, B. Deiseroth, K. Kersting, T. Suzuki, B. Hie, S. Ermon, C. Ré, C. Zhang, and S. Massaroli · 2024
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling
L. Ren, Y. Liu, Y. Lu, Y. Shen, C. Liang, and W. Chen · 2024
Closest in time.
Associative recurrent memory transformer
I. Rodkin, Y. Kuratov, A. Bulatov, and M. Burtsev · 2024
Closest in time.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Y. Schiff, C.-H. Kao, A. Gokaslan, T. Dao, A. Gu, and V. Kuleshov · 2024
Closest in time.
Power scheduler: A batch size and token number agnostic learning rate scheduler
Y. Shen, M. Stallone, M. Mishra, G. Zhang, S. Tan, A. Prasad, A. M. Soria, D. D. Cox, and R. Panda · 2024
Closest in time.
L. Team · 2024
Closest in time.
Graph-mamba: Towards long-range graph sequence modeling with selective state spaces
C. X. Wang, O. Tsepa, J. Ma, and B. Wang · 2024
Closest in time.
Gated delta networks: Improving mamba2 with delta rule, 2024
S. Yang, J. Kautz, and A. Hatamizadeh · 2024
Closest in time.
Gated slot attention for efficient linear-time sequence modeling
Y. Zhang, S. Yang, R. Zhu, Y. Zhang, L. Cui, Y. Wang, B. Wang, F. Shi, B. Wang, W. Bi, P. Zhou, and G. Fu · 2024
Closest in time.
A unified implicit attention formulation for gated-linear recurrent sequence models
I. Zimerman, A. Ali, and L. Wolf · 2024
Closest in time.