Fetching the paper…
Reading the bibliography…
Sequence models lie at the heart of modern deep learning.
Augmenting Self-attention with Persistent Memory, July 2019
S. Sukhbaatar, E. Grave, G. Lample, H. Jegou, and A. Joulin · 1907
Earlier work this paper cites.
Fast Transformer Decoding: One Write-Head is All You Need, Nov. 2019
N. Shazeer · 1911
Earlier work this paper cites.
Angenaherte auflosung von systemen linearer glei-chungen
S. Kaczmarz · 1937
Earlier work this paper cites.
Non-Holographic Associative Memory
D. J. Willshaw, O. P. Buneman, and H. C. Longuet-Higgins · 1969
Earlier work this paper cites.
Projection method for solving a singular system of linear equations and its applications
K. Tanabe · 1971
Earlier work this paper cites.
Correlation Matrix Memories
T. Kohonen · 1972
Earlier work this paper cites.
On optimal nonlinear associative recall
T. Poggio · 1975
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
J. J. Hopfield · 1982
Earlier work this paper cites.
Exponential convergence of recursive least squares with exponential forgetting factor
R. M. Johnstone, C. R. Johnson, R. R. Bitmead, and B. D. O. Anderson · 1982
Earlier work this paper cites.
Adaptive switching circuits
B. Widrow and M. E. Hoff · 1988
Earlier work this paper cites.
Parallel Models of Associative Memory
G. E. Hinton and J. A. Anderson · 1989
Earlier work this paper cites.
Self-Organization and Associative Memory , volume 8 of
T. Kohonen · 1989
Earlier work this paper cites.
Holography, Associative Memory, and Inductive Generalization
D. Willshaw · 1989
Earlier work this paper cites.
Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks
J. Schmidhuber · 1992
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Scaling Laws for Neural Language Models, Jan. 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
B. Schölkopf and A. J. Smola · 2001
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
N. Cesa-Bianchi, A. Conconi, and C. Gentile · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners, July 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
All of Nonparametric Statistics
L. Wasserman · 2006
Earlier work this paper cites.
Large Associative Memory Problem in Neurobiology and Machine Learning, Apr. 2021
D. Krotov and J. Hopfield · 2008
Earlier work this paper cites.
Minimum-disturbance description for the development of adaptation algorithms and a new leakage least squares algorithm
F. T. Castoldi and M. L. R. de Campos · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Adaptive Filter Theory
S. S. Haykin · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Predicting effects of noncoding variants with deep learning–based sequence model
J. Zhou and O. G. Troyanskaya · 2015
Earlier work this paper cites.
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber · 2016
Earlier work this paper cites.
Dense Associative Memory for Pattern Recognition
D. Krotov and J. J. Hopfield · 2016
Earlier work this paper cites.
Quasi-recurrent neural networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2017
Earlier work this paper cites.
Convolutional Sequence to Sequence Learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Earlier work this paper cites.
Attention is All you Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Local Polynomial Modelling and Its Applications: Monographs on Statistics and Applied Probability 66
J. Fan · 2018
Earlier work this paper cites.
The unreasonable effectiveness of the forget gate, Sept. 2018
J. van der Westhuizen and J. Lasenby · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
DeepAR: Probabilistic forecasting with autoregressive recurrent networks
D. Salinas, V. Flunkert, J. Gasthaus, and T. Januschowski · 2019
Cited alongside, same era.
The Bitter Lesson, 2019
R. Sutton · 2019
Cited alongside, same era.
Exact Gaussian Processes on a Million Data Points
K. A. Wang, G. Pleiss, J. Gardner, S. Tyree, K. Q. Weinberger, and A. G. Wilson · 2019
Cited alongside, same era.
Rethinking Attention with Performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2020
Cited alongside, same era.
HiPPO: Recurrent Memory with Optimal Polynomial Projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Cited alongside, same era.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Transformers Learn In-Context by Gradient Descent
J. V. Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Later among the works it cites.
Hyena hierarchy: towards larger convolutional language models
M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. Ré · 2023
Later among the works it cites.
Toeplitz neural network for sequence modeling
Z. Qin, X. Han, W. Sun, B. He, D. Li, D. Li, Y. Dai, L. Kong, and Y. Zhong · 2023
Later among the works it cites.
Sparse modular activation for efficient sequence modeling
L. Ren, Y. Liu, S. Wang, Y. Xu, C. Zhu, and C. Zhai · 2023
Later among the works it cites.
Sequence modeling with multiresolution convolutional memory
J. Shi, K. A. Wang, and E. B. Fox · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. Linderman · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Random Feature Attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong · 2020
Cited alongside, same era.
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
Y. Sun, X. Wang, Z. Liu, J. Miller, A. Efros, and M. Hardt · 2020
Cited alongside, same era.
Is Space-Time Attention All You Need for Video Understanding?
G. Bertasius, H. Wang, and L. Torresani · 2021
Cited alongside, same era.
Skyformer: remodel self-attention with Gaussian kernel and nyström method
Y. Chen, Q. Zeng, H. Ji, and Y. Yang · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
A. Gu, K. Goel, and C. Re · 2021
Cited alongside, same era.
Later among the works it cites.
Retentive Network: A Successor to Transformer for Large Language Models, Aug. 2023
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in Transformers
J. von Oswald, E. Niklasson, M. Schlegel, S. Kobayashi, N. Zucchet, N. Scherrer, N. Miller, M. Sandler, B. A. y Arcas, M. Vladymyrov, R. Pascanu, and J. Sacramento · 2023
Later among the works it cites.
Small-scale proxies for large-scale Transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. E. Everett, A. A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, J. Pennington, J. Sohl-Dickstein, K. Xu, J. Lee, J. Gilmer, and S. Kornblith · 2023
Later among the works it cites.
Linear Transformers with Learnable Kernel Functions are Better In-Context Models, Feb. 2024
Y. Aksenov, N. Balagansky, S. M. L. C. Vaina, B. Shaposhnikov, A. Gorbatovski, and D. Gavrilov · 2024
Later among the works it cites.
The Surprising Effectiveness of Test-Time Training for Abstract Reasoning, Nov. 2024
E. Akyürek, M. Damani, L. Qiu, H. Guo, Y. Kim, and J. Andreas · 2024
Later among the works it cites.
Chronos: Learning the language of time series
A. F. Ansari, L. Stella, A. C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. S. Rangapuram, S. P. Arango, S. Kapoor, J. Zschiegner, D. C. Maddix, H. Wang, M. W. Mahoney, K. Torkkola, A. G. Wilson, M. Bohlke-Schneider, and B. Wang · 2024
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff
S. Arora, S. Eyuboglu, M. Zhang, A. Timalsina, S. Alberti, J. Zou, A. Rudra, and C. Re · 2024
Later among the works it cites.
xLSTM: Extended Long Short-Term Memory, May 2024
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2024
Later among the works it cites.
Titans: Learning to Memorize at Test Time, Dec. 2024
A. Behrouz, P. Zhong, and V. Mirrokni · 2024
Later among the works it cites.
DiJiang: Efficient Large Language Models through Compact Kernelization
H. Chen, Liuzhicheng, X. Wang, Y. Tian, and Y. Wang · 2024
Later among the works it cites.
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
T. Dao and A. Gu · 2024
Later among the works it cites.
S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, G. Desjardins, A. Doucet, D. Budden, Y. W. Teh, R. Pascanu, N. D. Freitas, and C. Gulcehre · 2024
Later among the works it cites.
A Survey on In-context Learning, Oct. 2024
Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, T. Liu, B. Chang, X. Sun, L. Li, and Z. Sui · 2024
Later among the works it cites.
Can a transformer represent a Kalman filter?
G. Goel and P. Bartlett · 2024
Later among the works it cites.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
A. Gu and T. Dao · 2024
Later among the works it cites.
Demystify Mamba in Vision: A Linear Attention Perspective, May 2024
D. Han, Z. Wang, Z. Xia, Y. Han, Y. Pu, C. Ge, J. Song, S. Song, B. Zheng, and G. Huang · 2024
Later among the works it cites.
GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling, Jan. 2024
T. Katsch · 2024
Later among the works it cites.
Longhorn: State Space Models are Amortized Online Learners, Oct. 2024
B. Liu, R. Wang, L. Wu, Y. Feng, P. Stone, and Q. Liu · 2024
Later among the works it cites.
Sequence modeling and design from molecular to genome scale with Evo
E. Nguyen, M. Poli, M. G. Durrant, B. Kang, D. Katrekar, D. B. Li, L. J. Bartie, A. W. Thomas, S. H. King, G. Brixi, J. Sullivan, M. Y. Ng, A. Lewis, A. Lou, S. Ermon, S. A. Baccus, T. Hernandez-Boussard, C. Ré, P. D. Hsu, and B. L. Hie · 2024
Later among the works it cites.
Understanding Factual Recall in Transformers via Associative Memories, Dec. 2024
E. Nichani, J. D. Lee, and A. Bietti · 2024
Later among the works it cites.
OpenAI o1 System Card, 2024
OpenAI · 2024
Later among the works it cites.
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence, Sept. 2024
B. Peng, D. Goldstein, Q. Anthony, A. Albalak, E. Alcaide, S. Biderman, E. Cheah, X. Du, T. Ferdinan, H. Hou, P. Kazienko, K. K. GV, J. Kocoń, B. Koptyra, S. Krishna, R. M. Jr, J. Lin, N. Muennighoff, F. Obeid, A. Saito, G. Song, H. Tu, C. Wirawan, S. Woźniak, R. Zhang, B. Zhao, Q. Zhao, P. Zhou, J. Zhu, and R.-J. Zhu · 2024
Later among the works it cites.
HGRN2: Gated Linear RNNs with State Expansion
Z. Qin, S. Yang, W. Sun, X. Shen, D. Li, W. Sun, and Y. Zhong · 2024
Later among the works it cites.
Can Mamba Always Enjoy the "Free Lunch"?, Oct. 2024
R. Ren, Z. Li, and Y. Liu · 2024
Later among the works it cites.
There are like 4 more linear RNN papers out today, Apr. 2024
S. Rush · 2024
Later among the works it cites.
Learning to (Learn at Test Time): RNNs with Expressive Hidden States, Aug. 2024
Y. Sun, X. Li, K. Dalal, J. Xu, A. Vikram, G. Zhang, Y. Dubois, X. Chen, X. Wang, S. Koyejo, T. Hashimoto, and C. Guestrin · 2024
Later among the works it cites.
Mimetic Initialization Helps State Space Models Learn to Recall, Oct. 2024
A. Trockman, H. Harutyunyan, J. Z. Kolter, S. Kumar, and S. Bhojanapalli · 2024
Later among the works it cites.
KV Shifting Attention Enhances Language Modeling, Dec. 2024
M. Xu, W. Cheng, B. Wang, and W. Chen · 2024
Later among the works it cites.
An Analysis of Attention via the Lens of Exchangeability and Latent Variable Models, Apr. 2024
Y. Zhang, B. Liu, Q. Cai, L. Wang, and Z. Wang · 2024
Later among the works it cites.
DeltaProduct: Increasing the Expressivity of DeltaNet Through Products of Householders, Feb. 2025
J. Siems, T. Carstensen, A. Zela, F. Hutter, M. Pontil, and R. Grazzi · 2025
Closest in time.
Memory mosaics
J. Zhang, N. Nolte, R. Sadhukhan, B. Chen, and L. Bottou · 2025
Closest in time.