Fetching the paper…
Reading the bibliography…
Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale.
“Adjustment of an inverse matrix corresponding to a change in one element of a given matrix”
Jack Sherman and Winifred Morrison · 1950
Earlier work this paper cites.
“Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition”
Thomas. Cover · 1965
Earlier work this paper cites.
“Non-holographic associative memory”
David Willshaw, O Buneman and Hugh Longuet-Higgins · 1969
Earlier work this paper cites.
“Neural networks and physical systems with emergent collective computational abilities.”
John Hopfield · 1982
Earlier work this paper cites.
“On the capabilities of multilayer perceptrons”
Eric Baum · 1988
Earlier work this paper cites.
“Adaptive switching circuits”
Bernard Widrow and Marcian Hoff · 1988
Earlier work this paper cites.
“Neural network capacity using delta rule”
DL Prados and SC Kak · 1989
Earlier work this paper cites.
“Learning to control fast-weight memories: An alternative to recurrent nets. Accepted for publication in”
JH Schmidhuber · 1992
Earlier work this paper cites.
“Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets”
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
“Long Short-term Memory”
Jürgen Schmidhuber and Sepp Hochreiter · 1997
Earlier work this paper cites.
“Learning capability and storage capacity of two-hidden-layer feedforward networks”
Guang-Bin Huang · 2003
Earlier work this paper cites.
“The organization of behavior: A neuropsychological theory”
Donald Hebb · 2005
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
“On the number of linear regions of deep neural networks”
Guido Montufar, Razvan Pascanu, Kyunghyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
Razvan Pascanu, Guido Montufar and Yoshua Bengio · 2014
Earlier work this paper cites.
“Gaussian error linear units (gelus)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“Dense associative memory for pattern recognition”
Dmitry Krotov and John Hopfield · 2016
Earlier work this paper cites.
“The LAMBADA dataset: Word prediction requiring a broad discourse context”
Denis Paperno, German Kruszewski, Angeliki Lazaridou, Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernandez · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension”
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi and Hannaneh Hajishirzi · 2017
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2017
Earlier work this paper cites.
“Neural semantic encoders”
Tsendsuren Munkhdalai and Hong Yu · 2017
Earlier work this paper cites.
“Learning and memory: Basic principles, processes, and procedures”
W Terry · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Think you have solved question answering? try arc, the ai2 reasoning challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“Local polynomial modelling and its applications: monographs on statistics and applied probability 66”
Jianqing Fan · 2018
Earlier work this paper cites.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins and Kristina Toutanova · 2019
Earlier work this paper cites.
“DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs”
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh and Matt Gardner · 2019
Earlier work this paper cites.
“Natural questions: a benchmark for question answering research”
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin and Kenton Lee · 2019
Earlier work this paper cites.
“Openceres: When open information extraction meets the semi-structured web”
Colin Lockard, Prashant Shiralkar and Xin Dong · 2019
Earlier work this paper cites.
“Metalearned neural memory”
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang and Adam Trischler · 2019
Earlier work this paper cites.
“Social IQa: Commonsense Reasoning about Social Interactions”
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le and Yejin Choi · 2019
Cited alongside, same era.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Cited alongside, same era.
“Low-rank bottleneck in multi-head attention models”
Srinadh Bhojanapalli, Chulhee Yun, Ankit Rawat, Sashank Reddi and Sanjiv Kumar · 2020
Cited alongside, same era.
“Piqa: Reasoning about physical commonsense in natural language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Cited alongside, same era.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Cited alongside, same era.
“Going beyond linear transformers with recurrent fast weight programmers”
“Muon: An optimizer for hidden layers in neural networks”, 2024
Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse and Jeremy Bernstein · 2024
Later among the works it cites.
“PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels”
Praneeth Kacham, Vahab Mirrokni and Peilin Zhong · 2024
Later among the works it cites.
“PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels”
Praneeth Kacham, Vahab Mirrokni and Peilin Zhong · 2024
Later among the works it cites.
“BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack”
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin and Mikhail Burtsev · 2024
Later among the works it cites.
“A survey on long video generation: Challenges, methods, and prospects”
Chengxuan Li, Di Huang, Zeyu Lu, Yang Xiao, Qingqi Pei and Lei Bai · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kazuki Irie, Imanol Schlag, Robert Csordas and Jurgen Schmidhuber · 2021
Cited alongside, same era.
“Finetuning Pretrained Transformers into RNNs”
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen and Noah. Smith · 2021
Cited alongside, same era.
“Hierarchical associative memory”
Dmitry Krotov · 2021
Cited alongside, same era.
“Hopfield Networks is All You Need”
Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Lukas Gruber, Markus Holzleitner, Thomas Adler, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2021
Cited alongside, same era.
“Winogrande: An adversarial winograd schema challenge at scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2021
Cited alongside, same era.
“The dynamics of gradient descent for overparametrized neural networks”
Siddhartha Satpathi and Rayadurgam Srikant · 2021
Cited alongside, same era.
“Linear transformers are secretly fast weight programmers”
Imanol Schlag, Kazuki Irie and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Xiaoyu Li, Yuanpeng Li, Yingyu Liang, Zhenmei Shi and Zhao Song · 2024
Later among the works it cites.
“Parallelizing non-linear sequential models over the sequence length”
Yi Lim, Qi Zhu, Joshua Selfridge and Muhammad Kasim · 2024
Later among the works it cites.
“Longhorn: State space models are amortized online learners”
Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone and Qiang Liu · 2024
Later among the works it cites.
“Lost in the middle: How language models use long contexts”
Nelson Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang · 2024
Later among the works it cites.
“Exponential capacity of dense associative memories”
Carlo Lucibello and Marc Mézard · 2024
Later among the works it cites.
“The Illusion of State in State-Space Models”
William Merrill, Jackson Petty and Ashish Sabharwal · 2024
Later among the works it cites.
“The fineweb datasets: Decanting the web for the finest text data at scale”
Guilherme Penedo, Hynek Kydlicek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von and Thomas Wolf · 2024
Later among the works it cites.
“Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence”
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Xingjian Du, Teddy Ferdinan and Haowen Hou · 2024
Later among the works it cites.
“Mechanistic design and scaling of hybrid architectures”
Michael Poli, Armin Thomas, Eric Nguyen, Pragaash Ponnusamy, Bjorn Deiseroth, Kristian Kersting, Taiji Suzuki, Brian Hie, Stefano Ermon and Christopher Re · 2024
Later among the works it cites.
“Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling”
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang and Weizhu Chen · 2024
Later among the works it cites.
“Learning to (learn at test time): Rnns with expressive hidden states”
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang and Sanmi Koyejo · 2024
Later among the works it cites.
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Tommaso Guidi, Marco Gori and Stefano Melacci · 2024
Later among the works it cites.
“Rnns are not transformers (yet): The key bottleneck on in-context retrieval”
Kaiyue Wen, Xingyu Dang and Kaifeng Lyu · 2024
Later among the works it cites.
“Gated Delta Networks: Improving Mamba2 with Delta Rule”
Songlin Yang, Jan Kautz and Ali Hatamizadeh · 2024
Later among the works it cites.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Later among the works it cites.
“Parallelizing linear transformers with the delta rule over sequence length”
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen and Yoon Kim · 2024
Later among the works it cites.
“Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers” https://ssrn.com/abstract=5240330
Zeyuan Allen-Zhu · 2025
Closest in time.
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong and Vahab Mirrokni · 2025
Closest in time.
“One-Minute Video Generation with Test-Time Training”
Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Cheung, Jan Kautz and Carlos Guestrin · 2025
Closest in time.
Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé and Morgane Rivière · 2025
Closest in time.
“Lattice: Learning to Efficiently Compress the Memory”
M. Karami and V. Mirrokni · 2025
Closest in time.
“Rwkv-7" goose" with expressive dynamic state evolution”
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Haowen Hou, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan and Saiteja Utpala · 2025
Closest in time.
“Implicit Language Models are RNNs: Balancing Parallelization and Expressivity”
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani and Jannes Gladrow · 2025
Closest in time.
“DeltaProduct: Increasing the Expressivity of DeltaNet Through Products of Householders”
Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil and Riccardo Grazzi · 2025
Closest in time.
“Test-time regression: a unifying framework for designing sequence models with associative memory”
Ke Wang, Jiaxin Shi and Emily Fox · 2025
Closest in time.