Fetching the paper…
Reading the bibliography…
Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention.
“Neural networks and physical systems with emergent collective computational abilities.”
John Hopfield · 1982
Earlier work this paper cites.
“Adaptive switching circuits”
Bernard Widrow and Marcian Hoff · 1988
Earlier work this paper cites.
“Multilayer feedforward networks are universal approximators”
Kurt Hornik, Maxwell Stinchcombe and Halbert White · 1989
Earlier work this paper cites.
“Neural network capacity using delta rule”
DL Prados and SC Kak · 1989
Earlier work this paper cites.
“A learning algorithm for continually running fully recurrent neural networks”
Ronald Williams and David Zipser · 1989
Earlier work this paper cites.
“Local learning algorithms”
Léon Bottou and Vladimir Vapnik · 1992
Earlier work this paper cites.
“Learning to control fast-weight memories: An alternative to recurrent nets. Accepted for publication in”
JH Schmidhuber · 1992
Earlier work this paper cites.
“Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets”
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
“Long Short-term Memory”
Jürgen Schmidhuber and Sepp Hochreiter · 1997
Earlier work this paper cites.
“Systems of memory in the human brain”
Daniel Willingham · 1997
Earlier work this paper cites.
“Learning to forget: Continual prediction with LSTM”
Felix Gers, Jürgen Schmidhuber and Fred Cummins · 2000
Earlier work this paper cites.
“Learning and memory”
Hideyuki Okano, Tomoo Hirano and Evan Balaban · 2000
Earlier work this paper cites.
“The organization of behavior: A neuropsychological theory”
Donald Hebb · 2005
Earlier work this paper cites.
“SVM-KNN: Discriminative nearest neighbor classification for visual category recognition”
Hao Zhang, Alexander Berg, Michael Maire and Jitendra Malik · 2006
Earlier work this paper cites.
“What are the differences between long-term, short-term, and working memory?”
Nelson Cowan · 2008
Earlier work this paper cites.
“Online domain adaptation of a pre-trained cascade of classifiers”
Vidit Jain and Erik Learned-Miller · 2011
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate”
Dzmitry Bahdanau · 2014
Earlier work this paper cites.
“Neural Turing Machines”, 2014
Alex Graves, Greg Wayne and Ivo Danihelka · 2014
Earlier work this paper cites.
“The structure of value: Accounting for taste”
George Mandler · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra and Antoine Bordes · 2014
Earlier work this paper cites.
“End-to-end memory networks”
Sainbayar Sukhbaatar, Jason Weston and Rob Fergus · 2015
Earlier work this paper cites.
“Learning to learn by gradient descent by gradient descent”
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew Hoffman, David Pfau, Tom Schaul, Brendan Shillingford and Nando De · 2016
Earlier work this paper cites.
“LSTM: A search space odyssey”
Klaus Greff, Rupesh Srivastava, Jan Koutník, Bas Steunebrink and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
“The LAMBADA dataset: Word prediction requiring a broad discourse context”
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernández · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2017
Earlier work this paper cites.
“Neural semantic encoders”
Tsendsuren Munkhdalai and Hong Yu · 2017
Earlier work this paper cites.
“Learning and memory: Basic principles, processes, and procedures”
W Terry · 2017
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Think you have solved question answering? try arc, the ai2 reasoning challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“Sigmoid-weighted linear units for neural network function approximation in reinforcement learning”
Stefan Elfwing, Eiji Uchibe and Kenji Doya · 2018
Earlier work this paper cites.
“On first-order meta-learning algorithms”
A Nichol · 2018
Earlier work this paper cites.
“The unreasonable effectiveness of the forget gate”
Jos Van and Joan Lasenby · 2018
Earlier work this paper cites.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins and Kristina Toutanova · 2019
Earlier work this paper cites.
“Transformer-XL: Attentive Language Models beyond a Fixed-Length Context”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime. Carbonell, Quoc Le and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
“Online model distillation for efficient video inference”
Ravi Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan and Kayvon Fatahalian · 2019
Earlier work this paper cites.
“Metalearned neural memory”
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang and Adam Trischler · 2019
Earlier work this paper cites.
“Social IQa: Commonsense Reasoning about Social Interactions”
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le and Yejin Choi · 2019
Earlier work this paper cites.
“Augmenting self-attention with persistent memory”
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou and Armand Joulin · 2019
Earlier work this paper cites.
“R-transformer: Recurrent neural network enhanced transformer”
Zhiwei Wang, Yao Ma, Zitao Liu and Jiliang Tang · 2019
Earlier work this paper cites.
“Long-term feature banks for detailed video understanding”
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl and Ross Girshick · 2019
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Earlier work this paper cites.
“Fast context adaptation via meta-learning”
Luisa Zintgraf, Kyriacos Shiarli, Vitaly Kurin, Katja Hofmann and Shimon Whiteson · 2019
Earlier work this paper cites.
“Piqa: Reasoning about physical commonsense in natural language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Earlier work this paper cites.
“The pile: An 800gb dataset of diverse text for language modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite and Noa Nabeshima · 2020
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Cited alongside, same era.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Cited alongside, same era.
“Generalization through Memorization: Nearest Neighbor Language Models”
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer and Mike Lewis · 2020
Cited alongside, same era.
“Self-attentive associative memory”
Hung Le, Truyen Tran and Svetha Venkatesh · 2020
Cited alongside, same era.
“Retrieval-augmented generation for knowledge-intensive nlp tasks”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih and Tim Rocktäschel · 2020
Cited alongside, same era.
“xLSTM: Extended Long Short-Term Memory”
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
“Mambamixer: Efficient selective state space models with dual token and channel selection”
Ali Behrouz, Michele Santacatterina and Ramin Zabih · 2024
Closest in time.
Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer and Gargi Gosh · 2024
Closest in time.
“Birth of a transformer: A memory viewpoint”
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou and Leon Bottou · 2024
Closest in time.
“RecurrentGemma: Moving Past Transformers for Efficient Open Language Models”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard and Zhou Yu · 2020
Cited alongside, same era.
“Scatterbrain: Unifying sparse and low-rank attention”
Beidi Chen, Tri Dao, Eric Winsor, Zhao Song, Atri Rudra and Christopher Ré · 2021
Cited alongside, same era.
“Rethinking Attention with Performers”
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell and Adrian Weller · 2021
Cited alongside, same era.
“Going beyond linear transformers with recurrent fast weight programmers”
Kazuki Irie, Imanol Schlag, Róbert Csordás and Jürgen Schmidhuber · 2021
Cited alongside, same era.
“RWKV-LM”, 2021
Bo Peng · 2021
Cited alongside, same era.
“Efficient content-based sparse attention with routing transformers”
Aurko Roy, Mohammad Saffar, Ashish Vaswani and David Grangier · 2021
Cited alongside, same era.
“Winogrande: An adversarial winograd schema challenge at scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2021
Cited alongside, same era.
Aleksandar Botev, Soham De, Samuel Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Sessa and Robert Dadashi · 2024
Closest in time.
“An Evolved Universal Transformer Memory”
Edoardo Cetin, Qi Sun, Tianyu Zhao and Yujin Tang · 2024
Closest in time.
“FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”
Tri Dao · 2024
Closest in time.
Tri Dao and Albert Gu · 2024
Closest in time.
“Griffin: Mixing gated linear recurrences with local attention for efficient language models”
Soham De, Samuel Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen and Srivatsan Srinivasan · 2024
Closest in time.
“Flex Attention: A Programming Model for Generating Optimized Attention Kernels”
Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang and Horace He · 2024
Closest in time.
“Hymba: A Hybrid-head Architecture for Small Language Models”
Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Mahabaleshwarkar, Shih-Yang Liu, Matthijs Van, Min-Hung Chen and Yoshi Suhara · 2024
Closest in time.
“Mamba: Linear-Time Sequence Modeling with Selective State Spaces”
Albert Gu and Tri Dao · 2024
Closest in time.
“LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models”
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji and Sinong Wang · 2024
Closest in time.
“CAMELoT: Towards Large Language Models with Training-Free Consolidated Associative Memory”
Zexue He, Leonid Karlinsky, Donghyun Kim, Julian McAuley, Dmitry Krotov and Rogerio Feris · 2024
Closest in time.
“RULER: What’s the Real Context Size of Your Long-Context Language Models?”
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia and Boris Ginsburg · 2024
Closest in time.
“PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels”
Praneeth Kacham, Vahab Mirrokni and Peilin Zhong · 2024
Closest in time.
“BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack”
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin and Mikhail Burtsev · 2024
Closest in time.
“Learning, Forgetting, Remembering: Insights From Tracking LLM Memorization During Training”
Danny Leybzon and Corentin Kervadec · 2024
Closest in time.
“Longhorn: State space models are amortized online learners”
Bo Liu, Rui Wang, Lemeng Wu, Yihao Feng, Peter Stone and Qiang Liu · 2024
Closest in time.
“Lost in the middle: How language models use long contexts”
Nelson Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang · 2024
Closest in time.
“The Illusion of State in State-Space Models”
William Merrill, Jackson Petty and Ashish Sabharwal · 2024
Closest in time.
“Leave no context behind: Efficient infinite context transformers with infini-attention”
Tsendsuren Munkhdalai, Manaal Faruqui and Siddharth Gopal · 2024
Closest in time.
“Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution”
Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Michael Wornow, Callum Birch-Sykes, Stefano Massaroli, Aman Patel, Clayton Rabideau and Yoshua Bengio · 2024
Closest in time.
“SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series”, 2024
Badri. Patro and Vijay. Agneeswaran · 2024
Closest in time.
“The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale”
Guilherme Penedo, Hynek Kydlíček, Loubna allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Werra and Thomas Wolf · 2024
Closest in time.
“Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence”
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Xingjian Du, Teddy Ferdinan and Haowen Hou · 2024
Closest in time.
“Exploring Transformer Extrapolation”
Zhen Qin, Yiran Zhong and Hui Deng · 2024
Closest in time.
“Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling”
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang and Weizhu Chen · 2024
Closest in time.
“Associative recurrent memory transformer”
Ivan Rodkin, Yuri Kuratov, Aydar Bulatov and Mikhail Burtsev · 2024
Closest in time.
“Rethinking llm memorization through the lens of adversarial compression”
Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary Lipton and J Kolter · 2024
Closest in time.
“Beyond Memorization: Violating Privacy via Inference with Large Language Models”
Robin Staab, Mark Vero, Mislav Balunovic and Martin Vechev · 2024
Closest in time.
“Learning to (learn at test time): Rnns with expressive hidden states”
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang and Sanmi Koyejo · 2024
Closest in time.
“Gemma: Open models based on gemini research and technology”
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Kale and Juliette Love · 2024
Closest in time.
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Tommaso Guidi, Marco Gori and Stefano Melacci · 2024
Closest in time.
“LongSSM: On the Length Extension of State-space Models in Language Modelling”
Shida Wang · 2024
Closest in time.
“MEMORYLLM: Towards Self-Updatable Large Language Models”
Yu Wang, Yifan Gao, Xiusi Chen, Haoming Jiang, Shiyang Li, Jingfeng Yang, Qingyu Yin, Zheng Li, Xian Li, Bing Yin, Jingbo Shang and Julian McAuley · 2024
Closest in time.
“Towards LifeSpan Cognitive Systems”
Yu Wang, Chi Han, Tongtong Wu, Xiaoxin He, Wangchunshu Zhou, Nafis Sadeq, Xiusi Chen, Zexue He, Wei Wang and Gholamreza Haffari · 2024
Closest in time.
“Efficient Streaming Language Models with Attention Sinks”
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han and Mike Lewis · 2024
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang and Haoran Wei · 2024
Closest in time.
“Gated Delta Networks: Improving Mamba2 with Delta Rule”
Songlin Yang, Jan Kautz and Ali Hatamizadeh · 2024
Closest in time.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Closest in time.
“Parallelizing Linear Transformers with the Delta Rule over Sequence Length”
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen and Yoon Kim · 2024
Closest in time.
“B’MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory”
Luca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar, Yantao Shen, Benjamin Bowman, Matthew Trager, Alessandro Achille and Stefano Soatto · 2024
Closest in time.
Jianyu Zhang, Niklas Nolte, Ranajoy Sadhukhan, Beidi Chen and Léon Bottou · 2024
Closest in time.