Fetching the paper…
Reading the bibliography…
This paper investigates the gap in representation powers of Recurrent Neural Networks (RNNs) and Transformers in the context of solving algorithmic problems.
Non-holographic associative memory
David J Willshaw, O Peter Buneman, and Hugh Christopher Longuet-Higgins · 1969
Earlier work this paper cites.
Selection and sorting with limited storage
J Ian Munro and Mike S Paterson · 1980
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
John J Hopfield · 1982
Earlier work this paper cites.
The space complexity of approximating the frequency moments
Noga Alon, Yossi Matias, and Mario Szegedy · 1996
Earlier work this paper cites.
Computing on data streams
Monika Rauch Henzinger, Prabhakar Raghavan, and Sridhar Rajagopalan · 1998
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Parallel models of associative memory: updated edition
Geoffrey E Hinton and James A Anderson · 2014
Earlier work this paper cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate, 2016
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Modeling recurrence for transformer
Jie Hao, Xing Wang, Baosong Yang, Longyue Wang, Jinfeng Zhang, and Zhaopeng Tu · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages, 2020
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training, 2020
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Earlier work this paper cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity, 2020
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
On the curse of memory in recurrent neural networks: Approximation and optimization analysis
Zhong Li, Jiequn Han, Weinan E, and Qianxiao Li · 2021
Earlier work this paper cites.
Stable, fast and accurate: Kernelized attention with relative positional encoding, 2021
Shengjie Luo, Shanda Li, Tianle Cai, Di He, Dinglan Peng, Shuxin Zheng, Guolin Ke, Liwei Wang, and Tie-Yan Liu · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Earlier work this paper cites.
Random feature attention, 2021
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong · 2021
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens, 2022
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre · 2022
Cited alongside, same era.
Recurrent memory transformer, 2022
Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Re · 2022
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity, 2022
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Cited alongside, same era.
Retentive network: A successor to transformer for large language models, 2023
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer, 2023
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer
Zhengbao Jiang, Luyu Gao, Zhiruo Wang, Jun Araki, Haibo Ding, Jamie Callan, and Graham Neubig · 2022
Cited alongside, same era.
On the approximation properties of recurrent encoder-decoder architectures
Zhong Li, Haotian Jiang, and Qianxiao Li · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits, 2022
William Merrill, Ashish Sabharwal, and Noah A. Smith · 2022
Cited alongside, same era.
Sumformer: Universal approximation for efficient transformers, 2023
Silas Alberti, Niclas Dern, Laura Thesing, and Gitta Kutyniok · 2023
Cited alongside, same era.
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models, 2023
Daniel Y. Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré · 2023
Cited alongside, same era.
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks, 2023
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages, 2023
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms, 2024
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas · 2024
Closest in time.
xlstm: Extended long short-term memory, 2024
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Closest in time.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models, 2024
Aviv Bick, Kevin Y. Li, Eric P. Xing, J. Zico Kolter, and Albert Gu · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach · 2024
Closest in time.
In search of needles in a 11m haystack: Recurrent memory finds what llms miss, 2024
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems, 2024
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model, 2024
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham · 2024
Closest in time.
Transformers are multi-state rnns, 2024
Matanel Oren, Michael Hassid, Yossi Adi, and Roy Schwartz · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks, 2024
Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling, 2024
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen · 2024
Closest in time.
Transformers, parallel computation, and logarithmic depth, 2024
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
An empirical study of mamba-based language models, 2024
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, Garvit Kulshreshtha, Vartika Singh, Jared Casper, Jan Kautz, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Closest in time.
Chain-of-thought reasoning without prompting, 2024
Xuezhi Wang and Denny Zhou · 2024
Closest in time.
Do efficient transformers really save computation?, 2024
Kai Yang, Jan Ackermann, Zhenyu He, Guhao Feng, Bohang Zhang, Yunzhen Feng, Qiwei Ye, Di He, and Liwei Wang · 2024
Closest in time.