Fetching the paper…
Reading the bibliography…
Transformers have surpassed RNNs in popularity due to their superior abilities in parallel training and long-term dependency modeling.
The exponentially weighted moving average
J. Stuart Hunter · 1986
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter and Yoshua Bengio · 2001
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Çaglar Gülçehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
A clockwork RNN
Jan Koutník, Klaus Greff, Faustino J. Gomez, and Jürgen Schmidhuber · 2014
Earlier work this paper cites.
Lstm: A search space odyssey
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton · 2015
Earlier work this paper cites.
Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding
Yajie Miao, Mohammad Gowayyed, and Florian Metze · 2015
Earlier work this paper cites.
Weather forecasting using deep learning techniques
Afan Galih Salman, Bayu Kanigoro, and Yaya Heryadi · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martín Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Strongly-typed recurrent neural networks
David Balduzzi and Muhammad Ghifary · 2016
Earlier work this paper cites.
Dilated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A. Hasegawa-Johnson, and Thomas S. Huang · 2017
Earlier work this paper cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Stock price prediction using lstm, rnn and cnn-sliding window model
Sreelekshmy Selvin, R Vinayakumar, EA Gopalakrishnan, Vijay Krishna Menon, and KP Soman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Simple recurrent units for highly parallelizable recurrence
Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, and Yoav Artzi · 2018
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
Can recurrent neural networks warp time?
Corentin Tallec and Yann Ollivier · 2018
Earlier work this paper cites.
The unreasonable effectiveness of the forget gate
Jos van der Westhuizen and Joan Lasenby · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron C. Courville · 2019
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Improving the gating mechanism of recurrent neural networks
Albert Gu, Çaglar Gülçehre, Thomas Paine, Matt Hoffman, and Razvan Pascanu · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
GLU variants improve transformer
Noam Shazeer · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Cited alongside, same era.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V Le · 2022
Later among the works it cites.
FNet: Mixing tokens with Fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon · 2022
Later among the works it cites.
What makes convolutional models great on long sequence modeling?
Yuhong Li, Tianle Cai, Yi Zhang, De huai Chen, and Debadeepta Dey · 2022
Later among the works it cites.
Neural architecture search on efficient transformers and beyond
Zexiang Liu, Dong Li, Kaiyue Lu, Zhen Qin, Weixuan Sun, Jiacheng Xu, and Yiran Zhong · 2022
Later among the works it cites.
Linear video transformer with feature fixation
Kaiyue Lu, Zexiang Liu, Jianyuan Wang, Weixuan Sun, Zhen Qin, Dong Li, Xuyang Shen, Hui Deng, Xiaodong Han, Yuchao Dai, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Cited alongside, same era.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state-space layers, 2021
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Rethinking positional encoding in language pre-training
Guolin Ke, Di He, and Tie-Yan Liu · 2021
Cited alongside, same era.
Later among the works it cites.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2022
Later among the works it cites.
Fine-tuning pre-trained transformers into decaying fast weights
Huanru Henry Mao · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2022
Later among the works it cites.
The devil in linear transformer
Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong · 2022
Later among the works it cites.
cosformer: Rethinking softmax in attention
Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman · 2022
Later among the works it cites.
Locality matters: A locality-biased linear attention for automatic speech recognition
Jingyu Sun, Guiping Zhong, Dinghao Zhou, Baoxiang Li, and Yiran Zhong · 2022
Later among the works it cites.
Junxiong Wang, Jing Nathan Yan, Albert Gu, and Alexander M. Rush · 2022
Later among the works it cites.
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2022
Later among the works it cites.
Simple hardware-efficient long convolutions for sequence modeling
Daniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré · 2023
Closest in time.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2023
Closest in time.
A unified view of long-sequence models towards modeling million-scale dependencies
Hongyu He and Marko Kabic · 2023
Closest in time.
Encoding recurrence into transformers
Feiqing Huang, Kexin Lu, Yuxi CAI, Zhen Qin, Yanwen Fang, Guangjian Tian, and Guodong Li · 2023
Closest in time.
On the universality of linear recurrences followed by nonlinear projections
Antonio Orvieto, Soham De, Çaglar Gülçehre, Razvan Pascanu, and Samuel L. Smith · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Çaglar Gülçehre, Razvan Pascanu, and Soham De · 2023
Closest in time.
RWKV: reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran G. V., Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan S. Wind, Stanislaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, and Rui-Jie Zhu · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
Toeplitz neural network for sequence modeling
Zhen Qin, Xiaodong Han, Weixuan Sun, Bowen He, Dong Li, Dongxu Li, Yuchao Dai, Lingpeng Kong, and Yiran Zhong · 2023
Closest in time.
Scaling transnormer to 175 billion parameters
Zhen Qin, Dong Li, Weigao Sun, Weixuan Sun, Xuyang Shen, Xiaodong Han, Yunshen Wei, Baohong Lv, Fei Yuan, Xiao Luo, Yu Qiao, and Yiran Zhong · 2023
Closest in time.
Linearized relative positional encoding
Zhen Qin, Weixuan Sun, Kaiyue Lu, Hui Deng, Dongxu Li, Xiaodong Han, Yuchao Dai, Lingpeng Kong, and Yiran Zhong · 2023
Closest in time.
Accelerating toeplitz neural network with constant-time inference complexity
Zhen Qin and Yiran Zhong · 2023
Closest in time.
Vicinity vision transformer
Weixuan Sun, Zhen Qin, Hui Deng, Jianyuan Wang, Yi Zhang, Kaihao Zhang, Nick Barnes, Stan Birchfield, Lingpeng Kong, and Yiran Zhong · 2023
Closest in time.