Fetching the paper…
Reading the bibliography…
Recent hybrid models combining Linear State Space Models (SSMs) with self-attention mechanisms have demonstrated impressive results across a range of sequence modeling tasks.
Human memory: A proposed system and its control processes
Richard C. Atkinson and Richard M. Shiffrin · 1968
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
The human knowledge compression contest
Marcus Hutter · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
The acl anthology network corpus
Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
A. Graves · 2016
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, E. Uchibe, and K. Doya · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre · 2018
Earlier work this paper cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R Bowman · 2018
Earlier work this paper cites.
Towards flatter loss surface via nonmonotonic learning rate scheduling
Sihyeon Seong, Yegang Lee, Youngwook Kee, Dongyoon Han, and Junmo Kim · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, S. Levine, Yoshua Bengio, and B. Scholkopf · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Cited alongside, same era.
Etc: Encoding long and structured inputs in transformers
J. Ainslie, Santiago Ontañón, Chris Alberti, V. Cvicek, Zachary Kenneth Fisher, Philip Pham, Anirudh Ravula, Sumit K. Sanghai, Qifan Wang, and Li Yang · 2020
Cited alongside, same era.
Not all memories are created equal: Learning to forget by expiring
Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur D. Szlam, J. Weston, and Angela Fan · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Primer: Searching for efficient transformers for language modeling
David R. So, Wojciech Ma’nke, Hanxiao Liu, Zihang Dai, Noam M. Shazeer, and Quoc V. Le · 2021
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
Why neural networks find simple solutions: The many regularizers of geometric complexity
Benoit Dherin, Michael Munn, Mihaela Rosca, and David Barrett · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit gradient regularization
D. Barrett and B. Dherin · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Sparse gpu kernels for deep learning
Trevor Gale, M. Zaharia, C. Young, and Erich Elsen · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, M. Krikun, Noam M. Shazeer, and Z. Chen · 2020
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Later among the works it cites.
Diagonal state spaces are as effective as structured state spaces
Ankit Gupta and Jonathan Berant · 2022
Later among the works it cites.
On the parameterization and initialization of diagonal state space models
Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré · 2022
Later among the works it cites.
Transkimmer: Transformer learns to layer-wise skim
Yue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin, and Minyi Guo · 2022
Later among the works it cites.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V. Le · 2022
Later among the works it cites.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2022
Later among the works it cites.
Dynamic sparse attention for scalable transformer acceleration
Liu Liu, Zheng Qu, Zhaodong Chen, Fengbin Tu, Yufei Ding, and Yuan Xie · 2022
Later among the works it cites.
Language model pre-training with sparse latent typing
Liliang Ren, Zixuan Zhang, Han Wang, Clare Voss, ChengXiang Zhai, and Heng Ji · 2022
Later among the works it cites.
A survey on dynamic neural networks for natural language processing
Canwen Xu and Julian McAuley · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al · 2022
Later among the works it cites.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re · 2023
Closest in time.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Closest in time.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman · 2023
Closest in time.
State spaces aren’t enough: Machine translation needs attention
Ali Vardasbi, Telmo Pires, Robin M. Schmidt, and Stephan Peitz · 2023
Closest in time.