Fetching the paper…
Reading the bibliography…
Models using structured state space sequence (S4) layers have achieved state-of-the-art performance on long-range sequence modeling tasks.
Parallel prefix computation
Richard Ladner and Michael Fischer · 1980
Earlier work this paper cites.
Prefix sums and their applications
Guy Blelloch · 1990
Earlier work this paper cites.
Parallel computing using the prefix problem
Sivaramakrishnan Lakshmivarahan and Sudarshan Dhall · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A first course in the numerical analysis of differential equations
Arieh Iserles · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The ACL anthology network corpus
Dragomir Radev, Pradeep Muthukrishnan, and Vahed Qazvinian · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond Daly, Peter Pham, Dan Huang, Andrew Ng, and Christopher Potts · 2011
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Dilated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Ricky Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud · 2018
Earlier work this paper cites.
Simple recurrent units for highly parallelizable recurrence
Tao Lei, Yu Zhang, Sida Wang, Hui Dai, and Yoav Artzi · 2018
Earlier work this paper cites.
Independently recurrent neural network (INDRNN): Building a longer and deeper RNN
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao · 2018
Earlier work this paper cites.
Learning long-range spatial dependencies with horizontal gated recurrent units
Drew Linsley, Junkyung Kim, Vijay Veerabadran, Charles Windolf, and Thomas Serre · 2018
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
ListOps: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel Bowman · 2018
Cited alongside, same era.
Learning longer-term dependencies in RNNs with auxiliary losses
Trieu Trinh, Andrew Dai, Thang Luong, and Quoc Le · 2018
Cited alongside, same era.
Speech Commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Cited alongside, same era.
Trellis networks for sequence modeling
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2019
Cited alongside, same era.
Recurrent kalman networks: Factorized inference in high-dimensional deep feature spaces
Philipp Becker, Harit Pandya, Gregor Gebhardt, Cheng Zhao, C James Taylor, and Gerhard Neumann · 2019
Cited alongside, same era.
AntisymmetricRNN: A dynamical system view on recurrent neural networks
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy Colwell, and Adrian Weller · 2021
Later among the works it cites.
Lipschitz recurrent neural networks
N. Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael Mahoney · 2021
Later among the works it cites.
Luna: Linear unified nested attention
Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma, and Luke Zettlemoyer · 2021
Later among the works it cites.
In-depth benchmarking of deep neural network architectures for ecg diagnosis
Naoki Nonaka and Jun Seita · 2021
Later among the works it cites.
Flexconv: Continuous kernel convolutions with differentiable kernel sizes
David Romero, Robert-Jan Bruintjes, Jakub Mikolaj Tomczak, Erik Bekkers, Mark Hoogendoorn, and Jan van Gemert · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bo Chang, Minmin Chen, Eldad Haber, and Ed Chi · 2019
Cited alongside, same era.
GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series
Edward De Brouwer, Jaak Simm, Adam Arany, and Yves Moreau · 2019
Cited alongside, same era.
Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group
Mario Lezcano-Casado and David Martınez-Rubio · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Latent ordinary differential equations for irregularly-sampled time series
Yulia Rubanova, Ricky Chen, and David Duvenaud · 2019
Cited alongside, same era.
Legendre Memory Units: Continuous-time representation in recurrent neural networks
Aaron Voelker, Ivana Kajić, and Chris Eliasmith · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew Peters, and Arman Cohan · 2020
Cited alongside, same era.
T. Konstantin Rusch and Siddhartha Mishra · 2021
Later among the works it cites.
Multi-time attention networks for irregularly sampled time series
Satya Narayan Shukla and Benjamin Marlin · 2021
Later among the works it cites.
Long Range Arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Later among the works it cites.
Nyströmformer: A Nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
H-transformer-1d: Fast one-dimensional hierarchical attention for sequences
Zhenhai Zhu and Radu Soricut · 2021
Later among the works it cites.
It’s raw! Audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Re · 2022
Closest in time.
On the parameterization and initialization of diagonal state space models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré · 2022
Closest in time.
Diagonal state spaces are as effective as structured state spaces
Ankit Gupta, Albert Gu, and Jonathan Berant · 2022
Closest in time.
Long movie clip classification with state-space video models
Md Mohaiminul Islam and Gedas Bertasius · 2022
Closest in time.
FNet: Mixing tokens with Fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon · 2022
Closest in time.
The Annotated S4
Sasha Rush and Sidd Karamcheti · 2022
Closest in time.
Modeling irregular time series with continuous recurrent units
Mona Schirmer, Mazin Eltayeb, Stefan Lessmann, and Maja Rudolph · 2022
Closest in time.
How to train your HIPPO: State space models with generalized orthogonal basis projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Re · 2023
Closest in time.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2023
Closest in time.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Closest in time.