Fetching the paper…
Reading the bibliography…
A central problem in learning from sequential data is representing cumulative history in an incremental fashion as more data is processed.
Orthogonal Polynomials
G. Szegö · 1967
Earlier work this paper cites.
Oscillation and chaos in physiological control systems
Michael C Mackey and Leon Glass · 1977
Earlier work this paper cites.
Linear systems: A state variable approach with numerical implementation
Raymond A DeCarlo · 1989
Earlier work this paper cites.
Fourier analysis
Thomas William Körner · 1989
Earlier work this paper cites.
Generalized sliding FFT and its application to implementation of block LMS adaptive filters
Behrouz Farhang-Boroujeny and Saeed Gazor · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Chebyshev and Fourier spectral methods
John P Boyd · 2001
Earlier work this paper cites.
Digital signal processing: principles algorithms and applications
John G Proakis · 2001
Earlier work this paper cites.
The sliding DFT
Eric Jacobsen and Richard Lyons · 2003
Earlier work this paper cites.
An update to the sliding DFT
Eric Jacobsen and Richard Lyons · 2004
Earlier work this paper cites.
Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication
Herbert Jaeger and Harald Haas · 2004
Earlier work this paper cites.
Fast algorithms for the computation of sliding discrete sinusoidal transforms
Vitaly Kober · 2004
Earlier work this paper cites.
Mathematical methods for physicists
George B Arfken and Hans J Weber · 2005
Earlier work this paper cites.
Efficient computation of the running discrete Haar transform
Jose A Rosendo Macias and Antonio Gomez Exposito · 2005
Earlier work this paper cites.
Fast algorithms for the computation of sliding discrete Hartley transforms
Vitaly Kober · 2007
Earlier work this paper cites.
An efficient recursive algorithm and an explicit formula for calculating update vectors of running Walsh-Hadamard transform
Barzan Mozafari and Mohammad H Savoji · 2007
Earlier work this paper cites.
Performance recovery in digital implementation of analogue systems
Guofeng Zhang, Tongwen Chen, and Xiang Chen · 2007
Earlier work this paper cites.
A first course in the numerical analysis of differential equations
Arieh Iserles · 2009
Earlier work this paper cites.
Fast algorithm for Walsh Hadamard transform on sliding windows
Wanli Ouyang and Wai-Kuen Cham · 2009
Earlier work this paper cites.
Accurate, guaranteed stable, sliding discrete Fourier transform [DSP tips & tricks]
Krzysztof Duda · 2010
Earlier work this paper cites.
An introduction to orthogonal polynomials
T. S. Chihara · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Sliding conjugate symmetric sequency-ordered complex Hadamard transform: fast algorithm and applications
Jiasong Wu, Lu Wang, Guanyu Yang, Lotfi Senhadji, Limin Luo, and Huazhong Shu · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Fast computation of sliding discrete Tchebichef moments and its application in duplicated regions detection
Beijing Chen, Gouenou Coatrieux, Jiasong Wu, Zhifang Dong, Jean Louis Coatrieux, and Huazhong Shu · 2015
Earlier work this paper cites.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst · 2016
Cited alongside, same era.
LSTM: A search space odyssey
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber · 2016
Towards non-saturating recurrent units for modelling long-term dependencies
Sarath Chandar, Chinnadhurai Sankar, Eugene Vorontsov, Samira Ebrahimi Kahou, and Yoshua Bengio · 2019
Later among the works it cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Learning fast algorithms for linear transforms using butterfly factorizations
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré · 2019
Later among the works it cites.
Augmented neural ODEs
Emilien Dupont, Arnaud Doucet, and Yee Whye Teh · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noisy activation functions
Caglar Gulcehre, Marcin Moczulski, Misha Denil, and Yoshua Bengio · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Cited alongside, same era.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Chris Pal · 2016
Cited alongside, same era.
Dilated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang · 2017
Cited alongside, same era.
Gaussian quadrature for kernel features
Tri Dao, Christopher M De Sa, and Christopher Ré · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group
Mario Lezcano-Casado and David Martínez-Rubio · 2019
Later among the works it cites.
Latent ordinary differential equations for irregularly-sampled time series
Yulia Rubanova, Tian Qi Chen, and David K Duvenaud · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Later among the works it cites.
Approximation theory and approximation practice , volume 164
Lloyd N Trefethen · 2019
Later among the works it cites.
Legendre memory units: Continuous-time representation in recurrent neural networks
Aaron Voelker, Ivana Kajić, and Chris Eliasmith · 2019
Later among the works it cites.
Dynamical systems in spiking neuromorphic hardware
Aaron Russell Voelker · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 2019
Later among the works it cites.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2019
Later among the works it cites.
Butterfly transform: An efficient FFT based neural architecture design
Keivan Alizadeh, Ali Farhadi, and Mohammad Rastegari · 2020
Closest in time.
Accelerated gossip in networks of given dimension using Jacobi polynomial iterations
Raphaël Berthier, Francis Bach, and Pierre Gaillard · 2020
Closest in time.
Kaleidoscope: An efficient, learnable representation for all structured linear maps
Tri Dao, Nimit Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra, and Christopher Ré · 2020
Closest in time.
How to train your neural ODE: the world of Jacobian and kinetic regularization
Chris Finlay, Jörn-Henrik Jacobsen, Levon Nurbekyan, and Adam M Oberman · 2020
Closest in time.
Improving the gating mechanism of recurrent neural networks
Albert Gu, Caglar Gulcehre, Tom Le Paine, Matt Hoffman, and Razvan Pascanu · 2020
Closest in time.
Neural controlled differential equations for irregular time series
Patrick Kidger, James Morrill, James Foster, and Terry Lyons · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Closest in time.
Stefano Massaroli, Michael Poli, Michelangelo Bin, Jinkyoo Park, Atsushi Yamashita, and Hajime Asama · 2020
Closest in time.
SNODE: Spectral discretization of neural ODEs for system identification
Alessio Quaglino, Marco Gallieri, Jonathan Masci, and Jan Koutník · 2020
Closest in time.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2020
Closest in time.
Weak supervision as an efficient approach for automated seizure detection in electroencephalography
Khaled Saab, Jared Dunnmon, Christopher Ré, Daniel Rubin, and Christopher Lee-Messer · 2020
Closest in time.
Approximation capabilities of neural ordinary differential equations
Han Zhang, Xi Gao, Jacob Unterman, and Tom Arodz · 2020
Closest in time.