Fetching the paper…
Reading the bibliography…
Structured state-space models (SSMs) such as S4, stemming from the seminal work of Gu et al., are gaining popularity as effective approaches for modeling sequential data.
An inequality of the hölder type, connected with stieltjes integration
Laurence C Young · 1936
Earlier work this paper cites.
Integration of paths–a faithful representation of paths by noncommutative formal power series
Kuo-Tsai Chen · 1958
Earlier work this paper cites.
Fading memory and the problem of approximating nonlinear operators with volterra series
Stephen Boyd and Leon Chua · 1985
Earlier work this paper cites.
On the computational power of neural nets
Hava T Siegelmann and Eduardo D Sontag · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Differential equations driven by rough signals (i): An extension of an inequality of lc young
Terry Lyons · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and J"urgen Schmidhuber · 1997
Earlier work this paper cites.
An elementary proof of a theorem of johnson and lindenstrauss
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
Differential equations driven by rough paths
Terry J Lyons, Michael Caruana, and Thierry Lévy · 2007
Earlier work this paper cites.
Reservoir computing approaches to recurrent neural network training
Mantas Lukoveviius and Herbert Jaeger · 2009
Earlier work this paper cites.
Uniqueness for the signature of a path of bounded variation and the reduced path group
Ben Hambly and Terry Lyons · 2010
Earlier work this paper cites.
Multidimensional Stochastic Processes as Rough Paths: Theory and Applications
Peter K. Friz and Nicolas B. Victoir · 2010
Earlier work this paper cites.
Reproducing Kernel Hilbert Spaces in Probability and Statistics
A. Berlinet and C. Thomas-Agnan · 2011
Earlier work this paper cites.
Efficient BackProp , pages 9–48
Yann A. LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2017
Earlier work this paper cites.
Can recurrent neural networks warp time?, 2018
Corentin Tallec and Yann Ollivier · 2018
Earlier work this paper cites.
On the computational power of rnns
Samuel A Korsky and Robert C Berwick · 2019
Earlier work this paper cites.
Universal approximation of input-output maps by temporal convolutional nets
Joshua Hanson and Maxim Raginsky · 2019
Earlier work this paper cites.
Deep signature transforms
Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Neil Houlsby, Sylvain Gelly, Xiaohua Zhang, and Jakob Uszkoreit · 2020
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2020
Earlier work this paper cites.
Neural controlled differential equations for irregular time series
Patrick Kidger, James Morrill, James Foster, and Terry Lyons · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Universal simulation of stable dynamical systems by recurrent neural nets
Joshua Hanson and Maxim Raginsky · 2020
Cited alongside, same era.
Embedding and learning with signatures, 2020
Adeline Fermanian · 2020
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Re · 2021
Cited alongside, same era.
Neural rough differential equations for long time series
James Morrill, Cristopher Salvi, Patrick Kidger, and James Foster · 2021
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models, 2023
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
Gateloop: Fully data-controlled linear recurrence for sequence modeling, 2023
Tobias Katsch · 2023
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Framing rnn as a kernel method: A neural ode approach
Adeline Fermanian, Pierre Marion, Jean-Philippe Vert, and Gérard Biau · 2021
Cited alongside, same era.
Discrete-time signatures and randomness in reservoir computing
Christa Cuchiero, Lukas Gonon, Lyudmila Grigoryeva, Juan-Pablo Ortega, and Josef Teichmann · 2021
Cited alongside, same era.
Distribution regression for sequential data, 2021
Maud Lemercier, Cristopher Salvi, Theodoros Damoulas, Edwin V. Bonilla, and Terry Lyons · 2021
Cited alongside, same era.
Sk-tree: a systematic malware detection algorithm on streaming trees via the signature kernel
Thomas Cochrane, Peter Foster, Varun Chhabra, Maud Lemercier, Terry Lyons, and Cristopher Salvi · 2021
Cited alongside, same era.
On words of non-Hermitian random matrices
Guillaume Dubach and Yuval Peled · 2021
Cited alongside, same era.
S4nd: Modeling images and videos as multidimensional signals using state spaces
Eric Nguyen, Karan Goel, Albert Gu, Gordon W. Downs, Preey Shah, Tri Dao, Stephen A. Baccus, and Christopher Ré · 2022
Cited alongside, same era.
A neural rde approach for continuous-time non-markovian stochastic control problems
Melker Hoglund, Emilio Ferrucci, Camilo Hernandez, Aitor Muguruza Gonzalez, Cristopher Salvi, Leandro Sanchez-Betancourt, and Yufei Zhang · 2023
Later among the works it cites.
Shida Wang and Beichen Xue · 2023
Later among the works it cites.
On the effectiveness of randomized signatures as reservoir for learning rough dynamics
Enea Monzio Compagnoni, Anna Scampicchio, Luca Biggio, Antonio Orvieto, Thomas Hofmann, and Josef Teichmann · 2023
Later among the works it cites.
New directions in the applications of rough path theory
Adeline Fermanian, Terry Lyons, James Morrill, and Cristopher Salvi · 2023
Later among the works it cites.
Neural signature kernels as infinite-width-depth-limits of controlled resnets, 2023
Nicola Muca Cirone, Maud Lemercier, and Cristopher Salvi · 2023
Later among the works it cites.
Non-adversarial training of neural sdes with signature kernel scores
Zacharia Issa, Blanka Horvath, Maud Lemercier, and Cristopher Salvi · 2023
Later among the works it cites.
Hgrn2: Gated linear rnns with state expansion
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong · 2024
Closest in time.
Griffin: Mixing gated linear recurrences with local attention for efficient language models, 2024
Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre · 2024
Closest in time.
Were rnns all we needed?, 2024
Leo Feng, Frederick Tung, Mohamed Osama Ahmed, Yoshua Bengio, and Hossein Hajimirsadegh · 2024
Closest in time.
Log neural controlled differential equations: The lie brackets make a difference
Benjamin Walker, Andrew D. McLeod, Tiexin Qin, Yichuan Cheng, Haoliang Li, and Terry Lyons · 2024
Closest in time.
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Nicolas Zucchet and Antonio Orvieto · 2024
Closest in time.
Repeat after me: Transformers are better than state space models at copying
Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Closest in time.
Lecture notes on rough paths and applications to machine learning
Thomas Cass and Cristopher Salvi · 2024
Closest in time.
A path-dependent pde solver based on signature kernels
Alexandre Pannier and Cristopher Salvi · 2024
Closest in time.
Signature kernel conditional independence tests in causal discovery for stochastic processes
Georg Manten, Cecilia Casolo, Emilio Ferrucci, Søren Wengel Mogensen, Cristopher Salvi, and Niki Kilbertus · 2024
Closest in time.
Expressive power of randomized signature
Christa Cuchiero, Lukas Gonon, Lyudmila Grigoryeva, Juan-Pablo Ortega, and Josef Teichmann · 2024
Closest in time.