Fetching the paper…
Reading the bibliography…
Sequential models, such as Recurrent Neural Networks and Neural Ordinary Differential Equations, have long suffered from slow training due to their inherent sequential nature.
Parallel algorithms for initial-value problems for difference and differential equations
Alfredo Bellen and Marino Zennaro · 1989
Earlier work this paper cites.
Prefix sums and their applications
Guy E Blelloch · 1990
Earlier work this paper cites.
An introduction to numerical analysis
Kendall Atkinson · 1991
Earlier work this paper cites.
A parallel shooting technique for solving dissipative ode’s
P Chartier and B Philippe · 1993
Earlier work this paper cites.
Parallel multiple shooting for the solution of initial value problems
Martin Kiehl · 1994
Earlier work this paper cites.
Iterative methods for linear and nonlinear equations
Carl T Kelley · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Résolution d’edp par un schéma en temps “pararéel”
Jacques-Louis Lions, Yvon Maday, and Gabriel Turinici · 2001
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Analysis of the parareal time-parallel time-integration method
Martin J Gander and Stefan Vandewalle · 2007
Earlier work this paper cites.
A dictionary of behavioral motifs reveals clusters of genes affecting caenorhabditis elegans locomotion
André EX Brown, Eviatar I Yemini, Laura J Grundy, Tadas Jucikas, and William R Schafer · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Tensorflow: learning functions at scale
Martín Abadi · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Simple recurrent units for highly parallelizable recurrence
Tao Lei, Yu Zhang, Sida I Wang, Hui Dai, and Yoav Artzi · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The uea multivariate time series classification archive, 2018
Anthony Bagnall, Hoang Anh Dau, Jason Lines, Michael Flynn, James Large, Aaron Bostrom, Paul Southam, and Eamonn Keogh · 2018
Cited alongside, same era.
Trellis networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
Neural ordinary differential equations
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Cited alongside, same era.
Compiling machine learning programs via high-level tracing
Roy Frostig, Matthew James Johnson, and Chris Leary · 2018
Cited alongside, same era.
Ffjord: Free-form continuous dynamics for scalable reversible generative models
Will Grathwohl, Ricky TQ Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud · 2018
Cited alongside, same era.
Neural rough differential equations for long time series
James Morrill, Cristopher Salvi, Patrick Kidger, and James Foster · 2021
Later among the works it cites.
Flexconv: Continuous kernel convolutions with differentiable kernel sizes
David W Romero, Robert-Jan Bruintjes, Jakub M Tomczak, Erik J Bekkers, Mark Hoogendoorn, and Jan C van Gemert · 2021
Later among the works it cites.
Moser flow: Divergence-based generative modeling on manifolds
Noam Rozen, Aditya Grover, Maximilian Nickel, and Yaron Lipman · 2021
Later among the works it cites.
Unicornn: A recurrent model for learning very long time dependencies
T Konstantin Rusch and Siddhartha Mishra · 2021
Later among the works it cites.
Long expressive memory for sequence modeling
T Konstantin Rusch, Siddhartha Mishra, N Benjamin Erichson, and Michael W Mahoney · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2018
Cited alongside, same era.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu Trinh, Andrew Dai, Thang Luong, and Quoc Le · 2018
Cited alongside, same era.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Cited alongside, same era.
Symplectic recurrent neural networks
Zhengdao Chen, Jianyu Zhang, Martin Arjovsky, and Léon Bottou · 2019
Cited alongside, same era.
Hamiltonian neural networks
Samuel Greydanus, Misko Dzamba, and Jason Yosinski · 2019
Cited alongside, same era.
Latent ordinary differential equations for irregularly-sampled time series
Yulia Rubanova, Ricky TQ Chen, and David K Duvenaud · 2019
Cited alongside, same era.
Multiscale deep equilibrium models
Shaojie Bai, Vladlen Koltun, and J Zico Kolter · 2020
Cited alongside, same era.
Accelerating feedforward computation via parallel nonlinear equation solving
Yang Song, Chenlin Meng, Renjie Liao, and Stefano Ermon · 2021
Later among the works it cites.
Approximate fixed-points in recurrent neural networks
Zhengxiong Wang and Anton Ragni · 2021
Later among the works it cites.
Matching normalizing flows and probability paths on manifolds
Heli Ben-Hamu, Samuel Cohen, Joey Bose, Brandon Amos, Maximillian Nickel, Aditya Grover, Ricky TQ Chen, and Yaron Lipman · 2022
Later among the works it cites.
On the parameterization and initialization of diagonal state space models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré · 2022
Later among the works it cites.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2022
Later among the works it cites.
Encoding recurrence into transformers
Feiqing Huang, Kexin Lu, CAI Yuxi, Zhen Qin, Yanwen Fang, Guangjian Tian, and Guodong Li · 2022
Later among the works it cites.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2022
Later among the works it cites.
Mgnni: Multiscale graph neural networks with implicit layers
Juncheng Liu, Bryan Hooi, Kenji Kawaguchi, and Xiaokui Xiao · 2022
Later among the works it cites.
Finde: Neural differential equations for finding and preserving invariant quantities
Takashi Matsubara and Takaharu Yaguchi · 2022
Later among the works it cites.
Parallel training of gru networks with a multi-grid solver for long sequences
Gordon Euhyun Moon and Eric C Cyr · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy TH Smith, Andrew Warrington, and Scott W Linderman · 2022
Later among the works it cites.
Gradient regularization of newton method with bregman distances
Nikita Doikov and Yurii Nesterov · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Closest in time.
Sequence modeling with multiresolution convolutional memory
Jiaxin Shi, Ke Alexander Wang, and Emily Fox · 2023
Closest in time.
Parallel sampling of diffusion models
Andy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh, and Nima Anari · 2023
Closest in time.