Fetching the paper…
Reading the bibliography…
The recently-introduced class of ordinary differential equation networks (ODE-Nets) establishes a fruitful connection between deep learning and dynamical systems.
Calculus of finite differences
C. Jordan and K. Jordán · 1965
Earlier work this paper cites.
Solving ordinary differential equations. I. Nonstiff problems
E. Hairer, S. P. Nørsett, and G. Wanner · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Norm inequalities for derivatives and differences
M. K. Kwong and A. Zettl · 2006
Earlier work this paper cites.
Fully automatic h p hp -adaptivity in three dimensions
W. Rachowicz, D. Pardo, and L. Demkowicz · 2006
Earlier work this paper cites.
Nonlinear finite element methods
P. Wriggers · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Approximate computation and implicit regularization for very large-scale data analysis
M. W. Mahoney · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Stable architectures for deep neural networks
E. Haber and L. Ruthotto · 2017
Earlier work this paper cites.
Regularization for deep learning: A taxonomy
J. Kukacka, V. Golkov, and D. Cremers · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
B. Neyshabur · 2017
Earlier work this paper cites.
A proposal on machine learning via dynamical systems
E. Weinan · 2017
Earlier work this paper cites.
The GUM corpus: Creating multilayer resources in the classroom
A. Zeldes · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Cited alongside, same era.
Reversible architectures for arbitrarily deep residual neural networks
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham · 2018
Cited alongside, same era.
Multi-level residual networks from dynamical systems view
B. Chang, L. Meng, E. Haber, F. Tung, and D. Begert · 2018
Cited alongside, same era.
Neural ordinary differential equations
T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud · 2018
Cited alongside, same era.
FFJORD: Free-form continuous dynamics for scalable reversible generative models
W. Grathwohl, R. T. Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud · 2018
Cited alongside, same era.
Lagrangian neural networks
M. Cranmer, S. Greydanus, S. Hoyer, P. Battaglia, D. Spergel, and S. Ho · 2020
Later among the works it cites.
Depth-adaptive transformer
M. Elbayad, J. Gu, E. Grave, and M. Auli · 2020
Later among the works it cites.
Lipschitz recurrent neural networks
N. B. Erichson, O. Azencot, A. Queiruga, and M. W. Mahoney · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
A. Fan, E. Grave, and A. Joulin · 2020
Later among the works it cites.
Towards understanding normalization in neural ODEs
J. Gusak, L. Markeeva, T. Daulbaev, A. Katrutsa, A. Cichocki, and I. Oseledets · 2020
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2020
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Y. Lu, A. Zhong, Q. Li, and B. Dong · 2018
Cited alongside, same era.
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Y. Lu, A. Zhong, Q. Li, and B. Dong · 2018
Cited alongside, same era.
Deep equilibrium models
S. Bai, J. Z. Kolter, and V. Koltun · 2019
Cited alongside, same era.
HDT-UD: A very large Universal Dependencies treebank for German
E. Borges Völker, M. Wendt, F. Hennig, and A. Köhn · 2019
Cited alongside, same era.
AntisymmetricRNN: A dynamical system view on recurrent neural networks
B. Chang, M. Chen, E. Haber, and E. H. Chi · 2019
Cited alongside, same era.
Augmented neural odes
E. Dupont, A. Doucet, and Y. W. Teh · 2019
Cited alongside, same era.
ANODE: unconditionally accurate memory-efficient gradients for neural ODEs
A. Gholami, K. Keutzer, and G. Biros · 2019
Cited alongside, same era.
Later among the works it cites.
Deep transformers with latent depth
X. Li, A. Cooper Stickland, Y. Tang, and X. Kong · 2020
Later among the works it cites.
Understanding recurrent neural networks using nonequilibrium response theory
S. H. Lim · 2020
Later among the works it cites.
Dissecting neural odes
S. Massaroli, M. Poli, J. Park, A. Yamashita, and H. Asma · 2020
Later among the works it cites.
Universal dependencies v2: An evergrowing multilingual treebank collection
J. Nivre, M.-C. de Marneffe, F. Ginter, J. Hajič, C. D. Manning, S. Pyysalo, S. Schuster, F. Tyers, and D. Zeman · 2020
Later among the works it cites.
Continuous-in-depth neural networks
A. F. Queiruga, N. B. Erichson, D. Taylor, and M. W. Mahoney · 2020
Later among the works it cites.
Adaptive checkpoint adjoint method for gradient estimation in neural ODE
J. Zhuang, N. Dvornek, X. Li, S. Tatikonda, X. Papademetris, and J. Duncan · 2020
Later among the works it cites.
Noisy recurrent neural networks
S. H. Lim, N. B. Erichson, L. Hodgkinson, and M. W. Mahoney · 2021
Closest in time.
Coupled oscillatory recurrent neural network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies
T. K. Rusch and S. Mishra · 2021
Closest in time.
Unicornn: A recurrent model for learning very long time dependencies
T. K. Rusch and S. Mishra · 2021
Closest in time.
Infinitely deep bayesian neural networks with stochastic differential equations
W. Xu, R. T. Chen, X. Li, and D. Duvenaud · 2021
Closest in time.
Neural delay differential equations
Q. Zhu, Y. Guo, and W. Lin · 2021
Closest in time.