Fetching the paper…
Reading the bibliography…
Neural Ordinary Differential Equations (Neural ODEs) are the continuous analog of Residual Neural Networks (ResNets).
Mathematical theory of optimal processes
Lev Semenovich Pontryagin · 1987
Earlier work this paper cites.
Functional analysis, Sobolev spaces and partial differential equations , volume 2
Haim Brezis and Haim Brézis · 2011
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Analyse numérique et équations différentielles-4ème Ed
Jean-Pierre Demailly · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large kernel matters–improve semantic segmentation by global convolutional network
Chao Peng, Xiangyu Zhang, Gang Yu, Guiming Luo, and Jian Sun · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Earlier work this paper cites.
The reversible residual network: Backpropagation without storing activations
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Earlier work this paper cites.
A proposal on machine learning via dynamical systems
E Weinan · 2017
Earlier work this paper cites.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Earlier work this paper cites.
Neural ordinary differential equations
Tian Qi Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Earlier work this paper cites.
Superneurons: Dynamic gpu memory management for training deep neural networks
Linnan Wang, Jinmian Ye, Yiyang Zhao, Wei Wu, Ang Li, Shuaiwen Leon Song, Zenglin Xu, and Tim Kraska · 2018
Earlier work this paper cites.
Stochastic training of residual networks: a differential equation viewpoint
Qi Sun, Yunzhe Tao, and Qiang Du · 2018
Earlier work this paper cites.
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Yiping Lu, Aoxiao Zhong, Quanzheng Li, and Bin Dong · 2018
Earlier work this paper cites.
Ffjord: Free-form continuous dynamics for scalable reversible generative models
Will Grathwohl, Ricky TQ Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud · 2018
Cited alongside, same era.
i-revnet: Deep invertible networks
Jörn-Henrik Jacobsen, Arnold W.M. Smeulders, and Edouard Oyallon · 2018
Cited alongside, same era.
Automatic differentiation in machine learning: a survey
Atilim Gunes Baydin, Barak A Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
Universal approximation property of neural ordinary differential equations
Takeshi Teshima, Koichi Tojo, Masahiro Ikeda, Isao Ishikawa, and Kenta Oono · 2020
Later among the works it cites.
Deep neural networks, generic universal interpolation, and controlled odes
Christa Cuchiero, Martin Larsson, and Josef Teichmann · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
A mean field analysis of deep resnet and beyond: Towards provably optimization via overparameterization from depth
Yiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu, and Lexing Ying · 2020
Later among the works it cites.
Adaptive checkpoint adjoint method for gradient estimation in neural ode
Juntang Zhuang, Nicha Dvornek, Xiaoxiao Li, Sekhar Tatikonda, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Augmented neural odes
Y Teh, Arnaud Doucet, and E Dupont · 2019
Cited alongside, same era.
Deep learning via dynamical systems: An approximation perspective
Qianxiao Li, Ting Lin, and Zuowei Shen · 2019
Cited alongside, same era.
A mean-field optimal control formulation of deep learning
E Weinan, Jiequn Han, and Qianxiao Li · 2019
Cited alongside, same era.
Deep neural networks motivated by partial differential equations
Lars Ruthotto and Eldad Haber · 2019
Cited alongside, same era.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2019
Cited alongside, same era.
Hamiltonian neural networks
Samuel Greydanus, Misko Dzamba, and Jason Yosinski · 2019
Cited alongside, same era.
Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho · 2019
Cited alongside, same era.
On the global convergence of training deep linear resnets
Difan Zou, Philip M Long, and Quanquan Gu · 2020
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Revisiting resnets: Improved training and scaling strategies
Irwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, and Barret Zoph · 2021
Later among the works it cites.
Mali: A memory efficient and reverse accurate integrator for neural odes
Juntang Zhuang, Nicha C Dvornek, Sekhar Tatikonda, and James S Duncan · 2021
Later among the works it cites.
Scaling properties of deep residual networks
Alain-Sam Cohen, Rama Cont, Alain Rossier, and Renyuan Xu · 2021
Later among the works it cites.
Global convergence of resnets: From finite to infinite width using linear parameterization
Raphaël Barboni, Gabriel Peyré, and François-Xavier Vialard · 2021
Later among the works it cites.
Heunnet: Extending resnet using heun’s method
Mehrdad Maleki, Mansura Habiba, and Barak A Pearlmutter · 2021
Later among the works it cites.
On neural differential equations
Patrick Kidger · 2022
Closest in time.
Sinkformers: Transformers with doubly stochastic attention
Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré · 2022
Closest in time.
Convergence and implicit regularization properties of gradient descent for deep residual networks
Rama Cont, Alain Rossier, and RenYuan Xu · 2022
Closest in time.