Fetching the paper…
Reading the bibliography…
We study the convergence of gradient flow for the training of deep neural networks.
“Ordinary differential equations”
Jack Hale · 1969
Earlier work this paper cites.
“Approximation by superpositions of a sigmoidal function”
George Cybenko · 1989
Earlier work this paper cites.
“New problems on minimizing movements”
Ennio De · 1993
Earlier work this paper cites.
“Learning long-term dependencies with gradient descent is difficult”
Y. Bengio, P. Simard and P. Frasconi · 1994
Earlier work this paper cites.
“Error estimates and condition numbers for radial basis function interpolation”
Robert Schaback · 1995
Earlier work this paper cites.
“Scattered data approximation”
Holger Wendland · 2004
Earlier work this paper cites.
“Measure Theory”
Vladimir Bogachev · 2007
Earlier work this paper cites.
“Gradient flows: in metric spaces and in the space of probability measures”
Luigi Ambrosio, Nicola Gigli and Giuseppe Savaré · 2008
Earlier work this paper cites.
“Transport equation and Cauchy problem for non-smooth vector fields”
Luigi Ambrosio · 2008
Earlier work this paper cites.
“Kernel methods for deep learning”
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
“Optimal transport: old and new”
C\’edric Villani · 2009
Earlier work this paper cites.
“Vector valued reproducing kernel Hilbert spaces and universality”
Claudio Carmeli, Ernesto De, Alessandro Toigo and Veronica Umanit\’a · 2010
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks”
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
“Universality, Characteristic Kernels and RKHS Embedding of Measures.”
Bharath Sriperumbudur, Kenji Fukumizu and Gert Lanckriet · 2011
Earlier work this paper cites.
“Deep learning made easier by linear transformations in perceptrons”
Tapani Raiko, Harri Valpola and Yann LeCun · 2012
Earlier work this paper cites.
“A user’s guide to optimal transport”
Luigi Ambrosio et al · 2013
Earlier work this paper cites.
“Variational analysis in Sobolev and BV spaces: applications to PDEs and optimization”
Hedy Attouch, Giuseppe Buttazzo and G\’erard Michaille · 2014
Earlier work this paper cites.
“Introduction to measure theory and functional analysis”
Piermarco Cannarsa and Teresa D’Aprile · 2015
Earlier work this paper cites.
“Optimal transport for applied mathematicians”
Filippo Santambrogio · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Identity mappings in deep residual networks”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Identity Matters in Deep Learning”
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
“On the equivalence between kernel quadrature rules and random feature expansions”
Francis Bach · 2017
Earlier work this paper cites.
“Convergence Analysis of Two-layer Neural Networks with ReLU Activation”
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
“ { \{ Euclidean, metric, and Wasserstein } \} gradient flows: an overview”
Filippo Santambrogio · 2017
Cited alongside, same era.
“Inception-v4, inception-resnet and the impact of residual connections on learning”
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke and Alexander Alemi · 2017
Cited alongside, same era.
“Attention is all you need”
Ashish Vaswani et al · 2017
Cited alongside, same era.
“Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks”
Peter Bartlett, Dave Helmbold and Philip Long · 2018
Cited alongside, same era.
“On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport”
L\’enaïc Chizat and Francis Bach · 2018
Cited alongside, same era.
“Neural Ordinary Differential Equations”
“A shooting formulation of deep learning”
Francois-Xavier Vialard, Roland Kwitt, Susan Wei and Marc Niethammer · 2020
Later among the works it cites.
Stephan Wojtowytsch · 2020
Later among the works it cites.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2020
Later among the works it cites.
“Nonsmooth implicit differentiation for machine-learning and optimization”
J\’er\ˆome Bolte, Tam Le, Edouard Pauwels and Tony Silveti-Falls · 2021
Later among the works it cites.
“On the global convergence of gradient descent for multi-layer resnets in the mean-field regime”
Zhiyan Ding, Shi Chen, Qin Li and Stephen Wright · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ricky T.. Chen, Yulia Rubanova, Jesse Bettencourt and David Duvenaud · 2018
Cited alongside, same era.
“A mean field view of the landscape of two-layer neural networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Cited alongside, same era.
“Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks”
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Cited alongside, same era.
“Fixup Initialization: Residual Learning Without Normalization”
Hongyi Zhang, Yann Dauphin and Tengyu Ma · 2018
Cited alongside, same era.
“A convergence theory for deep learning via over-parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Cited alongside, same era.
“On Lazy Training in Differentiable Programming”
Lenaic Chizat, Edouard Oyallon and Francis Bach · 2019
Cited alongside, same era.
“Gradient descent finds global minima of deep neural networks”
Simon Du et al · 2019
Cited alongside, same era.
“The Barron Space and the Flow-Induced Function Spaces for Neural Network Models”
Weinan E, Chao Ma and Lei Wu · 2021
Later among the works it cites.
“On the proof of global convergence of gradient descent for deep relu networks with linear widths”
Quynh Nguyen · 2021
Later among the works it cites.
“Momentum residual neural networks”
Michael Sander, Pierre Ablin, Mathieu Blondel and Gabriel Peyr\’e · 2021
Later among the works it cites.
“Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers”
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege and Michael Westdickenberg · 2022
Later among the works it cites.
“On global convergence of ResNets: From finite to infinite width using linear parameterization”
Rapha\"el Barboni, Gabriel Peyr\’e and Francois-Xavier Vialard · 2022
Later among the works it cites.
“Convergence of gradient descent for deep neural networks”
Sourav Chatterjee · 2022
Later among the works it cites.
“Overparameterization of deep ResNet: zero loss and mean-field analysis”
Zhiyan Ding, Shi Chen, Qin Li and Stephen Wright · 2022
Later among the works it cites.
“Representation formulas and pointwise properties for Barron functions”
E Weinan and Stephan Wojtowytsch · 2022
Later among the works it cites.
“Generalization Guarantees of Deep ResNets in the Mean-Field Regime”
Yihang Chen et al · 2023
Later among the works it cites.
“Conditional Optimal Transport on Function Spaces”
Bamdad Hosseini, Alexander Hsu and Amirhossein Taghvaei · 2023
Later among the works it cites.
“A Convergence result of a continuous model of deep learning via Łojasiewicz–Simon inequality”
Noboru Isobe · 2023
Later among the works it cites.
“Implicit regularization of deep residual networks towards neural ODEs”, 2023
Pierre Marion, Yu-Han Wu, Michael Sander and G\’erard Biau · 2023
Later among the works it cites.
“A rigorous framework for the mean field limit of multilayer neural networks”
Phan-Minh Nguyen and Huy Pham · 2023
Later among the works it cites.
“Heterogeneous gradient flows in the topology of fibered optimal transport”
Jan Peszek and David Poyato · 2023
Later among the works it cites.
Lorenzo Schiavo, Jan Maas and Francesco Pedrotti · 2023
Later among the works it cites.
“Deep limits of residual neural networks”
Matthew Thorpe and Yves van Gennip · 2023
Later among the works it cites.
“Conditional Wasserstein Distances with Applications in Bayesian OT Flow Matching”
Jannis Chemseddine, Paul Hagemann, Christian Wald and Gabriele Steidl · 2024
Closest in time.
“Dynamic Conditional Optimal Transport through Simulation-Free Flows”
Gavin Kerrigan, Giosue Migliorini and Padhraic Smyth · 2024
Closest in time.