Fetching the paper…
Reading the bibliography…
In this paper we prove that, in the deep limit, the stochastic gradient descent on a ResNet type deep neural network, where each layer shares the same weight matrix, converges to the stochastic gradient descent for a Neural ODE and that the corresponding value/loss functions converge.
A simple baseline for Bayesian uncertainty in deep learning
Maddox, W., Garipov, T., Izmailov, P., Vetrov, D., and Wilson, A. G · 1902
Earlier work this paper cites.
Dupont, E., Doucet, A., and Teh, Y. W · 1904
Earlier work this paper cites.
Universality of deep convolutional neural networks
Zhou, D.-X · 1906
Earlier work this paper cites.
Stochastic differential equations
Gīhman, u. I., and Skorohod, A. V · 1972
Earlier work this paper cites.
Logarithmic Sobolev inequalities
Gross, L · 1975
Earlier work this paper cites.
An introduction to the theory of large deviations
Stroock, D. W · 1984
Earlier work this paper cites.
Dirichlet forms
Fabes, E., Fukushima, M., Gross, L., Kenig, C., Röckner, M., and Stroock, D. W · 1992
Earlier work this paper cites.
An Introduction to Γ \Gamma -convergence
Dal Maso, G · 1993
Earlier work this paper cites.
Second order parabolic differential equations
Lieberman, G. M · 1996
Earlier work this paper cites.
An initiation to logarithmic Sobolev inequalities
Royer, G · 1999
Earlier work this paper cites.
Topics in optimal transportation
Villani, C · 2003
Earlier work this paper cites.
Existence and uniqueness of solutions to Fokker-Planck type equations with irregular coefficients
Le Bris, C., and Lions, P.-L · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., and Hinton, G · 2009
Earlier work this paper cites.
Optimal transport
Villani, C · 2009
Cited alongside, same era.
Partial differential equations
Evans, L. C · 2010
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X., and Bengio, Y · 2010
Cited alongside, same era.
On uniqueness problems related to the Fokker-Planck-Kolmogorov equation for measures
Bogachev, V., Röckner, M., and Shaposhnikov, S · 2011
Cited alongside, same era.
On uniqueness of a probability solution to the cauchy problem for the Fokker-Planck-Kolmogorov equation
Shaposhnikov, S · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S., and Lan, G · 2013
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M · 2017
Later among the works it cites.
Searching for activation functions, 2017
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Later among the works it cites.
Deep relaxation: partial differential equations for optimizing deep neural networks
Chaudhari, P., Oberman, A., Osher, S., Soatto, S., and Carlier, G · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P., and Soatto, S · 2018
Later among the works it cites.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Later among the works it cites.
On the power of over-parametrization in neural networks with quadratic activation
Du, S., and Lee, J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Graham, B · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., and Szegedy, C · 2015
Cited alongside, same era.
Deeply-supervised nets
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., and Tu, Z · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Later among the works it cites.
A mean-field optimal control formulation of deep learning
E, W., Han, J., and Li, Q · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Later among the works it cites.
Maximum principle based algorithms for deep learning
Li, Q., Chen, L., Tai, C., and Weinan, E · 2018
Later among the works it cites.
Li, Q., Tai, C., and E, W · 2018
Later among the works it cites.
Spurious local minima are common in two-layer ReLU neural networks
Safran, I., and Shamir, O · 2018
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J · 2019
Closest in time.