Fetching the paper…
Reading the bibliography…
This paper presents a partial differential equation framework for deep residual neural networks and for the associated learning problem.
A stochastic approximation method
H. Robinds and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Applied Optimal Control: Optimization, Estimation and Control
A.E. Bryson · 1975
Earlier work this paper cites.
Deterministic and Stochastic Control
W. H. Fleming and R. W. Rishel · 1975
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R. Rockafellar · 1976
Earlier work this paper cites.
Method of successive approximations for solution of optimal control problems
F. L. Chernousko and A. A. Lyubushin · 1982
Earlier work this paper cites.
A learning rule for asynchronous perceptrons with feedback in a combinatorial environment
L.B. Almeida · 1987
Earlier work this paper cites.
Generalization of back propagation to recurrent and higher order neural networks
F. J. Pineda · 1987
Earlier work this paper cites.
Mathematical Theory of Optimal Processes
L. S. Pontryagin · 1987
Earlier work this paper cites.
A theoretical framework for back-propagation
LeCun, Y · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K. Hornik · 1989
Earlier work this paper cites.
Semigroups of Linear Operators and Applications to Partial Differential Equations
A. Pazy · 1992
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. R. Barron · 1993
Earlier work this paper cites.
Optimal Control and Viscosity Solutions of Hamilton-Jacobi-Bellman Equations
M. Bardi and I. Capuzzo-Dolcetta · 1997
Earlier work this paper cites.
Optimal Control and Viscosity Solutions of Hamilton-Jacobi Equations
M. Bardi and D. I. Capuzzo · 1997
Earlier work this paper cites.
Are loss functions all the same?
L. Rosasco, E. D. De Vito, A. Caponnetto, M. Piana, A. Verri · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
Learning deep architectures for AI
Y. Bengio · 2009
Earlier work this paper cites.
A survey of numerical methods for optimal control
A. V Rao · 2009
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: an approach based on the Kurdyka-Lojasiewicz inequality
H. Attouch, J. Bolte, P. Redont and A. Soubeyran · 2010
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Deep sparse rectified neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Optimal Control: An Introduction to the Theory and Its Applications
M. Athans and P.L. Falb · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2013
Cited alongside, same era.
Nonlinear Systems
H. K. Khalil · 2014
Cited alongside, same era.
Deep residual learning and PDEs on manifold
Li, Z., Shi, Z · 2017
Later among the works it cites.
Y. Lu, A. Zhong, Q. Li and B. Dong · 2017
Later among the works it cites.
Searching for activation functions
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Later among the works it cites.
Double continuum limit of deep neural networks
S. Sonoda, N. Murata · 2017
Later among the works it cites.
PolyNet: A pursuit of structural diversity in very deep networks
X. Zhang, Z. Li, C. Loy · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the design of loss functions for classification: theory, robustness to outliers, and savageBoost,
H. Masnadi-Shirazi and N. Vasconcelos · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
A differential equation for modeling Nesterov’s accelerated gradient method: theory and insights
W. Su, S. Boyd, and E. J. Candes · 2015
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2016
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
S. Arora, N. Cohen, N. Golowich, and W. Hu · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Optimal approximation with sparsely connected deep neural networks
H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen · 2018
Later among the works it cites.
Reversible architectures for arbitrarily deep residual neural networks
B. Chang, L. Meng, E. Haber, L. Ruthotto, D. Begert, and E. Holtham · 2018
Later among the works it cites.
Multi-level residual networks from dynamical systems view
B. Chang, L. Meng, E. Haber, F. Tung and D. Begert · 2018
Later among the works it cites.
Exponential convergence of the deep neural network approximation for analytic functions
W. E and Q. Wang · 2018
Later among the works it cites.
Maximum principle based algorithms for deep learning
Q. Li, L. Chen, C. Tai, and W. E · 2018
Later among the works it cites.
Beyond finite layer neural network: bridging deep architects and numerical differential equations
Y. Lu, A. Zhong, Q. Li, and B. Dong · 2018
Later among the works it cites.
An optimal control approach to deep learning and applications to discrete-weight neural networks
Q. Li and S. Hao · 2018
Later among the works it cites.
Deep limits of residual neural networks
T. Matthew, Y. van Gennip · 2018
Later among the works it cites.
Deep neural networks motivated by partial differential equations
L. Ruthotto and E. Haber · 2018
Later among the works it cites.
Spurious local minima are common in two-layer Relu neural networks
I. Safran and O. Shamir · 2018
Later among the works it cites.
Nonlocal neural networks, nonlocal diffusion and nonlocal modeling
Y. Tao, Q. Sun, Q. Du, and W. Liu · 2018
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Later among the works it cites.
Stochastic backward Euler: an implicit gradient descent algorithm for k k -means clustering
P. Yin, M. Pham, A. Oberman, and S. Osher · 2018
Later among the works it cites.
A mean-field optimal control formulation of deep learning
W. E, J. Han and Q. Li · 2019
Closest in time.
Deep neural network approximation theory
P. Grohs, D. Perekrestenko, D. Elbrächter, and H. Bölcskei · 2019
Closest in time.