Fetching the paper…
Reading the bibliography…
Deep neural networks have become invaluable tools for supervised machine learning, e.g., classification of text or images.
The use of adjoint systems in the problem of differential corrections for trajectories
G. A. Bliss · 1919
Earlier work this paper cites.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Mehtods of conjugate gradients for solving linear systems
M. Hestenes and E. Stieffel · 1952
Earlier work this paper cites.
Contributions to the theory of optimal control
R. E. Kalman et al · 1960
Earlier work this paper cites.
Learning representations by back-propagating errors
D. Rumelhart, G. Hinton, and J. Williams, R · 1986
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. E. Boser, and J. S. Denker · 1990
Earlier work this paper cites.
A scaled conjugate gradient algorithm for fast supervised learning
M. F. Møller · 1993
Earlier work this paper cites.
Learning Long-Term Dependencies with Gradient Descent Is Difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Numerical Solution of Boundary Value Problems for Ordinary Differential Equations
U. Ascher, R. Mattheij, and R. Russell · 1995
Earlier work this paper cites.
The cascadic multigrid method for elliptic problems
F. A. Bornemann and P. Deuflhard · 1996
Earlier work this paper cites.
Regularization of Inverse Problems
H. Engl, M. Hanke, and A. Neubauer · 1996
Earlier work this paper cites.
Rank-Deficient and Discrete Ill-Posed Problems
P. C. Hansen · 1997
Earlier work this paper cites.
Computer Methods for Ordinary Differential Equations and Differential-Algebraic Equations
U. Ascher and L. Petzold · 1998
Earlier work this paper cites.
Partial Differential Equations
L. C. Evans · 1998
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S. Wright · 1999
Earlier work this paper cites.
A GCV based methods for nonlinear inverse problems
E. Haber and D. Oldenburg · 2000
Earlier work this paper cites.
The elements of statistical learning
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Cited alongside, same era.
Computational methods for inverse problems
C. R. Vogel · 2001
Cited alongside, same era.
Adjoint Sensitity Analysis for Differential-Aglebraic Equations: The Adjoint DAE System and its Numerical Solution
Y. Cao, S. Li, L. Petzold, and R. Serban · 2003
Cited alongside, same era.
Iterative Methods for Sparse Linear Systems
Y. Saad · 2003
Cited alongside, same era.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
A. Y. Ng · 2004
Cited alongside, same era.
Priorconditioners for linear systems
D. Calvetti and E. Somersalo · 2005
Cited alongside, same era.
Stochastic gradient descent tricks
L. Bottou · 2012
Later among the works it cites.
Sample size selection in optimization methods for machine learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Later among the works it cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Later among the works it cites.
Training Deep and Recurrent Networks with Hessian-Free Optimization
J. Martens and I. Sutskever · 2012
Later among the works it cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nonlinear inverse scale space methods
M. Burger, G. Gilboa, S. Osher, and J. Xu · 2006
Cited alongside, same era.
Learning deep architectures for AI
Y. Bengio · 2009
Cited alongside, same era.
Measuring invariances in deep networks
I. Goodfellow, H. Lee, Q. V. Le, A. Saxe, and A. Y. Ng · 2009
Cited alongside, same era.
FAIR: Flexible Algorithms for Image Registration
J. Modersitzki · 2009
Cited alongside, same era.
Numerical methods for Evolutionary Differential Equations
U. Ascher · 2010
Cited alongside, same era.
Image Processing and Analysis
T. F. Chan and J. Shen · 2010
Cited alongside, same era.
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Later among the works it cites.
K. Rothauge, E. Haber, and U. Ascher · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Sub-Sampled Newton Methods II: Local Convergence Rates
F. Roosta-Khorasani and M. W. Mahoney · 2016
Later among the works it cites.
Learning across scales-a multiscale method for convolution neural networks
E. Haber, L. Ruthotto, and E. Holtham · 2017
Closest in time.