Fetching the paper…
Reading the bibliography…
The continuous dynamical system approach to deep learning is explored in order to devise alternative frameworks for training algorithms.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
The maximum principle of L.S. Pontryagin in optimal-system theory
Lev I Rozonoer · 1959
Earlier work this paper cites.
The theory of optimal processes. i. the maximum principle
Vladimir Grigor’evich Boltyanskii, Revaz Valer’yanovich Gamkrelidze, and Lev Semenovich Pontryagin · 1960
Earlier work this paper cites.
Gradient theory of optimal flight paths
Henry J Kelley · 1960
Earlier work this paper cites.
On the method of successive approximations for solution of optimal control problems
Ivan A Krylov and Felix L Chernousko · 1962
Earlier work this paper cites.
Necessary and sufficient optimality conditions for sampled-data control systems
Anatolii B Butkovsky · 1963
Earlier work this paper cites.
On discrete analogues of pontryagin’s maximum principle
R Jackson and F Horn · 1965
Earlier work this paper cites.
A maximum principle of the pontryagin type for systems described by nonlinear difference equations
Hubert Halkin · 1966
Earlier work this paper cites.
On the accumulation of perturbations in the linear systems with two coordinates
Vladimir V Aleksandrov · 1968
Earlier work this paper cites.
Multiplier and gradient methods
Magnus R Hestenes · 1969
Earlier work this paper cites.
Asymptotic properties of non-linear least squares estimators
Robert I Jennrich · 1969
Earlier work this paper cites.
Two-point boundary value problems: shooting methods
Sanford M Roberts and Jerome S Shipman · 1972
Earlier work this paper cites.
Applied optimal control: optimization, estimation and control
Arthur Earl Bryson · 1975
Earlier work this paper cites.
Method of successive approximations for solution of optimal control problems
Felix L Chernousko and Alexey A Lyubushin · 1982
Earlier work this paper cites.
Modifications of the method of successive approximations for solving optimal control problems
Alexey A Lyubushin · 1982
Earlier work this paper cites.
The discrete-time maximum principle: a survey and some new results
Zbigniew Nahorski, Hans F Ravn, and René Victor Valqui Vidal · 1984
Earlier work this paper cites.
Mathematical theory of optimal processes
Lev S Pontryagin · 1987
Earlier work this paper cites.
A theoretical framework for back-propagation
Yann LeCun · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun and Yoshua Bengio · 1995
Earlier work this paper cites.
Fokker-planck equation
Hannes Risken · 1996
Cited alongside, same era.
Survey of numerical methods for trajectory optimization
John T Betts · 1998
Cited alongside, same era.
The MNIST database of handwritten digits
Yann LeCun · 1998
Cited alongside, same era.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Cited alongside, same era.
Nonlinear programming
Dimitri P Bertsekas · 1999
Cited alongside, same era.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard HR Hahnloser, Rahul Sarpeshkar, Misha A Mahowald, Rodney J Douglas, and H Sebastian Seung · 2000
Cited alongside, same era.
Elementary principles in statistical mechanics
J Willard Gibbs · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Later among the works it cites.
Automatic differentiation in machine learning: a survey
Atilim G Baydin, Barak A Pearlmutter, Alexey A Radul, and Jeffrey M Siskind · 2015
Later among the works it cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Deep learning in neural networks: An overview
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Francis Clarke · 2005
Cited alongside, same era.
Introduction to mathematical control theory
Alberto Bressan and Benedetto Piccoli · 2007
Cited alongside, same era.
Modern heuristic optimization techniques: theory and applications to power systems , volume 39
Kwang Y Lee and Mohamed A El-Sharkawi · 2008
Cited alongside, same era.
Learning deep architectures for AI
Yoshua Bengio · 2009
Cited alongside, same era.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
A survey of numerical methods for optimal control
Anil V Rao · 2009
Cited alongside, same era.
Jürgen Schmidhuber · 2015
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Later among the works it cites.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Later among the works it cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Decoupled neural interfaces using synthetic gradients
Max Jaderberg, Wojciech M Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Optimal control of continuity equations
Nikolay Pogodaev · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable ADMM approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein · 2016
Later among the works it cites.
Reversible architectures for arbitrarily deep residual neural networks
Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, and Elliot Holtham · 2017
Closest in time.
Understanding synthetic gradients and decoupled neural interfaces
Wojciech M Czarnecki, Grzegorz Świrszcz, Max Jaderberg, Simon Osindero, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Closest in time.
A proposal on machine learning via dynamical systems
Weinan E · 2017
Closest in time.
Thomas Frerix, Thomas Möllenhoff, Michael Moeller, and Daniel Cremers · 2017
Closest in time.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Closest in time.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Closest in time.
Numerical investigation of a class of liouville control problems
Souvik Roy and Alfio Borzì · 2017
Closest in time.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Closest in time.