Fetching the paper…
Reading the bibliography…
Deep learning is formulated as a discrete-time optimal control problem.
Normality and abnormality in the calculus of variations
Bliss, G. A · 1938
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
The theory of optimal processes. I. The maximum principle
Boltyanskii, V. G., Gamkrelidze, R. V., and Pontryagin, L. S · 1960
Earlier work this paper cites.
On the method of successive approximations for solution of optimal control problems
Krylov, I. A. and Chernousko, F. L · 1962
Earlier work this paper cites.
Relaxed variational problems
Warga, J · 1962
Earlier work this paper cites.
A maximum principle of the pontryagin type for systems described by nonlinear difference equations
Halkin, H · 1966
Earlier work this paper cites.
Discretional convexity and the maximum principle for discrete systems
Holtzman, J. M. and Halkin, H · 1966
Earlier work this paper cites.
A discrete version of pontryagin’s maximum principle
Hwang, C. and Fan, L · 1967
Earlier work this paper cites.
On the accumulation of perturbations in the linear systems with two coordinates
Aleksandrov, V. V · 1968
Earlier work this paper cites.
Theory of optimal control and mathematical programming
Canon, M. D., Cullum Jr, C. D., and Polak, E · 1970
Earlier work this paper cites.
An algorithm for the method of successive approximations in optimal control problems
Krylov, I. A. and Chernousko, F. L · 1972
Earlier work this paper cites.
Applied optimal control: optimization, estimation and control
Bryson, A. E · 1975
Earlier work this paper cites.
Method of successive approximations for solution of optimal control problems
Chernousko, F. L. and Lyubushin, A. A · 1982
Earlier work this paper cites.
Modifications of the method of successive approximations for solving optimal control problems
Lyubushin, A. A · 1982
Earlier work this paper cites.
Mathematical theory of optimal processes
Pontryagin, L. S · 1987
Earlier work this paper cites.
A theoretical framework for back-propagation
LeCun, Y · 1988
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 1995
Earlier work this paper cites.
Discrete-time control systems , volume 2
Ogata, K · 1995
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Y · 1998
Cited alongside, same era.
Nonlinear programming
Bertsekas, D. P · 1999
Cited alongside, same era.
Large deviations , volume 14
Den Hollander, F · 2008
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, Eric and, F. R · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Later among the works it cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
Baldassi, C., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R · 2015
Later among the works it cites.
Automatic differentiation in machine learning: a survey
Baydin, A. G., Pearlmutter, B. A., Radul, A. A., and Siskind, J. M · 2015
Later among the works it cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P · 2015
Later among the works it cites.
Han, S., Mao, H., and Dally, W. J · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
Optimal control: an introduction to the theory and its applications
Athans, M. and Falb, P. L · 2013
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
Bach, F. and Moulines, E · 2013
Cited alongside, same era.
Nonlinear programming: theory and algorithms
Bazaraa, M. S., Sherali, H. D., and Shetty, C. M · 2013
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Cited alongside, same era.
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Later among the works it cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Later among the works it cites.
Li, F., Zhang, B., and Liu, B · 2016
Later among the works it cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A · 2016
Later among the works it cites.
Zhu, C., Han, S., Mao, H., and Dally, W. J · 2016
Later among the works it cites.
On the role of synaptic stochasticity in training low-precision neural networks
Baldassi, C., Gerace, F., Kappen, H. J., Lucibello, C., Saglietti, L., Tartaglione, E., and Zecchina, R · 2017
Later among the works it cites.
Reversible architectures for arbitrarily deep residual neural networks
Chang, B., Meng, L., Haber, E., Ruthotto, L., Begert, D., and Holtham, E · 2017
Later among the works it cites.
A proposal on machine learning via dynamical systems
E, W · 2017
Later among the works it cites.
Stable architectures for deep neural networks
Haber, E. and Ruthotto, L · 2017
Later among the works it cites.
How to train a compact binary neural network with high accuracy?
Tang, W., Hua, G., and Wang, L · 2017
Later among the works it cites.
Maximum principle based algorithms for deep learning
Li, Q., Chen, L., Tai, C., and E, W · 2018
Closest in time.
Binaryrelax: A relaxation approach for training deep neural networks with quantized weights
Yin, P., Zhang, S., Lyu, J., Osher, S., Qi, Y., and Xin, J · 2018
Closest in time.