Fetching the paper…
Reading the bibliography…
The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks.
A universal prior for integers and estimation by minimum description length
Jorma Rissanen · 1983
Earlier work this paper cites.
Weak sharp minima and penalty functions in mathematical programming
Michael Charles Ferris · 1988
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
David JC MacKay · 1992
Earlier work this paper cites.
Timit acoustic-phonetic continuous speech corpus
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, David S Pallett, Nancy L Dahlgren, and Victor Zue · 1993
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Robust optimization for unconstrained simulation-based problems
Dimitris Bertsimas, Omid Nohadani, and Kwong Meng Teo · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Improving the robustness of deep neural networks via stability training
Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Richard H Byrd, Gillian M Chin, Jorge Nocedal, and Yuchen Wu · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data fitting
Michael P Friedlander and Mark Schmidt · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Later among the works it cites.
Uri Shaham, Yutaro Yamada, and Sahand Negahban · 2015
Later among the works it cites.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Later among the works it cites.
Deep learning
Yoshua Bengio, Ian Goodfellow, and Aaron Courville · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Closest in time.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2016
Closest in time.
Distributed deep learning using synchronous stochastic gradient descent
Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere, Karthikeyan Vaidynathan, Srinivas Sridharan, Dhiraj Kalamkar, Bharat Kaul, and Pradeep Dubey · 2016
Closest in time.
adaQN: An Adaptive Quasi-Newton Algorithm for Training RNNs , pp. 1–16
Nitish Shirish Keskar and Albert S. Berahas · 2016
Closest in time.
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Closest in time.
Training recurrent neural networks by diffusion
Hossein Mobahi · 2016
Closest in time.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Closest in time.