Fetching the paper…
Reading the bibliography…
Deep neural networks are commonly trained using stochastic non-convex optimization procedures, which are driven by gradient information estimated on fractions (batches) of the dataset.
A new approach to variable metric algorithms
Fletcher, Roger · 1970
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Liu, Dong C and Nocedal, Jorge · 1989
Earlier work this paper cites.
Evolutionary algorithms in theory and practice: evolution strategies, evolutionary programming, genetic algorithms
Bäck, Thomas · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, Kenji and Amari, Shun-ichi · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, Nikolaus and Ostermeier, Andreas · 2001
Earlier work this paper cites.
Max-margin markov networks
Roller, Ben Taskar Carlos Guestrin Daphne · 2004
Earlier work this paper cites.
Sequential parameter optimization
Bartz-Beielstein, Thomas, Lasarczyk, Christian WG, and Preuß, Mike · 2005
Earlier work this paper cites.
To recognize shapes, first learn to generate images
Hinton, Geoffrey E · 2007
Earlier work this paper cites.
Curriculum learning
Bengio, Yoshua, Louradour, Jérôme, Collobert, Ronan, and Weston, Jason · 2009
Earlier work this paper cites.
Sgd-qn: Careful quasi-newton stochastic gradient descent
Bordes, Antoine, Bottou, Léon, and Gallinari, Patrick · 2009
Earlier work this paper cites.
Benchmarking a weighted negative covariance matrix update on the bbob-2010 noiseless testbed
Hansen, Nikolaus and Ros, Raymond · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H., and Leyton-Brown, K · 2011
Cited alongside, same era.
Hybrid deterministic-stochastic methods for data fitting
Friedlander, Michael P and Schmidt, Mark · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G · 2012
Cited alongside, same era.
Efficiency of coordinate descent methods on huge-scale optimization problems
Nesterov, Yu · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Later among the works it cites.
Maximum likelihood-based online adaptation of hyper-parameters in CMA-ES
Loshchilov, Ilya, Schoenauer, Marc, Sebag, Michele, and Hansen, Nikolaus · 2014
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Needell, Deanna, Ward, Rachel, and Srebro, Nati · 2014
Later among the works it cites.
Stochastic optimization with importance sampling
Zhao, Peilin and Zhang, Tong · 2014
Later among the works it cites.
Variance reduction in sgd by distributed importance sampling
Alain, Guillaume, Lamb, Alex, Sankar, Chinnadhurai, Courville, Aaron, and Bengio, Yoshua · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schaul, Tom, Zhang, Sixin, and LeCun, Yann · 2012
Cited alongside, same era.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
Zeiler, Matthew D · 2012
Cited alongside, same era.
New types of deep neural network learning for speech recognition and related applications: An overview
Deng, L., Hinton, G., and Kingsbury, B · 2013
Cited alongside, same era.
The loss surface of multilayer networks
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2014
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Yann N, Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun, Ganguli, Surya, and Bengio, Yoshua · 2014
Cited alongside, same era.
Closest in time.
Stop wasting my gradients: Practical svrg
Babanezhad, Reza, Ahmed, Mohamed Osama, Virani, Alim, Schmidt, Mark, Konečnỳ, Jakub, and Sallinen, Scott · 2015
Closest in time.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, Matthieu, Bengio, Yoshua, and David, Jean-Pierre · 2015
Closest in time.
Rmsprop and equilibrated adaptive learning rates for non-convex optimization
Dauphin, Yann N, de Vries, Harm, Chung, Junyoung, and Bengio, Yoshua · 2015
Closest in time.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, Elad, Levy, Kfir Y, and Shalev-Shwartz, Shai · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Closest in time.
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2015
Closest in time.
Non-uniform stochastic average gradient method for training conditional random fields
Schmidt, Mark, Babanezhad, Reza, Ahmed, Mohamed Osama, Defazio, Aaron, Clifton, Ann, and Sarkar, Anoop · 2015
Closest in time.
” oddball sgd”: Novelty driven stochastic gradient descent for training deep neural networks
Simpson, Andrew JR · 2015
Closest in time.