Fetching the paper…
Reading the bibliography…
Restart techniques are common in gradient-free optimization to deal with multimodal functions.
Function minimization by conjugate gradients
Reeves Fletcher and Colin M Reeves · 1964
Earlier work this paper cites.
Restart procedures for the conjugate gradient method
Michael James David Powell · 1977
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Kenji Fukumizu and Shun-ichi Amari · 2000
Earlier work this paper cites.
Evaluating the cma evolution strategy on multimodal test functions
Nikolaus Hansen and Stefan Kern · 2004
Earlier work this paper cites.
Sgd-qn: Careful quasi-newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Earlier work this paper cites.
Benchmarking a BI-population CMA-ES on the BBOB-2009 function testbed
Nikolaus Hansen · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Benchmarking the bfgs algorithm on the bbob-2009 function testbed
Raymond Ros · 2009
Earlier work this paper cites.
Niching the CMA-ES via nearest-better clustering
Mike Preuss · 2010
Earlier work this paper cites.
Alternative restart strategies for CMA-ES
Ilya Loshchilov, Marc Schoenauer, and Michele Sebag · 2012
Earlier work this paper cites.
Adaptive restart for accelerated gradient schemes
Brendan O’Donoghue and Emmanuel Candes · 2012
Cited alongside, same era.
Adadelta: An adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
New types of deep neural network learning for speech recognition and related applications: An overview
L. Deng, G. Hinton, and B. Kingsbury · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
The loss surface of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2014
Cited alongside, same era.
Niching methods and multimodal optimization performance
Mike Preuss · 2015
Later among the works it cites.
No more pesky learning rate guessing games
Leslie N Smith · 2015
Later among the works it cites.
Stochastic subgradient methods with linear convergence for polyhedral convex optimization
Tianbao Yang and Qihang Lin · 2015
Later among the works it cites.
Deep pyramidal residual networks
Dongyoon Han, Jiwhan Kim, and Junmo Kim · 2016
Closest in time.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Rmsprop and equilibrated adaptive learning rates for non-convex optimization
Yann N Dauphin, Harm de Vries, Junyoung Chung, and Yoshua Bengio · 2015
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Tiny imagenet visual recognition challenge
Hadi Pouransari and Saman Ghili · 2015
Cited alongside, same era.
SGDR: Stochastic Gradient Descent with Restarts
Ilya Loshchilov and Frank Hutter · 2016
Closest in time.
Cyclical learning rates for training neural networks
Leslie N Smith · 2016
Closest in time.
Sergey Zagoruyko and Nikos Komodakis · 2016
Closest in time.
Residual Networks of Residual Networks: Multilevel Residual Networks
K. Zhang, M. Sun, T. X. Han, X. Yuan, L. Guo, and T. Liu · 2016
Closest in time.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger · 2017
Closest in time.
Robin Tibor Schirrmeister, Jost Tobias Springenberg, Lukas Dominique Josef Fiederer, Martin Glasstetter, Katharina Eggensperger, Michael Tangermann, Frank Hutter, Wolfram Burgard, and Tonio Ball · 2017
Closest in time.