Fetching the paper…
Reading the bibliography…
Stagewise training strategy is widely used for learning neural networks, which runs a stochastic algorithm (e.g., SGD) starting with a relatively large step size (aka learning rate) and geometrically decreasing the step size after a number of iterations.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization : a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Beyond the regret minimization barrier: an optimal algorithm for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
From error bounds to the complexity of first-order descent methods for convex functions
Jerome Bolte, Trong Phong Nguyen, Juan Peypouquet, and Bruce Suter · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark W. Schmidt · 2016
Cited alongside, same era.
Fast rate analysis of some stochastic optimization algorithms
Chao Qu, Huan Xu, and Chong Ong · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
On exponential convergence of sgd in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Closest in time.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2018
Closest in time.
Universal stagewise learning for non-convex problems with convergence on averaged solutions
Zaiyi Chen, Tianbao Yang, Jinfeng Yi, Bowen Zhou, and Enhong Chen · 2018
Closest in time.
Damek Davis and Dmitriy Drusvyatskiy · 2018
Closest in time.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Cited alongside, same era.
Diversity leads to generalization in neural networks
Bo Xie, Yingyu Liang, and Le Song · 2016
Cited alongside, same era.
Non-convex finite-sum optimization via SCSG methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I. Jordan · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Stochastic convex optimization: Faster local growth implies faster global convergence
Yi Xu, Qihang Lin, and Tianbao Yang · 2017
Cited alongside, same era.
Characterization of gradient dominance and regularity conditions for neural networks
Yi Zhou and Yingbin Liang · 2017
Cited alongside, same era.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph H. Lampert · 2018
Closest in time.
Accelerated stochastic algorithms for nonconvex finite-sum and multi-block optimization
Guanghui Lan and Yu Yang · 2018
Closest in time.
Yi Xu, Qi Qi, Qihang Lin, Rong Jin, and Tianbao Yang · 2018
Closest in time.
Generalization error bounds with probabilistic guarantee for SGD in nonconvex optimization
Yi Zhou, Yingbin Liang, and Huishuai Zhang · 2018
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
Rethinking learning rate schedules for stochastic optimization, 2019
Rong Ge, Sham M. Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Closest in time.