Fetching the paper…
Reading the bibliography…
The learning rate (LR) is one of the most important hyper-parameters in stochastic gradient descent (SGD) algorithm for training deep neural networks (DNN).
L. B. Ward, “Reminiscence and rote learning.” Psychological Monographs , vol. 49, no. 4, 1937
1937
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
B. T. Polyak, “Gradient methods for minimizing functionals,” Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki , vol. 3, no. 4, pp. 643–653, 1963
1963
Earlier work this paper cites.
S. Lojasiewicz, “A topological property of real analytic subsets,” Coll. du CNRS, Les équations aux dérivées partielles , vol. 117, pp. 87–89, 1963
1963
Earlier work this paper cites.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
Y. Bengio, S. Bengio, and J. Cloutier, “Learning a synaptic learning rule,” in IJCNN , vol. 2. IEEE, 1991, pp. 969–vol
1991
Earlier work this paper cites.
J. Schmidhuber, “Learning to control fast-weight memories: An alternative to dynamic recurrent networks,” Neural Computation , vol. 4, no. 1, pp. 131–139, 1992
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Flat minima,” Neural Computation , vol. 9, no. 1, pp. 1–42, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
P. Tseng, “An incremental gradient (-projection) method with momentum term and adaptive stepsize rule,” SIAM Journal on Optimization , vol. 8, no. 2, pp. 506–531, 1998
1998
Earlier work this paper cites.
S. Hochreiter, A. S. Younger, and P. R. Conwell, “Learning to learn using gradient descent,” in International Conference on Artificial Neural Networks . Springer, 2001, pp. 87–94
2001
Earlier work this paper cites.
J. Nocedal and S. Wright, Numerical optimization . Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , vol. 22, no. 10, pp. 1345–1359, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of machine learning research , vol. 12, no. Jul, pp. 2121–2159, 2011
2011
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in NeurIPS Workshop on Deep Learning and Unsupervised Feature Learning , 2011
2011
Earlier work this paper cites.
M. D. Zeiler, “Adadelta: an adaptive learning rate method,” arXiv:1212.5701 , 2012
2012
Earlier work this paper cites.
T. Tieleman and G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” Neural networks for machine learning , 2012
2012
Earlier work this paper cites.
Y. Bengio, “Practical recommendations for gradient-based training of deep architectures,” in Neural networks: Tricks of the trade . Springer, 2012, pp. 437–478
2012
Earlier work this paper cites.
J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” JMLR , 2012
2012
Earlier work this paper cites.
J. Snoek, H. Larochelle, and R. P. Adams, “Practical bayesian optimization of machine learning algorithms,” in NeurIPS , 2012
2012
Earlier work this paper cites.
T. Schaul, S. Zhang, and Y. LeCun, “No more pesky learning rates,” in ICML , 2013
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas, “Learning to learn by gradient descent by gradient descent,” in NeurIPS , 2016
2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” in BMVC , 2016
2016
Earlier work this paper cites.
H. Karimi, J. Nutini, and M. Schmidt, “Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2016, pp. 795–811
2016
Cited alongside, same era.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic variance reduction for nonconvex optimization,” in ICML , 2016
2016
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” in ICLR , 2017
2017
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang, “On large-batch training for deep learning: Generalization gap and sharp minima,” in ICLR , 2017
2017
Cited alongside, same era.
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio, “Sharp minima can generalize for deep nets,” in ICML , 2017
Y. Wu, M. Ren, R. Liao, and R. Grosse, “Understanding short-horizon bias in stochastic meta-optimization,” in ICLR , 2018
2018
Later among the works it cites.
J. Shu, Z. Xu, and D. Meng, “Small sample learning in big data era,” arXiv:1808.04572 , 2018
2018
Later among the works it cites.
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in ECCV , 2018
2018
Later among the works it cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR , 2018
2018
Later among the works it cites.
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le, “Learning transferable architectures for scalable image recognition,” in CVPR , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts,” in ICLR , 2017
2017
Cited alongside, same era.
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building machines that learn and think like people,” Behavioral and brain sciences , vol. 40, 2017
2017
Cited alongside, same era.
S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” in ICLR , 2017
2017
Cited alongside, same era.
Y. Chen, M. W. Hoffman, S. G. Colmenarejo, M. Denil, T. P. Lillicrap, M. Botvinick, and N. De Freitas, “Learning to learn without gradient descent by gradient descent,” in ICML , 2017
2017
Cited alongside, same era.
O. Wichrowska, N. Maheswaranathan, M. W. Hoffman, S. G. Colmenarejo, M. Denil, N. de Freitas, and J. Sohl-Dickstein, “Learned optimizers that scale and generalize,” in ICML , 2017
2017
Cited alongside, same era.
K. Li and J. Malik, “Learning to optimize neural nets,” in ICLR , 2017
2017
Cited alongside, same era.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” Siam Review , vol. 60, no. 2, pp. 223–311, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
F. He, T. Liu, and D. Tao, “Control batch size and learning rate to generalize well: Theoretical and empirical evidence,” in NeurIPS , 2019
2019
Later among the works it cites.
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik, “Sgd: General analysis and improved rates,” in ICML , 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” NeurIPS , vol. 32, pp. 8026–8037, 2019
2019
Later among the works it cites.
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli, “The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares,” in NeurIPS , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J. D. Lee, I. Panageas, G. Piliouras, M. Simchowitz, M. I. Jordan, and B. Recht, “First-order methods almost always avoid saddle points,” Mathematical Programming , 2019
2019
Later among the works it cites.
I. Panageas, G. Piliouras, and X. Wang, “First-order methods almost always avoid saddle points: The case of vanishing step-sizes,” in NeurIPS , 2019
2019
Later among the works it cites.
L. Berrada, A. Zisserman, and M. P. Kumar, “Deep frank-wolfe for neural network optimization,” in ICLR , 2019
2019
Later among the works it cites.
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien, “Painless stochastic gradient: Interpolation, line-search, and convergence rates,” in NeurIPS , 2019
2019
Later among the works it cites.
E. Park and J. B. Oliva, “Meta-curvature,” in NeurIPS , 2019
2019
Later among the works it cites.
F. Hutter, L. Kotthoff, and J. Vanschoren, Automated Machine Learning . Springer, 2019
2019
Later among the works it cites.
J. Shu, Q. Xie, L. Yi, Q. Zhao, S. Zhou, Z. Xu, and D. Meng, “Meta-weight-net: Learning an explicit mapping for sample weighting,” in NeurIPS , 2019
2019
Later among the works it cites.
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” in ICLR , 2019
2019
Later among the works it cites.
S. Vaswani, F. Bach, and M. Schmidt, “Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron,” in The 22nd International Conference on Artificial Intelligence and Statistics . PMLR, 2019, pp. 1195–1204
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Q. Yang, Y. Zhang, W. Dai, and S. J. Pan, Transfer learning . Cambridge University Press, 2020
2020
Closest in time.