Fetching the paper…
Reading the bibliography…
Nowadays, online learning is an appealing learning paradigm, which is of great interest in practice due to the recent emergence of large scale applications such as online advertising placement and online web ranking.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K. (1989) · 1989
Earlier work this paper cites.
On-line algorithms in machine learning
Blum, A. (1998) · 1998
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M. (2003) · 2003
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S. (2007) · 2007
Earlier work this paper cites.
Adaptive algorithms for online decision problems
Hazan, E. and Seshadhri, C. (2007) · 2007
Earlier work this paper cites.
Lecture notes on online learning
Rakhlin, A. and Tewari, A. (2009) · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Interior-point methods for full-information and bandit online learning
Abernethy, J. D., Hazan, E., and Rakhlin, A. (2012) · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Strongly adaptive online learning
Daniely, A., Gonen, A., and Shalev-Shwartz, S. (2015) · 2015
Cited alongside, same era.
ADAM: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Online gradient descent in function space
Zhu, C. and Xu, H. (2015) · 2015
Cited alongside, same era.
Deep Learning
First-order methods almost always avoid saddle points
Lee, J. D., Panageas, I., Piliouras, G., Simchowitz, M., Jordan, M. I., and Recht, B. (2017) · 2017
Later among the works it cites.
Convergence analysis of two-layer neural networks with ReLU activation
Li, Y. and Yuan, Y. (2017) · 2017
Later among the works it cites.
Online learning to rank in stochastic click models
Zoghi, M., Tunys, T., Ghavamzadeh, M., Kveton, B., Szepesvari, C., and Wen, Z. (2017) · 2017
Later among the works it cites.
On the convergence of ADAM and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2018) · 2018
Later among the works it cites.
No spurious local minima in a two hidden unit ReLU network
Wu, C., Luo, J., and Lee, J. D. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kawaguchi, K. (2016) · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Lee, J. D., Simchowitz, M., Jordan, M. I., and Recht, B. (2016) · 2016
Cited alongside, same era.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y. (2015a)
Cited in the paper.
Open problem: The landscape of the loss surfaces of multilayer networks
Choromanska, A., LeCun, Y., and Arous, G. B. (2015b)
Cited in the paper.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
Du, S., Lee, J., Tian, Y., Singh, A., and Poczos, B. (2018a)
Cited in the paper.
When is a convolutional filter easy to learn?
Du, S. S., Lee, J. D., and Tian, Y. (2018b)
Cited in the paper.
On the convergence of a class of ADAM-type algorithms for non-convex optimization
Chen, X., Liu, S., Sun, R., and Hong, M. (2019) · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., and Liu, Y. (2019) · 2019
Closest in time.
Adaptive regret of convex and smooth functions
Zhang, L., Liu, T.-Y., and Zhou, Z.-H. (2019) · 2019
Closest in time.