Fetching the paper…
Reading the bibliography…
We develop a new algorithm for non-convex stochastic optimization that finds an $\epsilon$-critical point in the optimal $O(\epsilon^{-3})$ stochastic gradient and Hessian-vector product computations.
On the convergence of adam and beyond. arxiv 2019
Reddi, S., Kale, S., and Kumar, S · 1904
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Fast exact multiplication by the hessian
Pearlmutter, B. A · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
McMahan, H. B. and Streeter, M. J · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Darknet: Open source neural networks in c
Redmon, J · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Agarwal, N., Allen-Zhu, Z., Bullins, B., Hazan, E., and Ma, T · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M · 2017
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Srebro, N., and Woodworth, B · 2019
Later among the works it cites.
Momentum-based variance reduction in non-convex sgd
Cutkosky, A. and Orabona, F · 2019
Later among the works it cites.
On the ineffectiveness of variance reduced optimization for deep learning
Defazio, A. and Bottou, L · 2019
Later among the works it cites.
A limited-memory quasi-newton algorithm for bound-constrained non-smooth optimization
Keskar, N. and Wächter, A · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
A hybrid stochastic optimization framework for stochastic composite nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerating hessian-free gauss-newton full-waveform inversion via l-bfgs preconditioned conjugate-gradient algorithm
Pan, W., Innanen, K. A., and Liao, W · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Stochastic cubic regularization for fast nonconvex optimization
Tripuraneni, N., Stern, M., Jin, C., Regier, J., and Jordan, M. I · 2017
Cited alongside, same era.
A progressive batching l-bfgs method for machine learning
Bollapragada, R., Nocedal, J., Mudigere, D., Shi, H.-J., and Tang, P. T. P · 2018
Cited alongside, same era.
Accelerated methods for nonconvex optimization
Carmon, Y., Duchi, J. C., Hinder, O., and Sidford, A · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
Second-order information in non-convex stochastic optimization: Power and limitations
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Sekhari, A., and Sridharan, K
Cited in the paper.
Tran-Dinh, Q., Pham, N. H., Phan, D. T., and Nguyen, L. M · 2019
Later among the works it cites.
Stochastic variance-reduced cubic regularization methods
Zhou, D., Xu, P., and Gu, Q · 2019
Later among the works it cites.
Momentum improves normalized sgd
Cutkosky, A. and Mehta, H · 2020
Later among the works it cites.
Ma, X · 2020
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Xu, P., Roosta, F., and Mahoney, M. W · 2020
Later among the works it cites.
Adahessian: An adaptive second order optimizer for machine learning
Yao, Z., Gholami, A., Shen, S., Keutzer, K., and Mahoney, M. W · 2020
Later among the works it cites.