Fetching the paper…
Reading the bibliography…
Adaptive gradient-based optimization methods such as \textsc{Adagrad}, \textsc{Rmsprop}, and \textsc{Adam} are widely used in solving large-scale machine learning problems including deep learning.
H. Robbins and S. Monro, “A stochastic approximation method,” in Herbert Robbins Selected Papers
1985
Earlier work this paper cites.
J. Tsitsiklis, D. Bertsekas, and M. Athans, “Distributed asynchronous deterministic and stochastic gradient optimization algorithms,” IEEE transactions on automatic control
1986
Earlier work this paper cites.
Cambridge university press, 1990
R. A. Horn, R. A. Horn, and C. R. Johnson, Matrix analysis · 1990
Earlier work this paper cites.
D. Li, K. D. Wong, Y. H. Hu, and A. M. Sayeed, “Detection, classification, and tracking of targets,” IEEE signal processing magazine
2002
Earlier work this paper cites.
X. Zhang, J. Zhang, and L. Liao, “An adaptive trust region method and its convergence,” Science in China Series A: Mathematics
2002
Earlier work this paper cites.
M. Zinkevich, “Online convex programming and generalized infinitesimal gradient ascent,” 2003
2003
Earlier work this paper cites.
A. Beck and M. Teboulle, “Mirror descent and nonlinear projected subgradient methods for convex optimization,” Operations Research Letters
2003
Earlier work this paper cites.
M. Rabbat and R. Nowak, “Distributed optimization in sensor networks,” in Proceedings of the 3rd international symposium on Information processing in sensor networks
2004
Earlier work this paper cites.
S. Boyd, P. Diaconis, and L. Xiao, “Fastest mixing markov chain on a graph,” SIAM review
2004
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research
2011
Earlier work this paper cites.
S. Shalev-Shwartz et al
2012
Earlier work this paper cites.
Springer Science & Business Media, 2012
V. Lesser, C. L. Ortiz Jr, and M. Tambe, Distributed sensor networks: A multiagent perspective · 2012
Earlier work this paper cites.
J. C. Duchi, A. Agarwal, and M. J. Wainwright, “Dual averaging for distributed optimization: Convergence analysis and network scaling,” IEEE Transactions on Automatic control
2012
Earlier work this paper cites.
E. Wei and A. Ozdaglar, “Distributed alternating direction method of multipliers,” 2012
2012
Earlier work this paper cites.
M. D. Zeiler, “Adadelta: an adaptive learning rate method,” arXiv preprint arXiv:1212.5701
2012
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
T.-H. Chang, A. Nedić, and A. Scaglione, “Distributed constrained optimization by consensus-based primal-dual perturbation method,” IEEE Transactions on Automatic Control
2014
Cited alongside, same era.
W. Shi, Q. Ling, K. Yuan, G. Wu, and W. Yin, “On the linear convergence of the admm in decentralized consensus optimization,” IEEE Transactions on Signal Processing
2014
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980
2014
Cited alongside, same era.
D. Ataee Tarzanagh, M. R. Peyghami, and H. Mesgarani, “A new nonmonotone trust region method for unconstrained optimization equipped by an efficient adaptive radius,” Optimization Methods and Software
2014
Cited alongside, same era.
2017
Later among the works it cites.
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent,” in Advances in Neural Information Processing Systems
2017
Later among the works it cites.
T. Tieleman and G. Hinton, “Divide the gradient by a running average of its recent magnitude. coursera: Neural networks for machine learning,” tech. rep., Technical Report. Available online: https://zh. coursera. org/learn/neuralnetworks/lecture/YQHki/rmsprop-divide-the-gradient-by-a-running-average-of-its-recent-magnitude (accessed on 21 April 2017)
2017
Later among the works it cites.
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht, “The marginal value of adaptive gradient methods in machine learning,” in Advances in Neural Information Processing Systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
D. Mateos-Núnez and J. Cortés, “Distributed online convex optimization over jointly connected digraphs,” IEEE Transactions on Network Science and Engineering
2014
Cited alongside, same era.
E. C. Hall and R. M. Willett, “Online convex optimization in dynamic environments,” IEEE Journal of Selected Topics in Signal Processing
2015
Cited alongside, same era.
O. Besbes, Y. Gur, and A. Zeevi, “Non-stationary stochastic optimization,” Operations Research
2015
Cited alongside, same era.
W. Shi, Q. Ling, G. Wu, and W. Yin, “Extra: An exact first-order algorithm for decentralized consensus optimization,” SIAM Journal on Optimization
2015
Cited alongside, same era.
D. A. Tarzanagh, M. R. Peyghami, and F. Bastin, “A new nonmonotone adaptive retrospective trust region method for unconstrained optimization problems,” Journal of Optimization Theory and Applications
2015
Cited alongside, same era.
E. Hazan et al
2016
Cited alongside, same era.
A. Mokhtari, S. Shahrampour, A. Jadbabaie, and A. Ribeiro, “Online optimization in dynamic environments: Improved regret rates for strongly convex problems,” in Decision and Control (CDC), 2016 IEEE 55th Conference on
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Shahrampour and A. Jadbabaie, “An online optimization approach for multi-agent tracking of dynamic parameters in the presence of adversarial noise,” in American Control Conference (ACC), 2017
2017
Later among the works it cites.
S. Shahrampour and A. Jadbabaie, “Distributed online optimization in dynamic environments using mirror descent,” IEEE Transactions on Automatic Control
2018
Later among the works it cites.
J. Zeng and W. Yin, “On nonconvex decentralized gradient descent,” IEEE Transactions on Signal Processing
2018
Later among the works it cites.
2018
Later among the works it cites.
S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of adam and beyond,” in International Conference on Learning Representations
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
S. De, A. Mukherjee, and E. Ullah, “Convergence guarantees for rmsprop and adam in non-convex optimization and an empirical comparison to nesterov acceleration,” 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Zaheer, S. Reddi, D. Sachan, S. Kale, and S. Kumar, “Adaptive methods for nonconvex optimization,” in Advances in Neural Information Processing Systems
2018
Later among the works it cites.
A. Nedić, A. Olshevsky, and M. G. Rabbat, “Network topology and communication-computation tradeoffs in decentralized optimization,” Proceedings of the IEEE
2018
Later among the works it cites.
H. Kasai, “Sgdlibrary: A matlab library for stochastic optimization algorithms,” Journal of Machine Learning Research
2018
Later among the works it cites.