Fetching the paper…
Reading the bibliography…
Distributed adaptive stochastic gradient methods have been widely used for large-scale nonconvex optimization, such as training deep learning models.
C. M. Bishop, Pattern recognition and machine learning . springer, 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” Technical Report TR-2009, University of Toronto, Toronto , 2009
2009
Earlier work this paper cites.
A. Smola and S. Narayanamurthy, “An architecture for parallel topic models,” Proceedings of the VLDB Endowment , vol. 3, no. 1-2, pp. 703–710, 2010
2010
Earlier work this paper cites.
A. Agarwal and J. C. Duchi, “Distributed delayed stochastic optimization,” in Advances in Neural Information Processing Systems , 2011, pp. 873–881
2011
Earlier work this paper cites.
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies . Portland, Oregon, USA: Association for Computational Linguistics, June 2011, pp. 142–150. [Online]. Available: http://www.aclweb.org/anthology/P11-1015
2011
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
M. Li, T. Zhang, Y. Chen, and A. J. Smola, “Efficient mini-batch training for stochastic optimization,” in Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining , 2014, pp. 661–670
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Hastie, R. Tibshirani, and M. Wainwright, Statistical learning with sparsity: the lasso and generalizations . CRC press, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “Qsgd: Communication-efficient sgd via gradient quantization and encoding,” in Advances in Neural Information Processing Systems , 2017, pp. 1709–1720
2017
Earlier work this paper cites.
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu, “Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent,” in Advances in Neural Information Processing Systems , 2017, pp. 5330–5340
2017
Earlier work this paper cites.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in Advances in neural information processing systems , 2017, pp. 1509–1519
2017
Cited alongside, same era.
2017
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP , 2018, pp. 353–355
2018
Cited alongside, same era.
F. Zou, L. Shen, Z. Jie, W. Zhang, and W. Liu, “A sufficient condition for convergences of adam and rmsprop,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 11 127–11 135
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. U. Stich, J.-B. Cordonnier, and M. Jaggi, “Sparsified sgd with memory,” in Advances in Neural Information Processing Systems , 2018, pp. 4447–4458
2018
Cited alongside, same era.
2018
Cited alongside, same era.
P. Jiang and G. Agrawal, “A linear speedup analysis of distributed deep learning with sparse and quantized communication,” in Advances in Neural Information Processing Systems , 2018, pp. 2525–2536
2018
Cited alongside, same era.
J. Wangni, J. Wang, J. Liu, and T. Zhang, “Gradient sparsification for communication-efficient distributed optimization,” in Advances in Neural Information Processing Systems , 2018, pp. 1299–1309
2018
Cited alongside, same era.
S. Magnússon, C. Enyioha, N. Li, C. Fischione, and V. Tarokh, “Communication complexity of dual decomposition methods for distributed resource allocation optimization,” IEEE Journal of Selected Topics in Signal Processing , vol. 12, no. 4, pp. 717–732, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Later among the works it cites.
C. Chen, L. Shen, H. Huang, and W. Liu, “Quantized adam with error feedback,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 12, no. 5, pp. 1–26, 2021
2021
Later among the works it cites.
N. Shi, D. Li, M. Hong, and R. Sun, “RMSprop converges with proper hyper-parameter,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=3UDSdyIcBDA
2021
Later among the works it cites.
C. Chen, L. Shen, F. Zou, and W. Liu, “Towards practical adam: Non-convexity, convergence theory, and mini-batch acceleration,” The Journal of Machine Learning Research , vol. 23, no. 1, pp. 10 411–10 457, 2022
2022
Closest in time.
Y. Wang, L. Lin, and J. Chen, “Communication-compressed adaptive gradient method for distributed nonconvex optimization,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2022, pp. 6292–6320
2022
Closest in time.
M. Doostmohammadian, A. Aghasi, A. I. Rikos, A. Grammenos, E. Kalyvianaki, C. N. Hadjicostis, K. H. Johansson, and T. Charalambous, “Distributed anytime-feasible resource allocation subject to heterogeneous time-varying delays,” IEEE Open Journal of Control Systems , vol. 1, pp. 255–267, 2022
2022
Closest in time.
M. Doostmohammadian, A. Aghasi, M. Pirani, E. Nekouei, U. A. Khan, and T. Charalambous, “Fast-convergent anytime-feasible dynamics for distributed allocation of resources over switching sparse networks with quantized communication links,” in 2022 European Control Conference (ECC) . IEEE, 2022, pp. 84–89
2022
Closest in time.
2022
Closest in time.
2023
Closest in time.
L. Shen, C. Chen, F. Zou, Z. Jie, J. Sun, and W. Liu, “A unified analysis of adagrad with weighted aggregation and momentum acceleration,” IEEE Transactions on Neural Networks and Learning Systems , 2023
2023
Closest in time.