Fetching the paper…
Reading the bibliography…
Ever since Reddi et al.
Gradient convergence in gradient methods with errors
Bertsekas, D. P. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
An improved analysis of stochastic gradient descent with momentum
Liu, Y., Gao, Y., and Yin, W · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Roux, N. L · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Dozat, T · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A · 2017
Cited alongside, same era.
De, S., Mukherjee, A., and Ullah, E · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
On the convergence of adam and adagrad
Défossez, A., Bottou, L., Bach, F., and Usunier, N · 2020
Later among the works it cites.
Asymptotic study of stochastic adaptive algorithm in non-convex landscape
Gadat, S. and Gavra, I · 2020
Later among the works it cites.
Rmsprop converges with proper hyper-parameter
Shi, N., Li, D., Hong, M., and Sun, R · 2020
Later among the works it cites.
Towards practical adam: Non-convexity, convergence theory, and mini-batch acceleration
Chen, C., Shen, L., Zou, F., and Liu, W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
A unified analysis of stochastic momentum methods for deep learning
Yan, Y., Yang, T., Li, Z., Lin, Q., and Yang, Y · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Zaheer, M., Reddi, S., Sachan, D., Kale, S., and Kumar, S · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., and Liu, Y · 2019
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M · 2019
Cited alongside, same era.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Yu, H., Jin, R., and Yang, S · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
A novel convergence analysis for algorithms of the adam family
Guo, Z., Xu, Y., Yin, W., Jin, R., and Yang, T · 2021
Later among the works it cites.
Super-adam: Faster and universal framework of adaptive gradients
Huang, F., Li, J., and Huang, H · 2021
Later among the works it cites.
Theoretical analysis of adam using hyperparameters close to one without lipschitz smoothness
Iiduka, H · 2022
Closest in time.
On the convergence of msgd and adagrad for stochastic optimization
Jin, R., Xing, Y., and He, X · 2022
Closest in time.
Adam: A method for stochastic optimization
Scholar, G · 2022
Closest in time.
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al · 2022
Closest in time.
Does adam converge and when?
Zhang, Y., Chen, C., and Luo, Z.-Q · 2022
Closest in time.