Fetching the paper…
Reading the bibliography…
In this paper, we provide a rigorous proof of convergence of the Adaptive Moment Estimate (Adam) algorithm for a wide class of optimization objectives.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Nicolas Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Optimization with first-order surrogate functions
Julien Mairal · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J. Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Earlier work this paper cites.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros · 2017
Earlier work this paper cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Earlier work this paper cites.
Convergence guarantees for rmsprop and adam in non-convex optimization and an empirical comparison to nesterov acceleration
Soham De, Anirbit Mukherjee, and Enayat Ullah · 2018
Earlier work this paper cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
A sufficient condition for convergences of adam and rmsprop
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2018
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Towards practical adam: Non-convexity, convergence theory, and mini-batch acceleration
Congliang Chen, Li Shen, Fangyu Zou, and Wei Liu · 2022
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang, and Zhenxun Zhuang · 2022
Later among the works it cites.
The power of adaptivity in sgd: Self-tuning step sizes with unbounded gradients and affine variance
Matthew Faw, Isidoros Tziotis, Constantine Caramanis, Aryan Mokhtari, Sanjay Shakkottai, and Rachel Ward · 2022
Later among the works it cites.
Asymptotic study of stochastic adaptive algorithms in non-convex landscape
Sébastien Gadat and Ioana Gavra · 2022
Later among the works it cites.
Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhirui Ma, Tie-Yan Liu, and Wei Chen · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Cited alongside, same era.
Hybrid stochastic gradient descent algorithms for stochastic nonconvex optimization
Quoc Tran-Dinh, Nhan H Pham, Dzung T Phan, and Lam M Nguyen · 2019
Cited alongside, same era.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
J. Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
A simple convergence proof of adam and adagrad
Alexandre D’efossez, Léon Bottou, Francis R. Bach, and Nicolas Usunier · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Cited alongside, same era.
An optimal hybrid variance-reduced algorithm for stochastic composite nonconvex optimization
Deyi Liu, Lam M Nguyen, and Quoc Tran-Dinh · 2020
Cited alongside, same era.
Ruiqi Wang and Diego Klabjan · 2022
Later among the works it cites.
Adam can converge without any modification on update rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhimin Luo · 2022
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, and Suvrit Sra · 2023
Closest in time.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2023
Closest in time.
Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization
Ziyi Chen, Yi Zhou, Yingbin Liang, and Zhaosong Lu · 2023
Closest in time.
Beyond uniform smoothness: A stopped analysis of adaptive sgd
Matthew Faw, Litu Rout, Constantine Caramanis, and Sanjay Shakkottai · 2023
Closest in time.
Theoretical analysis of adam using hyperparameters close to one without lipschitz smoothness
Hideaki Iiduka · 2023
Closest in time.
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
Convex and non-convex optimization under generalized smoothness
Haochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin, and Ali Jadbabaie · 2023
Closest in time.
Zijian Liu, Perry Dong, Srikanth Jagabathula, and Zhengyuan Zhou · 2023
Closest in time.
Variance-reduced clipping for non-convex optimization
Amirhossein Reisizadeh, Haochuan Li, Subhro Das, and Ali Jadbabaie · 2023
Closest in time.
Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions
Bohan Wang, Huishuai Zhang, Zhiming Ma, and Wei Chen · 2023
Closest in time.