Fetching the paper…
Reading the bibliography…
The adaptive momentum method (AdaMM), which uses past gradients to update descent directions and learning rates simultaneously, has become one of the most popular first-order optimization methods for solving machine learning problems.
“A genetic algorithm tutorial,”
D. Whitley, · 1994
Earlier work this paper cites.
“A direct search optimization method that models the objective and constraint functions by linear interpolation,”
Michael JD Powell, · 1994
Earlier work this paper cites.
“An elementary proof of a theorem of johnson and lindenstrauss,”
S. Dasgupta and A. Gupta, · 2003
Earlier work this paper cites.
“Mesh adaptive direct search algorithms for constrained optimization,”
Charles Audet and John E Dennis Jr, · 2006
Earlier work this paper cites.
“Global convergence of general derivative-free trust-region algorithms to first-and second-order critical points,”
A. R. Conn, K. Scheinberg, and L. Vicente, · 2009
Earlier work this paper cites.
Introduction to derivative-free optimization
A. R. Conn, K. Scheinberg, and L. N. Vicente, · 2009
Earlier work this paper cites.
“Pswarm: a hybrid solver for linearly constrained global derivative-free optimization,”
A Ismael F Vaz and Luís N Vicente, · 2009
Earlier work this paper cites.
“The bobyqa algorithm for bound constrained optimization without derivatives,”
Michael JD Powell, · 2009
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li, · 2009
Earlier work this paper cites.
“Algorithm 909: Nomad: Nonlinear optimization with the mads algorithm,”
Sébastien Le Digabel, · 2011
Earlier work this paper cites.
“Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,”
T. Tieleman and G. Hinton, · 2012
Earlier work this paper cites.
“Stochastic first-and zeroth-order methods for nonconvex stochastic programming,”
S. Ghadimi and G. Lan, · 2013
Earlier work this paper cites.
“Derivative-free optimization: a review of algorithms and comparison of software implementations,”
L. M. Rios and N. V. Sahinidis, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“On the information-adaptive variants of the ADMM: an iteration complexity perspective,”
X. Gao, B. Jiang, and S. Zhang, · 2014
Earlier work this paper cites.
“Efficient and robust automated machine learning,”
M. Feurer, A. Klein, K. Eggensperger, J. Springenberg, M. Blum, and F. Hutter, · 2015
Earlier work this paper cites.
“Random gradient-free minimization of convex functions,”
Y. Nesterov and V. Spokoiny, · 2015
Earlier work this paper cites.
“Optimal rates for zero-order convex optimization: The power of two function evaluations,”
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and A. Wibisono, · 2015
Cited alongside, same era.
“Taking the human out of the loop: A review of bayesian optimization,”
B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. De Freitas, · 2016
Cited alongside, same era.
“A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order,”
X. Lian, H. Zhang, C.-J. Hsieh, Y. Huang, and J. Liu, · 2016
Cited alongside, same era.
“Zeroth-order asynchronous doubly stochastic algorithm with variance reduction,”
B. Gu, Z. Huo, and H. Huang, · 2016
Cited alongside, same era.
“Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization,”
S. Ghadimi, G. Lan, and H. Zhang, · 2016
Cited alongside, same era.
C.-C. Tu, P. Ting, P.-Y. Chen, S. Liu, H. Zhang, J. Yi, C.-J. Hsieh, and S.-M. Cheng, · 2018
Later among the works it cites.
“signsgd: compressed optimisation for non-convex problems,”
J. Bernstein, Y. Wang, K. Azizzadenesheli, and A. Anandkumar, · 2018
Later among the works it cites.
“On the convergence of adam and beyond,”
S. J. Reddi, S. Kale, and S. Kumar, · 2018
Later among the works it cites.
“Closing the generalization gap of adaptive gradient methods in training deep neural networks,”
J. Chen and Q. Gu, · 2018
Later among the works it cites.
“On the convergence of a class of adam-type algorithms for non-convex optimization,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Adversarial examples in the physical world,”
A. Kurakin, I. Goodfellow, and S. Bengio, · 2016
Cited alongside, same era.
“Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization,”
S. J. Reddi, S. Sra, B. Poczos, and A. J. Smola, · 2016
Cited alongside, same era.
“Rethinking the inception architecture for computer vision,”
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, · 2016
Cited alongside, same era.
“Auto-weka 2.0: Automatic model selection and hyperparameter optimization in weka,”
L. Kotthoff, C. Thornton, H. H. Hoos, F. Hutter, and K. Leyton-Brown, · 2017
Cited alongside, same era.
“Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models,”
P.-Y. Chen, H. Zhang, Y. Sharma, J. Yi, and C.-J. Hsieh, · 2017
Cited alongside, same era.
Derivative-free and blackbox optimization
Charles Audet and Warren Hare, · 2017
Cited alongside, same era.
“Towards deep learning models resistant to adversarial attacks,”
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu, · 2017
Cited alongside, same era.
X. Chen, S. Liu, R. Sun, and M. Hong, · 2018
Later among the works it cites.
“Dissecting adam: The sign, magnitude and variance of stochastic gradients,”
L. Balles and P. Hennig, · 2018
Later among the works it cites.
“Zeroth-order stochastic variance reduction for nonconvex optimization,”
S. Liu, B. Kailkhura, P.-Y. Chen, P. Ting, S. Chang, and L. Amini, · 2018
Later among the works it cites.
“Stochastic zeroth-order optimization via variance reduction method,”
L. Liu, M. Cheng, C.-J. Hsieh, and D. Tao, · 2018
Later among the works it cites.
“Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates,”
Krishnakumar Balasubramanian and Saeed Ghadimi, · 2018
Later among the works it cites.
“A Frank-Wolfe framework for efficient and effective adversarial attacks,”
J. Chen, J. Yi, and Q. Gu, · 2018
Later among the works it cites.
“On the convergence of adaptive gradient methods for nonconvex optimization,”
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu, · 2018
Later among the works it cites.
“Prior convictions: Black-box adversarial attacks with bandits and priors,”
A. Ilyas, L. Engstrom, and A. Madry, · 2018
Later among the works it cites.
“Black-box adversarial attacks with limited queries and information,”
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin, · 2018
Later among the works it cites.
“Query-efficient hard-label black-box attack: An optimization-based approach,”
M. Cheng, T. Le, P.-Y. Chen, J. Yi, H. Zhang, and C.-J. Hsieh, · 2018
Later among the works it cites.
“signSGD via zeroth-order oracle,”
S. Liu, P.-Y. Chen, X. Chen, and M. Hong, · 2019
Closest in time.
“On the convergence proof of amsgrad and a new version,”
T. T. Phuong and L. T. Phong, · 2019
Closest in time.
“Structured adversarial attack: Towards general implementation and better interpretability,”
Kaidi Xu, Sijia Liu, Pu Zhao, Pin-Yu Chen, Huan Zhang, Quanfu Fan, Deniz Erdogmus, Yanzhi Wang, and Xue Lin, · 2019
Closest in time.