Fetching the paper…
Reading the bibliography…
Zeroth-order optimization (ZO) algorithms have been recently used to solve black-box or simulation-based learning and control problems, where the gradient of the objective function cannot be easily computed but can be approximated using the objective function values.
Online convex optimization in the bandit setting: gradient descent without a gradient
Flaxman, A. D., Kalai, A. T., & McMahan, H. B. (2005) · 2005
Earlier work this paper cites.
Cooperative multi-agent reinforcement learning with partial observations
Zhang, Y., & Zavlanos, M. M. (2020) · 2006
Earlier work this paper cites.
Extremum seeking control: Convergence analysis
Nešić, D. (2009) · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Agarwal, A., Dekel, O., & Xiao, L. (2010) · 2010
Earlier work this paper cites.
Boosting one-point derivative-free online optimization via residual feedback
Zhang, Y., Zhou, Y., Ji, K., & Zavlanos, M. M. (2020) · 2010
Earlier work this paper cites.
Improved regret guarantees for online smooth convex optimization with bandit feedback
Saha, A., & Tewari, A. (2011) · 2011
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S., & Lan, G. (2013) · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Nesterov, Y. (2013) · 2013
Earlier work this paper cites.
On the complexity of bandit and derivative-free stochastic convex optimization
Shamir, O. (2013) · 2013
Earlier work this paper cites.
Bandit smooth convex optimization: Improving the bias-variance tradeoff
Dekel, O., Eldan, R., & Koren, T. (2015) · 2015
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
Duchi, J. C., Jordan, M. I., Wainwright, M. J., & Wibisono, A. (2015) · 2015
Cited alongside, same era.
Highly-smooth zero-th order online optimization
Bach, F., & Perchet, V. (2016) · 2016
Cited alongside, same era.
(bandit) convex optimization with biased noisy gradient oracles
Hu, X., Prashanth, L., György, A., & Szepesvári, C. (2016) · 2016
Cited alongside, same era.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., & Hsieh, C.-J. (2017) · 2017
Cited alongside, same era.
Stochastic online optimization. single-point and multi-point non-linear multi-armed bandits. convex and strongly-convex case
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Shamir, O. (2017) · 2017
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S., & Mesbahi, M. (2018) · 2018
Later among the works it cites.
Derivative-free methods for policy optimization: Guarantees for linear quadratic systems
Malik, D., Pananjady, A., Bhatia, K., Khamaru, K., Bartlett, P. L., & Wainwright, M. J. (2018) · 2018
Later among the works it cites.
Derivative-free optimization methods
Larson, J., Menickelly, M., & Wild, S. M. (2019) · 2019
Later among the works it cites.
Exploiting higher order smoothness in derivative-free optimization and continuous bandits
Akhavan, A., Pontil, M., & Tsybakov, A. (2020) · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gasnikov, A. V., Krymova, E. A., Lagunovskaya, A. A., Usmanova, I. N., & Fedorenko, F. A. (2017) · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., & Mordatch, I. (2017) · 2017
Cited alongside, same era.
Random gradient-free minimization of convex functions
Nesterov, Y., & Spokoiny, V. (2017) · 2017
Cited alongside, same era.
Socially-aware robot planning via bandit human feedback
Luo, X., Zhang, Y., & Zavlanos, M. M. (2020) · 2020
Closest in time.
Surrogate-based distributed optimisation for expensive black-box functions
Li, Z., Dong, Z., Liang, Z., & Ding, Z. (2021) · 2021
Closest in time.
Robust hybrid zero-order optimization algorithms with acceleration via averaging in time
Poveda, J. I., & Li, N. (2021) · 2021
Closest in time.