Fetching the paper…
Reading the bibliography…
Zeroth-order (a.k.a, derivative-free) methods are a class of effective optimization methods for solving complex machine learning problems, where gradients of the objective functions are not available or computationally prohibitive.
D. Gabay and B. Mercier, “A dual algorithm for the solution of nonlinear variational problems via finite element approximation,” Computers & Mathematics with Applications , vol. 2, no. 1, pp. 17–40, 1976
1976
Earlier work this paper cites.
J. Wright, A. Ganesh, S. R. Rao, Y. Peng, and Y. Ma, “Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization,” in NIPS , 2009
2009
Earlier work this paper cites.
S. Kim, K.-A. Sohn, and E. P. Xing, “A multivariate regression approach to association analysis of a quantitative trait network,” Bioinformatics , vol. 25, no. 12, pp. i204–i212, 2009
2009
Earlier work this paper cites.
A. Beck and M. Teboulle, “A fast iterative shrinkage-thresholding algorithm for linear inverse problems,” SIAM journal on imaging sciences , vol. 2, no. 1, pp. 183–202, 2009
2009
Earlier work this paper cites.
A. Agarwal, O. Dekel, and L. Xiao, “Optimal algorithms for online convex optimization with multi-point bandit feedback.” in COLT . Citeseer, 2010, pp. 28–40
2010
Earlier work this paper cites.
E. J. Candès, X. Li, Y. Ma, and J. Wright, “Robust principal component analysis?” Journal of the ACM (JACM) , vol. 58, no. 3, pp. 1–37, 2011
2011
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,” Foundations and Trends® in Machine Learning , vol. 3, no. 1, pp. 1–122, 2011
2011
Earlier work this paper cites.
A. Coates, A. Ng, and H. Lee, “An analysis of single-layer networks in unsupervised feature learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics , 2011, pp. 215–223
2011
Earlier work this paper cites.
B. He and X. Yuan, “On the o(1/n) convergence rate of the douglas–rachford alternating direction method,” SIAM Journal on Numerical Analysis , vol. 50, no. 2, pp. 700–709, 2012
2012
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in NIPS , 2013, pp. 315–323
2013
Earlier work this paper cites.
N. Simon, J. Friedman, T. Hastie, and R. Tibshirani, “A sparse-group lasso,” Journal of computational and graphical statistics , vol. 22, no. 2, pp. 231–245, 2013
2013
Earlier work this paper cites.
H. Ouyang, N. He, L. Tran, and A. G. Gray, “Stochastic alternating direction method of multipliers.” ICML , vol. 28, pp. 80–88, 2013
2013
Earlier work this paper cites.
T. Suzuki, “Dual averaging and proximal gradient descent for online alternating direction multiplier method.” in ICML , 2013, pp. 392–400
2013
Earlier work this paper cites.
——, “Stochastic dual coordinate ascent with alternating direction method of multipliers,” in ICML , 2014, pp. 736–744
2014
Earlier work this paper cites.
R. Nishihara, L. Lessard, B. Recht, A. Packard, and M. Jordan, “A general analysis of the convergence of admm,” in International Conference on Machine Learning , 2015, pp. 343–352
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Wang, W. Yin, and J. Zeng, “Global convergence of admm in nonconvex nonsmooth optimization,” Journal of Scientific Computing , pp. 1–35, 2015
2015
Earlier work this paper cites.
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and A. Wibisono, “Optimal rates for zero-order convex optimization: The power of two function evaluations,” IEEE TIT , vol. 61, no. 5, pp. 2788–2806, 2015
2015
Earlier work this paper cites.
X. Hu, L. Prashanth, A. György, and C. Szepesvari, “(bandit) convex optimization with biased noisy gradient oracles,” in Artificial Intelligence and Statistics . PMLR, 2016, pp. 819–828
2016
Earlier work this paper cites.
S. Ghadimi, G. Lan, and H. Zhang, “Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization,” Mathematical Programming , vol. 155, no. 1-2, pp. 267–305, 2016
2016
Cited alongside, same era.
Z. Allen-Zhu and E. Hazan, “Variance reduction for faster non-convex optimization,” in International conference on machine learning . PMLR, 2016, pp. 699–707
2016
Cited alongside, same era.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic variance reduction for nonconvex optimization,” in International conference on machine learning . PMLR, 2016, pp. 314–323
2016
Cited alongside, same era.
X. Lian, H. Zhang, C.-J. Hsieh, Y. Huang, and J. Liu, “A comprehensive linear speedup analysis for asynchronous stochastic parallel optimization from zeroth-order to first-order,” in Advances in Neural Information Processing Systems , 2016, pp. 3054–3062
2016
Cited alongside, same era.
2018
Later among the works it cites.
S. Liu, B. Kailkhura, P.-Y. Chen, P. Ting, S. Chang, and L. Amini, “Zeroth-order stochastic variance reduction for nonconvex optimization,” in NIPS , 2018, pp. 3731–3741
2018
Later among the works it cites.
S. Liu, J. Chen, P.-Y. Chen, and A. Hero, “Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications,” in AISTATS , vol. 84, 2018, pp. 288–297
2018
Later among the works it cites.
X. Gao, B. Jiang, and S. Zhang, “On the information-adaptive variants of the admm: an iteration complexity perspective,” Journal of Scientific Computing , vol. 76, no. 1, pp. 327–363, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. He, F. Ma, and X. Yuan, “Convergence study on the symmetric version of admm with larger step sizes,” SIAM Journal on Imaging Sciences , vol. 9, no. 3, pp. 1467–1501, 2016
2016
Cited alongside, same era.
S. Zheng and J. T. Kwok, “Fast and light stochastic admm,” in IJCAI , 2016
2016
Cited alongside, same era.
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, and T. Goldstein, “Training neural networks without gradients: a scalable admm approach,” in ICML , 2016, pp. 2722–2731
2016
Cited alongside, same era.
M. Hong, Z.-Q. Luo, and M. Razaviyayn, “Convergence analysis of alternating direction method of multipliers for a family of nonconvex problems,” SIAM Journal on Optimization , vol. 26, no. 1, pp. 337–364, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
C. Chen, B. He, Y. Ye, and X. Yuan, “The direct extension of admm for multi-block convex minimization problems is not necessarily convergent,” Mathematical Programming , vol. 155, no. 1-2, pp. 57–79, 2016
2016
Cited alongside, same era.
S. Reddi, S. Sra, B. Poczos, and A. J. Smola, “Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization,” in NIPS , 2016, pp. 1145–1153
2016
Cited alongside, same era.
P.-Y. Chen, H. Zhang, Y. Sharma, J. Yi, and C.-J. Hsieh, “Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models,” in Workshop on Artificial Intelligence and Security . ACM, 2017, pp. 15–26
2017
Cited alongside, same era.
K. Balasubramanian and S. Ghadimi, “Zeroth-order (non)-convex stochastic optimization via conditional gradient and gradient updates,” in Advances in Neural Information Processing Systems , 2018, pp. 3455–3464
2018
Later among the works it cites.
2018
Later among the works it cites.
C. Fang, C. J. Li, Z. Lin, and T. Zhang, “Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator,” in Advances in Neural Information Processing Systems , 2018, pp. 689–699
2018
Later among the works it cites.
B. Gu, Z. Huo, C. Deng, and H. Huang, “Faster derivative-free stochastic algorithm for shared memory machines,” in ICML , 2018, pp. 1807–1816
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Larson, M. Menickelly, and S. M. Wild, “Derivative-free optimization methods,” Acta Numerica , vol. 28, pp. 287–404, 2019
2019
Closest in time.
F. Huang, S. Gao, S. Chen, and H. Huang, “Zeroth-order stochastic alternating direction method of multipliers for nonconvex nonsmooth optimization,” in IJCAI , 2019, pp. 2549–2555
2019
Closest in time.
K. Ji, Z. Wang, Y. Zhou, and Y. Liang, “Improved zeroth-order variance reduced algorithms and analysis for nonconvex optimization,” in International Conference on Machine Learning , 2019, pp. 3100–3109
2019
Closest in time.
Z. Wang, K. Ji, Y. Zhou, Y. Liang, and V. Tarokh, “Spiderboost and momentum: Faster variance reduction algorithms,” in Advances in Neural Information Processing Systems , 2019, pp. 2403–2413
2019
Closest in time.
F. Huang, B. Gu, Z. Huo, S. Chen, and H. Huang, “Faster gradient-free proximal stochastic methods for nonconvex nonsmooth optimization,” in AAAI , 2019, pp. 1503–1510
2019
Closest in time.
B. Jiang, T. Lin, S. Ma, and S. Zhang, “Structured nonconvex and nonsmooth optimization: algorithms and iteration complexity analysis,” Computational Optimization and Applications , vol. 72, no. 1, pp. 115–157, 2019
2019
Closest in time.
F. Huang, S. Chen, and H. Huang, “Faster stochastic alternating direction method of multipliers for nonconvex optimization,” in International Conference on Machine Learning , 2019, pp. 2839–2848
2019
Closest in time.
Y. Liu, F. Shang, H. Liu, L. Kong, J. Licheng, and Z. Lin, “Accelerated variance reduction stochastic admm for large-scale machine learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2020
2020
Closest in time.
F. Huang, S. Gao, J. Pei, and H. Huang, “Accelerated zeroth-order and first-order momentum methods from mini to minimax optimization,” The Journal of Machine Learning Research , vol. 23, no. 1, pp. 1616–1685, 2022
2022
Closest in time.