Fetching the paper…
Reading the bibliography…
In this paper, a novel stochastic extra-step quasi-Newton method is developed to solve a class of nonsmooth nonconvex composite optimization problems.
Nguyen, L.M., van Dijk, M., Phan, D.T., Nguyen, P.H., Weng, T.W., Kalagnanam, J.R.: Finite-sum smooth optimization with SARAH (2019) · 1901
Earlier work this paper cites.
In: Adv. in Neural Inf. Process. Syst., pp. 3925–3936 (2018)
Zhou, D., Xu, P., Gu, Q.: Stochastic nested variance reduction for nonconvex optimization · 1910
Earlier work this paper cites.
Ann. Math. Stat. 22
Robbins, H., Monro, S.: A stochastic approximation method · 1951
Earlier work this paper cites.
Bull. Soc. Math. FR. 93
Moreau, J.J.: Proximité et dualité dans un espace hilbertien · 1965
Earlier work this paper cites.
Optimizing Methods in Statistics (Ohio State Univ., 1971) pp. 233–257 (1971)
Robbins, H., Siegmund, D.: A convergence theorem for non negative almost supermartingales and some applications · 1971
Earlier work this paper cites.
Math. Comp. 35
Nocedal, J.: Updating quasi-Newton matrices with limited storage · 1980
Earlier work this paper cites.
Math. Program. 45
Liu, D.C., Nocedal, J.: On the limited memory BFGS method for large scale optimization · 1989
Earlier work this paper cites.
Ann. Oper. Res. 46/47
Luo, Z.Q., Tseng, P.: Error bounds and convergence analysis of feasible descent methods: a general approach · 1993
Earlier work this paper cites.
SIAM J. Optim. 3
Pang, J.S., Qi, L.: Nonsmooth equations: motivation and algorithms · 1993
Earlier work this paper cites.
Math. Oper. Res. 18
Qi, L.: Convergence analysis of some algorithms for solving nonsmooth equations · 1993
Earlier work this paper cites.
Math. Program. 58
Qi, L., Sun, J.: A nonsmooth version of Newton’s method · 1993
Earlier work this paper cites.
Comput. Optim. Appl. 3
Chen, X., Qi, L.: A parameterized Newton method and a quasi-Newton method for nonsmooth equations · 1994
Earlier work this paper cites.
Oper. Res. Lett. 20
Qi, L.: On superlinear convergence of quasi-Newton methods for nonsmooth equations · 1997
Earlier work this paper cites.
SIAM J. Optim. 7
Sun, D., Han, J.: Newton and quasi-Newton methods for a class of nonsmooth equations and related problems · 1997
Earlier work this paper cites.
In: Proceedings of the 12th Int. Conf. on Neural Inf. Process. Syst., pp. 512–518 (1999)
Mason, L., Baxter, J., Bartlett, P., Frean, M.: Boosting algorithms as gradient descent in function space · 1999
Earlier work this paper cites.
MPS/SIAM Series on Optim. SIAM; MPS, Philadelphia, PA (2000)
Conn, A.R., Gould, N.I.M., Toint, P.L.: Trust-region methods · 2000
Earlier work this paper cites.
Springer Series in Statistics. Springer-Verlag, New York (2001)
Hastie, T., Tibshirani, R., Friedman, J.: The elements of statistical learning · 2001
Earlier work this paper cites.
Multiscale Model. Simul. 4
Combettes, P.L., Wajs, V.R.: Signal recovery by proximal forward-backward splitting · 2005
Earlier work this paper cites.
Information Science and Statistics. Springer, New York (2006)
Bishop, C.M.: Pattern recognition and machine learning · 2006
Earlier work this paper cites.
In: Proceedings of the 24th Int. Conf. on Mach. Learn., pp. 33–40 (2007)
Andrew, G., Gao, J.: Scalable training of ℓ 1 \ell_{1} -regularized log-linear models · 2007
Earlier work this paper cites.
In: Proceedings of the 11th Int. Conf. on Artif. Intelligence and Stat., pp. 436–443 (2007)
Schraudolph, N.N., Yu, J., Günter, S.: A stochastic quasi-Newton method for online convex optimization · 2007
Earlier work this paper cites.
J. Mach. Learn. Res. 9
Fan, R.E., Chang, K.W., Hsieh, C.J., Wang, X.R., Lin, C.J.: LIBLINEAR: A library for large linear classification · 2008
Earlier work this paper cites.
Found. Comput. Math. 9
Candès, E.J., Recht, B.: Exact matrix completion via convex optimization · 2009
Earlier work this paper cites.
In: 27th Ann. Allerton Conf. on Comm., Control and Computing, vol. 42, pp. 1493–1498 (2009)
Chandrasekaran, V., Sanghavi, S., Parrilo, P.A., Willsky, A.S.: Sparse and low-rank matrix decompositions · 2009
Earlier work this paper cites.
Appl. Math. Lett. 22
Dong, Y.: An extension of Luque’s growth condition · 2009
Earlier work this paper cites.
In: Proceedings of the 26th Int. Conf. on Mach. Learn., pp. 689–696 (2009)
Mairal, J., Bach, F., Ponce, J., Sapiro, G.: Online dictionary learning for sparse coding · 2009
Earlier work this paper cites.
In: Proceedings of the 27th Int. Conf. on Mach. Learn., vol. 27, pp. 735–742 (2010)
Martens, J.: Deep learning via Hessian-free optimization · 2010
Earlier work this paper cites.
J. Mach. Learn. Res. 11
Shi, J., Yin, W., Osher, S., Sajda, P.: A fast hybrid algorithm for large-scale ℓ 1 \ell_{1} -regularized logistic regression · 2010
Earlier work this paper cites.
SIAM J. Sci. Comput. 32
Wen, Z., Yin, W., Goldfarb, D., Zhang, Y.: A fast algorithm for sparse reconstruction based on shrinkage, subspace optimization, and continuation · 2010
Earlier work this paper cites.
Found. and Trends® in Mach. Learn. 4
Bach, F., Jenatton, R., Mairal, J., Obozinski, G.: Optimization with sparsity-inducing penalties · 2011
Earlier work this paper cites.
CMS Books in Mathematics/Ouvrages de Mathématiques de la SMC. Springer, New York (2011)
Bauschke, H.H., Combettes, P.L.: Convex analysis and monotone operator theory in Hilbert spaces · 2011
Earlier work this paper cites.
SIAM J. Optim. 21
Byrd, R.H., Chin, G.M., Neveitt, W., Nocedal, J.: On the use of stochastic Hessian information in optimization methods for machine learning · 2011
Earlier work this paper cites.
ACM Trans. on Intelligent Syst. and Technology (TIST) 2
Chang, C.C., Lin, C.J.: LIBSVM: a library for support vector machines · 2011
Earlier work this paper cites.
In: Fixed-point algorithms for inverse problems in science and engineering, Springer Optim. Appl. , vol. 49, pp. 185–212. Springer, New York (2011)
Combettes, P.L., Pesquet, J.C.: Proximal splitting methods in signal processing · 2011
Earlier work this paper cites.
J. Mach. Learn. Res. 12
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization · 2011
Earlier work this paper cites.
J. Mach. Learn. Res. 12
Shalev-Shwartz, S., Tewari, A.: Stochastic methods for ℓ 1 \ell_{1} -regularized loss minimization · 2011
Earlier work this paper cites.
J. Global Optim. 50
Wang, X., Ma, C., Li, M.: A globally and superlinearly convergent quasi-Newton method for general box constrained variational inequalities without smoothing approximation · 2011
Earlier work this paper cites.
J. Mach. Learn. Res. 13
Yuan, G.X., Ho, C.H., Lin, C.J.: An improved GLMNET for ℓ 1 \ell_{1} -regularized logistic regression · 2012
Earlier work this paper cites.
SIAM J. Optim. 23
Ghadimi, S., Lan, G.: Stochastic first- and zeroth-order methods for nonconvex stochastic programming · 2013
Earlier work this paper cites.
In: Adv. in Neural Inf. Process. Syst., pp. 315–323 (2013)
Johnson, R., Zhang, T.: Accelerating stochastic gradient descent using predictive variance reduction · 2013
Earlier work this paper cites.
Math. Program. 140
Nesterov, Y.: Gradient methods for minimizing composite functions · 2013
Earlier work this paper cites.
J. Mach. Learn. Res. 14
Shalev-Shwartz, S., Zhang, T.: Stochastic dual coordinate ascent methods for regularized loss minimization · 2013
Cited alongside, same era.
In: Proceedings of the 30th Int. Conf. on Mach. Learn., pp. 1139–1147 (2013)
Sutskever, I., Martens, J., Dahl, G., Hinton, G.: On the importance of initialization and momentum in deep learning · 2013
Cited alongside, same era.
Springer Science & Business Media (2013)
Vapnik, V.: The nature of statistical learning theory · 2013
Cited alongside, same era.
In: Adv. in Neural Inf. Process. Syst., pp. 1646–1654 (2014)
Defazio, A., Bach, F., Lacoste-Julien, S.: SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives · 2014
Cited alongside, same era.
Found. and Trends® Signal Process. 7
Deng, L., Yu, D.: Deep learning: Methods and applications · 2014
Cited alongside, same era.
J. Mach. Learn. Res. 15
Hsieh, C.J., Sustik, M.A., Dhillon, I.S., Ravikumar, P.: QUIC: quadratic approximation for sparse inverse covariance estimation · 2014
J. Mach. Learn. Res. 18
Agarwal, N., Bullins, B., Hazan, E.: Second-order stochastic optimization for machine learning in linear time · 2017
Later among the works it cites.
Akiba, T., Suzuki, S., Fukuda, K.: Extremely large minibatch SGD: training Resnet-50 on ImageNet in 15 minutes (2017) · 2017
Later among the works it cites.
In: Proceedings of the 49th Annual ACM SIGACT Symp. on Theory of Computing, pp. 1200–1205 (2017)
Allen-Zhu, Z.: Katyusha: The first direct acceleration of stochastic gradient methods · 2017
Later among the works it cites.
Berahas, A.S., Bollapragada, R., Nocedal, J.: An investigation of Newton-sketch and subsampled Newton methods (2017) · 2017
Later among the works it cites.
In: Proceedings of the 34th Int. Conf. on Mach. Learn., pp. 557–565 (2017)
Botev, A., Ritter, H., Barber, D.: Practical Gauss-Newton optimization for deep learning · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization · 2014
Cited alongside, same era.
SIAM J. Optim. 24
Lee, J.D., Sun, Y., Saunders, M.A.: Proximal Newton-type methods for minimizing composite functions · 2014
Cited alongside, same era.
IEEE Trans. Signal Process. 62
Mokhtari, A., Ribeiro, A.: RES: regularized stochastic BFGS algorithm · 2014
Cited alongside, same era.
Found. and Trends® in Optim. 1
Parikh, N., Boyd, S.: Proximal algorithms · 2014
Cited alongside, same era.
Patrinos, P., Stella, L., Bemporad, A.: Forward-backward truncated Newton methods for convex composite optimization (2014) · 2014
Cited alongside, same era.
Cambridge University Press (2014)
Shalev-Shwartz, S., Ben-David, S.: Understanding machine learning: From theory to algorithms · 2014
Cited alongside, same era.
Later among the works it cites.
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K.: Accurate, large minibatch SGD: Training ImageNet in 1 hour (2017) · 2017
Later among the works it cites.
SIAM J. Optim. 27
Iusem, A.N., Jofré, A., Oliveira, R.I., Thompson, P.: Extragradient method with variance reduction for stochastic variational inequalities · 2017
Later among the works it cites.
In: Proceedings of the 34th Int. Conf. on Mach. Learn., vol. 70, pp. 1895–1904 (2017)
Kohler, J.M., Lucchi, A.: Sub-sampled cubic regularization for non-convex optimization · 2017
Later among the works it cites.
In: Adv. in Neural Inf. Process. Syst., pp. 2348–2358 (2017)
Lei, L., Ju, C., Chen, J., Jordan, M.I.: Non-convex finite-sum optimization via SCSG methods · 2017
Later among the works it cites.
In: Proceedings of the 34th Int. Conf. on Mach. Learn., pp. 2613–2621 (2017)
Nguyen, L.M., Liu, J., Scheinberg, K., Takáč, M.: SARAH: A novel method for machine learning problems using stochastic recursive gradient · 2017
Later among the works it cites.
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A.: Automatic differentiation in pytorch (2017)
2017
Later among the works it cites.
SIAM J. Optim. 27
Pilanci, M., Wainwright, M.J.: Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence · 2017
Later among the works it cites.
Math. Program. 162
Schmidt, M., Le Roux, N., Bach, F.: Minimizing finite sums with the stochastic average gradient · 2017
Later among the works it cites.
Comput. Optim. Appl. 67
Stella, L., Themelis, A., Patrinos, P.: Forward–backward quasi-Newton methods for nonsmooth optimization problems · 2017
Later among the works it cites.
SIAM J. Optim. 27
Wang, X., Ma, S., Goldfarb, D., Liu, W.: Stochastic Quasi-Newton Methods for Nonconvex Stochastic Optimization · 2017
Later among the works it cites.
In: Proceedings of the 34th Int. Conf. on Mach. Learn., vol. 70, pp. 3931–3939 (2017)
Ye, H., Luo, L., Zhang, Z.: Approximate Newton methods and their local convergence · 2017
Later among the works it cites.
IEEE Trans. on Signal Process. (2017)
Zhao, R., Haskell, W.B., Tan, V.Y.: Stochastic L-BFGS: Improved convergence rates and practical acceleration strategies · 2017
Later among the works it cites.
IMA J. Numer. Anal. pp. 1–34 (2018)
Bollapragada, R., Byrd, R., Nocedal, J.: Exact and inexact subsampled Newton methods for optimization · 2018
Later among the works it cites.
SIAM Review 60
Bottou, L., Curtis, F.E., Nocedal, J.: Optimization methods for large-scale machine learning · 2018
Later among the works it cites.
Davis, D., Drusvyatskiy, D.: Stochastic model-based minimization of weakly convex functions (2018) · 2018
Later among the works it cites.
Davis, D., Drusvyatskiy, D.: Stochastic subgradient method converges at the rate O ( k − 1 / 4 ) (k^{-1/4}) · 2018
Later among the works it cites.
Math. Oper. Res. 43
Drusvyatskiy, D., Lewis, A.S.: Error bounds, quadratic growth, and linear convergence of proximal methods · 2018
Later among the works it cites.
In: Adv. in Neural Inf. Process. Syst., pp. 689–699 (2018)
Fang, C., Li, C.J., Lin, Z., Zhang, T.: Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator · 2018
Later among the works it cites.
Math. Program. (2018)
Liu, H., So, A.M.C., Wu, W.: Quadratic optimization with orthogonality constraint: explicit Łojasiewicz exponent and linear convergence of retraction-based line-search and stochastic variance-reduced gradient methods · 2018
Later among the works it cites.
In: Proceedings of the 35th Int. Conf. on Mach. Learn., pp. 3185–3194 (2018)
Liu, X., Hsieh, C.J.: Fast variance reduction method with stochastic batch size · 2018
Later among the works it cites.
https://imsc.uni-graz.at/mannel/sqn1.pdf
Mannel, F., Rund, A.: A hybrid semismooth quasi-Newton method for structured nonsmooth operator equations in Banach spaces (2018) · 2018
Later among the works it cites.
Milzarek, A., Xiao, X., Cen, S., Wen, Z., Ulbrich, M.: A stochastic semismooth Newton method for nonsmooth nonconvex optimization (2018) · 2018
Later among the works it cites.
SIAM J. Optim. 28
Mokhtari, A., Eisen, M., Ribeiro, A.: IQN: An incremental quasi-Newton method with local superlinear convergence rate · 2018
Later among the works it cites.
J. Optim. Theory Appl. 176
Nguyen, T.P., Pauwels, E., Richard, E., Suter, B.W.: Extragradient method in optimization: convergence and complexity · 2018
Later among the works it cites.
Osawa, K., Tsuji, Y., Ueno, Y., Naruse, A., Yokota, R., Matsuoka, S.: Large-scale distributed second-order optimization using Kronecker-factored approximate curvature for deep convolutional neural networks (2018) · 2018
Later among the works it cites.
In: Proceedings of the 35th Int. Conf. on Mach. Learn., vol. 80, pp. 4124–4132 (2018)
Poon, C., Liang, J., Schoenlieb, C.: Local convergence properties of SAGA/Prox-SVRG and acceleration · 2018
Later among the works it cites.
Math. Program. pp. 1–34 (2018)
Roosta-Khorasani, F., Mahoney, M.W.: Sub-sampled Newton methods · 2018
Later among the works it cites.
SIAM J. Optim. 28
Themelis, A., Stella, L., Patrinos, P.: Forward-backward envelope for the sum of two nonconvex functions: further properties and nonmonotone linesearch algorithms · 2018
Later among the works it cites.
Wang, Z., Ji, K., Zhou, Y., Liang, Y., Tarokh, V.: Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization · 2018
Later among the works it cites.
J. Sci. Comput. 76
Xiao, X., Li, Y., Wen, Z., Zhang, L.: A regularized semi-smooth Newton method with projection steps for composite convex programs · 2018
Later among the works it cites.
In: Proceedings of the 47th Int. Conf. on Parallel Process., pp. 1–10 (2018)
You, Y., Zhang, Z., Hsieh, C.J., Demmel, J., Keutzer, K.: ImageNet training in minutes · 2018
Later among the works it cites.
Pham, N.H., Nguyen, L.M., Phan, D.T., Tran-Dinh, Q.: ProxSARAH: An efficient algorithmic framework for stochastic composite nonconvex optimization · 2019
Closest in time.
J. Mach. Learn. Res. 20
Wang, J., Zhang, T.: Utilizing second order information in minibatch stochastic variance reduced proximal iterations · 2019
Closest in time.
Optim. Methods Softw. pp. 922–948 (2019)
Wang, X., Wang, X., Yuan, Y.x.: Stochastic proximal quasi-Newton methods for non-convex composite optimization · 2019
Closest in time.
Math. Program. (2019)
Xu, P., Roosta, F., Mahoney, M.W.: Newton-type methods for non-convex optimization under inexact Hessian information · 2019
Closest in time.