Fetching the paper…
Reading the bibliography…
We study the training of regularized neural networks where the regularizer can be non-smooth and non-convex.
On the stability of inverse problems
A. N. Tychonoff · 1943
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Penalized regressions: the bridge versus the lasso
Wenjiang J Fu · 1998
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
High-dimensional covariance estimation by minimizing ℓ 1 \ell_{1} -penalized log-determinant divergence
P. Ravikumar, M. J. Wainwright, G. Raskutti, and B. Yu · 2011
Earlier work this paper cites.
Bridge regression: adaptivity and group selection
Cheolwoo Park and Young Joo Yoon · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Proximal newton-type methods for convex optimization
Jason D Lee, Yuekai Sun, and Michael Saunders · 2012
Earlier work this paper cites.
Fast image deconvolution using closed-form thresholding formulas of lq (q= 12, 23) regularization
Wenfei Cao, Jian Sun, and Zongben Xu · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Cited alongside, same era.
Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization
Sashank J Reddi, Suvrit Sra, Barnabas Poczos, and Alexander J Smola · 2016
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Natasha: Faster non-convex stochastic optimization via strongly non-convex parameter
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Sparse+ group-sparse dirty models: Statistical guarantees without unreasonable conditions and a case for non-convexity
Proxsarah: An efficient algorithmic framework for stochastic composite nonconvex optimization
Nhan H Pham, Lam M Nguyen, Dzung T Phan, and Quoc Tran-Dinh · 2019
Later among the works it cites.
Stochastic optimization for DC functions and non-smooth non-convex regularizers with non-asymptotic convergence
Yi Xu, Qi Qi, Qihang Lin, Rong Jin, and Tianbao Yang · 2019
Later among the works it cites.
Non-asymptotic analysis of stochastic methods for non-smooth non-convex regularized problems
Yi Xu, Rong Jin, and Tianbao Yang · 2019
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
Damek Davis and Dmitriy Drusvyatskiy · 2019
Later among the works it cites.
On quasi-newton forward-backward splitting: Proximal calculus and convergence
Stephen Becker, Jalal Fadili, and Peter Ochs · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eunho Yang and Aurélie C Lozano · 2017
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Learning sparse neural networks through l 0 l_{0} regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2018
Cited alongside, same era.
Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Proxquant: Quantized neural networks via proximal operators
Yu Bai, Yu-Xiang Wang, and Edo Liberty · 2019
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Trimming the ℓ 1 \ell_{1} regularizer: Statistical analysis, optimization, and applications to deep learning
Jihun Yun, Peng Zheng, Eunho Yang, Aurelie Lozano, and Aleksandr Aravkin · 2019
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Later among the works it cites.
Stochastic gradient methods with block diagonal matrix adaptation
Jihun Yun, Aurelie C. Lozano, and Eunho Yang · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2019
Later among the works it cites.
Orthant based proximal stochastic gradient method for ℓ _ 1 \ell\_1 -regularized optimization
Tianyi Chen, Tianyu Ding, Bo Ji, Guanyi Wang, Yixin Shi, Sheng Yi, Xiao Tu, and Zhihui Zhu · 2020
Closest in time.
Proxsgd: Training structured neural networks under regularization and constraints
Yang Yang, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud JG van Sloun, Lei Lei, and Symeon Chatzinotas · 2020
Closest in time.