Fetching the paper…
Reading the bibliography…
Optimizing with group sparsity is significant in enhancing model interpretability in machining learning applications, e.g., feature selection, compressed sensing and model compression.
Probability and computing-randomized algorithms and probabilistic analysis
Michael Mitzenmacher · 2005
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
The group-lasso for generalized linear models: uniqueness of solutions and efficient algorithms
Volker Roth and Bernd Fischer · 2008
Earlier work this paper cites.
Efficient online and batch learning using forward backward splitting
John Duchi and Yoram Singer · 2009
Earlier work this paper cites.
Learning with dynamic group sparsity
Junzhou Huang, Xiaolei Huang, and Dimitris Metaxas · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Yurii Nesterov · 2009
Earlier work this paper cites.
The benefit of group sparsity
Junzhou Huang, Tong Zhang, et al · 2010
Earlier work this paper cites.
Structured sparse principal component analysis
Rodolphe Jenatton, Guillaume Obozinski, and Francis Bach · 2010
Earlier work this paper cites.
Dual averaging methods for regularized stochastic learning and online optimization
Lin Xiao · 2010
Earlier work this paper cites.
Online learning for group lasso
Haiqin Yang, Zenglin Xu, Irwin King, and Michael R Lyu · 2010
Earlier work this paper cites.
Libsvm: Data repository
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Learning with structured sparsity
Junzhou Huang, Tong Zhang, and Dimitris Metaxas · 2011
Earlier work this paper cites.
Structured sparsity through convex optimization
Francis Bach, Rodolphe Jenatton, Julien Mairal, Guillaume Obozinski, et al · 2012
Earlier work this paper cites.
See all by looking at a few: Sparse modeling for finding representative objects
Ehsan Elhamifar, Guillermo Sapiro, and Rene Vidal · 2012
Earlier work this paper cites.
Manifold identification in dual averaging for regularized stochastic online learning
Sangkyun Lee and Stephen J Wright · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Group-sparse signal denoising: non-convex regularization, convex optimization
Po-Yu Chen and Ivan W Selesnick · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.
A Fast Reduced-Space Algorithmic Framework for Sparse Optimization
Tianyi Chen · 2018
Later among the works it cites.
Farsa for ℓ 1 \ell_{1} -regularized convex optimization: local convergence and numerical experience
Tianyi Chen, Frank E Curtis, and Daniel P Robinson · 2018
Later among the works it cites.
Error bounds, quadratic growth, and linear convergence of proximal methods
Dmitriy Drusvyatskiy and Adrian S Lewis · 2018
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Cited alongside, same era.
Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization
Sashank J Reddi, Suvrit Sra, Barnabas Poczos, and Alexander J Smola · 2016
Cited alongside, same era.
A reduced-space algorithm for minimizing ℓ 1 \ell_{1} -regularized convex functions
Tianyi Chen, Frank E Curtis, and Daniel P Robinson · 2017
Cited alongside, same era.
Convergence theorems for gradient descent
Robert M. Gower · 2018
Later among the works it cites.
Combinatorial penalties: Which structures are preserved by convex relaxations?
Marwa El Halabi, Francis Bach, and Volkan Cevher · 2018
Later among the works it cites.
A simple proximal stochastic gradient method for nonsmooth nonconvex optimization
Zhize Li and Jian Li · 2018
Later among the works it cites.
Mri reconstruction via enhanced group sparsity and nonconvex regularization
Shujun Liu, Jianxin Cao, Hongqing Liu, Xichuan Zhou, Kui Zhang, and Zhengzhou Li · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
On the ineffectiveness of variance reduced optimization for deep learning
Aaron Defazio and Léon Bottou · 2019
Later among the works it cites.
“active-set complexity” of proximal gradient: How long does it take to find the sparsity pattern?
Julie Nutini, Mark Schmidt, and Warren Hare · 2019
Later among the works it cites.
Convergence of stochastic proximal gradient algorithm
Lorenzo Rosasco, Silvia Villa, and Bang Cng Vu · 2019
Later among the works it cites.
A stochastic extra-step quasi-newton method for nonsmooth nonconvex optimization
Minghan Yang, Andre Milzarek, Zaiwen Wen, and Tong Zhang · 2019
Later among the works it cites.
Multi-level composite stochastic optimization via nested variance reduction
Junyu Zhang and Lin Xiao · 2019
Later among the works it cites.
Orthant based proximal stochastic gradient method for ℓ 1 \ell_{1} -regularized optimization
Tianyi Chen, Tianyu Ding, Bo Ji, Guanyi Wang, Yixin Shi, Sheng Yi, Xiao Tu, and Zhihui Zhu · 2020
Closest in time.
Statistical adaptive stochastic gradient methods
Pengchuan Zhang, Hunter Lang, Qiang Liu, and Lin Xiao · 2020
Closest in time.