Fetching the paper…
Reading the bibliography…
We present a framework for smooth optimization of explicitly regularized objectives for (structured) sparsity.
An elementary counterexample to the open mapping principle for bilinear maps
Charles Horowitz · 1975
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Functional analysis 2nd ed
Walter Rudin · 1991
Earlier work this paper cites.
A statistical view of some chemometrics regression tools
L. E. Frank and Jerome H Friedman · 1993
Earlier work this paper cites.
Basis pursuit
Shaobing Chen and David Donoho · 1994
Earlier work this paper cites.
Sparse approximate solutions to linear systems
Balas Kausik Natarajan · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Penalized regressions: the bridge versus the lasso
Wenjiang J Fu · 1998
Earlier work this paper cites.
Least absolute shrinkage is equivalent to quadratic penalization
Yves Grandvalet · 1998
Earlier work this paper cites.
Asymptotics for lasso-type estimators
Wenjiang Fu and Keith Knight · 2000
Earlier work this paper cites.
Optimization transfer using surrogate objective functions
Kenneth Lange, David R Hunter, and Ilsoon Yang · 2000
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L Donoho, and Michael A Saunders · 2001
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Jianqing Fan and Runze Li · 2001
Earlier work this paper cites.
Optimally sparse representation in general (nonorthogonal) dictionaries via ℓ 1 \ell_{1} minimization
David L Donoho and Michael Elad · 2003
Earlier work this paper cites.
Variable selection using mm algorithms
David R Hunter and Runze Li · 2005
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie · 2005
Earlier work this paper cites.
High-dimensional graphs and variable selection with the lasso
Nicolai Meinshausen and Peter Bühlmann · 2006
Earlier work this paper cites.
Just relax: Convex programming methods for identifying sparse signals in noise
Joel A Tropp · 2006
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
On model selection consistency of lasso
Peng Zhao and Bin Yu · 2006
Earlier work this paper cites.
Exact reconstruction of sparse signals via nonconvex minimization
Rick Chartrand · 2007
Earlier work this paper cites.
Fast optimization methods for l1 regularization: A comparative study and two new approaches
Mark Schmidt, Glenn Fung, and Rmer Rosales · 2007
Earlier work this paper cites.
Restricted isometry properties and nonconvex compressive sensing
Rick Chartrand and Valentina Staneva · 2008
Earlier work this paper cites.
Iteratively reweighted algorithms for compressive sensing
Rick Chartrand and Wotao Yin · 2008
Earlier work this paper cites.
The sparsity and bias of the lasso selection in high-dimensional linear regression
Cun-Hui Zhang and Jian Huang · 2008
Earlier work this paper cites.
Set-valued analysis
Jean-Pierre Aubin and Hélène Frankowska · 2009
Earlier work this paper cites.
Learning with structured sparsity
Junzhou Huang, Tong Zhang, and Dimitris Metaxas · 2009
Earlier work this paper cites.
Information-theoretic limits on sparsity recovery in the high-dimensional and noisy setting
Martin J Wainwright · 2009
Earlier work this paper cites.
Regularization paths for generalized linear models via coordinate descent
Jerome Friedman, Trevor Hastie, and Rob Tibshirani · 2010
Earlier work this paper cites.
L1/2 regularization
Zongben Xu, Hai Zhang, Yao Wang, XiangYu Chang, and Yong Liang · 2010
Earlier work this paper cites.
Nearly unbiased variable selection under minimax concave penalty
Cun-Hui Zhang · 2010
Earlier work this paper cites.
Coordinate descent algorithms for nonconvex penalized regression, with applications to biological feature selection
Patrick Breheny and Jian Huang · 2011
Earlier work this paper cites.
A note on the complexity of ℓ p \ell_{p} minimization
Dongdong Ge, Xiaoye Jiang, and Yinyu Ye · 2011
Earlier work this paper cites.
Structured variable selection with sparsity-inducing norms
Rodolphe Jenatton, Jean-Yves Audibert, and Francis Bach · 2011
Earlier work this paper cites.
Optimization with sparsity-inducing penalties
Francis Bach, Rodolphe Jenatton, Julien Mairal, Guillaume Obozinski, et al · 2012
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Cited alongside, same era.
Introduction to smooth manifolds, 2012
John Lee · 2012
Cited alongside, same era.
On the minimization of a tikhonov functional with a non-convex sparsity constraint
Ronny Ramlau and Clemens A Zarzer · 2012
Cited alongside, same era.
L1/2 regularization: A thresholding representation theory and a fast solver
Zongben Xu, Xiangyu Chang, Fengmin Xu, and Hai Zhang · 2012
Cited alongside, same era.
Bilinear mappings–selected properties and problems
Marek Balcerzak, Filip Strobin, and Artur Wachowicz · 2013
Cited alongside, same era.
A comparison of typical ℓ p \ell_{p} minimization algorithms
Qin Lyu, Zhouchen Lin, Yiyuan She, and Chao Zhang · 2013
Cited alongside, same era.
Perspective maximum likelihood-type estimation via proximal decomposition
Patrick L Combettes and Christian L Müller · 2020
Later among the works it cites.
Towards optimization on varieties
Eitan Levin · 2020
Later among the works it cites.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Later among the works it cites.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
Mert Pilanci and Tolga Ergen · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Representation costs of linear neural networks: Analysis and design
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regularizers for structured sparsity
Charles A Micchelli, Jean M Morales, and Massimiliano Pontil · 2013
Cited alongside, same era.
A sparse-group lasso
Noah Simon, Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2013
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Cited alongside, same era.
Regularized m-estimators with nonconvexity: statistical and algorithmic theory for local optima
Po-Ling Loh and Martin J Wainwright · 2015
Cited alongside, same era.
Zhen Dai, Mina Karzand, and Nathan Srebro · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
A unified algorithm for the non-convex penalized estimation: The ncpen package
Dongshin Kim, Sangin Lee, and Sunghoon Kwon · 2021
Later among the works it cites.
Implicit sparse regularization: The impact of depth and early stopping
Jiangyuan Li, Thanh Nguyen, Chinmay Hegde, and Ka Wai Wong · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2021
Later among the works it cites.
Smooth bilevel programming for sparse regularization
Clarice Poon and Gabriel Peyré · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant Jayakumar, Razvan Pascanu, Peter Latham, and Yee Teh · 2021
Later among the works it cites.
Equivalences between sparse models and neural networks
Ryan Tibshirani · 2021
Later among the works it cites.
Non-negative least squares via overparametrization
Hung-Hsu Chou, Johannes Maly, and Claudio Mayrink Verdun · 2022
Later among the works it cites.
A critical review of lasso and its derivatives for variable selection under dependence among covariates
Laura Freijeiro-González, Manuel Febrero-Bande, and Wenceslao González-Manteiga · 2022
Later among the works it cites.
Feature learning in l 2 l_{2} -regularized dnns: Attraction/repulsion and sparsity
Arthur Jacot, Eugene Golikov, Clément Hongler, and Franck Gabriel · 2022
Later among the works it cites.
Inductive bias of multi-channel linear convolutional networks with bounded weight norm
Meena Jagadeesan, Ilya Razenshteyn, and Suriya Gunasekar · 2022
Later among the works it cites.
Jianhao Ma and Salar Fattahi · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Learning deep models: Critical points and local openness
Maher Nouiehed and Meisam Razaviyayn · 2022
Later among the works it cites.
Neuronized priors for bayesian sparse linear regression
Minsuk Shin and Jun S Liu · 2022
Later among the works it cites.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
Loucas Pillaud Vivien, Julien Reygner, and Nicolas Flammarion · 2022
Later among the works it cites.
A better way to decay: Proximal gradient training algorithms for neural nets
Liu Yang, Jifan Zhang, Joseph Shenouda, Dimitris Papailiopoulos, Kangwook Lee, and Robert D Nowak · 2022
Later among the works it cites.
High-dimensional linear regression via implicit regularization
Peng Zhao, Yun Yang, and Qiao-Chu He · 2022
Later among the works it cites.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Closest in time.
More is less: inducing sparsity via overparameterization
Hung-Hsu Chou, Johannes Maly, and Holger Rauhut · 2023
Closest in time.
(s) gd over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2023
Closest in time.
Deep learning meets sparse regularization: A signal processing perspective
Rahul Parhi and Robert D Nowak · 2023
Closest in time.
Smooth over-parameterized solvers for non-smooth structured optimization
Clarice Poon and Gabriel Peyré · 2023
Closest in time.
Implicit bias of sgd in ℓ 2 \ell_{2} -regularized linear dnns: One-way jumps from high to low rank
Zihan Wang and Arthur Jacot · 2023
Closest in time.
Symmetry leads to structured constraint of learning
Liu Ziyin · 2023
Closest in time.
spred: Solving l 1 l_{1} penalty with sgd
Liu Ziyin and Zihao Wang · 2023
Closest in time.
Stochastic collapse: How gradient noise attracts sgd dynamics towards simpler subnetworks
Feng Chen, Daniel Kunin, Atsushi Yamamura, and Surya Ganguli · 2024
Closest in time.
The effect of smooth parametrizations on nonconvex optimization landscapes
Eitan Levin, Joe Kileel, and Nicolas Boumal · 2024
Closest in time.
Kurdyka-lojasiewicz exponent via hadamard parametrization
Wenqing Ouyang, Yuncheng Liu, Ting Kei Pong, and Hao Wang · 2024
Closest in time.