Fetching the paper…
Reading the bibliography…
High-dimensional statistical learning (HDSL) has wide applications in data analysis, operations research, and decision-making.
Optimal nonlinear approximation
R. A. DeVore, R. Howard, and C. Micchelli · 1989
Earlier work this paper cites.
On affine scaling algorithms for non-convex quadratic programming
Y. Ye · 1992
Earlier work this paper cites.
A statistical view of some chemometrics regression tools
L. Frank and J. Friedman · 1993
Earlier work this paper cites.
Comparison of learning algorithms for handwritten digit recognition
Y. LeCun, L. Jackel, L. Bottou, A. Brunot, C. Cortes, J. Denker, H. Drucker, I. Guyon, U. Muller, E. Sackinger, and P. Simard · 1995
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
H. N. Mhaskar · 1996
Earlier work this paper cites.
On the complexity of approximating a kkt point of quadratic programming
Y. Ye · 1998
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
J. Fan and R. Li · 2001
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. Nesterov · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. Bartlett, M. Jordan, and J. McAuliffe · 2006
Earlier work this paper cites.
Modern statistical estimation via oracle inequalities
E. Candes · 2006
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
Gene selection using support vector machines with non-convex penalty
H. Zhang, J. Ahn, X. Lin, and C. Park · 2006
Earlier work this paper cites.
The adaptive lasso and its oracle properties
H. Zou · 2006
Earlier work this paper cites.
The dantzig selector: Statistical estimation when p is much larger than n
E. Candes and T. Tao · 2007
Earlier work this paper cites.
A robust hybrid of lasso and ridge regression
A. Owen · 2007
Earlier work this paper cites.
Ranking and empirical minimization of u-statistics
S. Clémençon, G. Lugosi, N. Vayatis, et al · 2008
Earlier work this paper cites.
Graph implementations for nonsmooth convex programs
M. C. Grant and S. P. Boyd · 2008
Earlier work this paper cites.
One-step sparse estimation in non-concave penalized likelihood models
H. Zou and R. Li · 2008
Earlier work this paper cites.
Simultaneous analysis of lasso and dantzig selector
P. Bickel, Y. Ritov, and A. Tsybakov · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
A. Rahimi and B. Recht · 2009
Earlier work this paper cites.
On the conditions used to prove oracle results for the lasso
S. A. van de Geer, P. Bühlmann, et al · 2009
Earlier work this paper cites.
Lower bound theory of nonzero entries in solutions of 2-p minimization
X. Chen, F. Xu, and Y. Ye · 2010
Earlier work this paper cites.
Safe feature elimination for the lasso and sparse supervised learning problems
L. E. Ghaoui, V. Viallon, and T. Rabbani · 2010
Earlier work this paper cites.
Rademacher complexities and bounding the excess risk in active learning
V. Koltchinskii · 2010
Earlier work this paper cites.
Nearly unbiased variable selection under minimax concave penalty
C. Zhang · 2010
Earlier work this paper cites.
ℓ \ell 1-penalized quantile regression in high-dimensional sparse models
A. Belloni and V. Chernozhukov · 2011
Earlier work this paper cites.
Statistics for high-dimensional data: methods theory and applications
P. Bühlmann and S. van de Geer · 2011
Earlier work this paper cites.
Non-concave penalty likelihood with np-dimensionality
J. Fan and J. Lv · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Minimax rates of estimation for high-dimensional linear regression over ℓ q \ell_{q} -balls
G. Raskutti, M. J. Wainwright, and B. Yu · 2011
Earlier work this paper cites.
Regression shrinkage and selection via the lasso: a retrospective
R. Tibshirani · 2011
Cited alongside, same era.
A unified framework for high-dimensional analysis of m m -estimators with decomposable regularizers
S. N. Negahban, P. Ravikumar, M. J. Wainwright, B. Yu, et al · 2012
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices. Chapter 5 of: Compressed Sensing, Theory and Applications
R. Vershynin · 2012
Cited alongside, same era.
A general theory of concave regularization for high dimensional sparse estimation problems
C. Zhang and T. Zhang · 2012
Cited alongside, same era.
Cvx: Matlab software for disciplined convex programming, version 2.0 beta, 2013
M. Grant and S. Boyd · 2013
Cited alongside, same era.
The mnist database of handwritten digits
Sgd learns the conjugate kernel class of the network
A. Daniely · 2017
Later among the works it cites.
Improved regularization of convolutional neural networks with cutout
T. DeVries and G. W. Taylor · 2017
Later among the works it cites.
X. Gastaldi · 2017
Later among the works it cites.
Global optimality in neural network training
B. Haeffele and R. Vidal · 2017
Later among the works it cites.
Optimality condition and complexity analysis for linearly-constrained optimization without differentiability on the boundary
G. Haeser, H. Liu, and Y. Ye · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. LeCun, C. Cortes, and C. Burges · 2013
Cited alongside, same era.
On constrained and regularized high-dimensional regression
X. Shen, W. Pan, Y. Zhu, and H. Zhou · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Cited alongside, same era.
The l1 penalized lad estimator for high dimensional linear regression
L. Wang · 2013
Cited alongside, same era.
Calibrating nonconvex penalized regression in ultra-high dimension
L. Wang, Y. Kim, and R. Li · 2013
Cited alongside, same era.
Strong oracle optimality of folded concave penalized estimation
J. Fan, L. Xue, and H. Zou · 2014
Cited alongside, same era.
Lectures on stochastic programming: modeling and theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2014
Cited alongside, same era.
Folded concave penalized sparse linear regression: sparsity, statistical performance, and algorithmic theories on local solutions
H. Liu, T. Yao, R. Li, and Y. Ye · 2017
Later among the works it cites.
Statistical consistency and asymptotic normality for high-dimensional robust m -estimators
P.-L. Loh · 2017
Later among the works it cites.
Learning sparse neural networks through l 0 l_{0} regularization
C. Louizos, M. Welling, and D. P. Kingma · 2017
Later among the works it cites.
Gap safe screening rules for sparsity enforcing penalties
E. Ndiaye, O. Fercoq, A. Gramfort, and J. Salmon · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Group sparse regularization for deep neural networks
S. Scardapane, D. Comminiello, A. Hussain, and A. Uncini · 2017
Later among the works it cites.
Error bounds for approximations with deep relu networks
D. Yarotsky · 2017
Later among the works it cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Later among the works it cites.
Pruned and Structurally Sparse Neural Networks
S. Alford, R. Robinett, L. Milechin, and J. Kepner · 2018
Later among the works it cites.
Approximation and estimation for high-dimensional deep learning networks
A. R. Barron and J. M. Klusowski · 2018
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2018
Later among the works it cites.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
X. Li, J. Lu, Z. Wang, J. Haupt, and T. Zhao · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Later among the works it cites.
Adding one neuron can eliminate all bad local minima
S. Liang, R. Sun, J. Lee, and R. Srikant · 2018
Later among the works it cites.
Sample average approximation with sparsity-inducing penalty for highdimensional stochastic programming
H. Liu, X. Wang, T. Yao, R. Li, and Y. Ye · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
R. Vershynin · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2019
Closest in time.
Towards a regularity theory for relu networks–chain rule and global error estimates
J. Berner, D. Elbrächter, P. Grohs, and A. Jentzen · 2019
Closest in time.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Closest in time.
Sparse networks from scratch: Faster training without losing performance
T. Dettmers and L. Zettlemoyer · 2019
Closest in time.
Generalization error in deep learning
D. Jakubovitz, R. Giryes, and M. R. Rodrigues · 2019
Closest in time.
Cifar-zoo: Pytorch implementation of cnns for cifar dataset
W. Li · 2019
Closest in time.
Training neural networks with local error signals
A. Nøkland and L. H. Eidnes · 2019
Closest in time.
Optimization for deep learning: theory and algorithms
R. Sun · 2019
Closest in time.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Y. Cao and Q. Gu · 2020
Closest in time.
Understanding and enhancing mixed sample data augmentation
E. Harris, A. Marcu, M. Painter, M. Niranjan, A. Prügel-Bennett, and J. Hare · 2020
Closest in time.