Fetching the paper…
Reading the bibliography…
In this paper, we leverage over-parameterization to design regularization-free algorithms for the high-dimensional single index model and provide theoretical guarantees for the induced implicit regularization phenomenon.
A generalization theory of gradient descent for learning over-parameterized deep ReLU networks
Cao, Y · 1902
Earlier work this paper cites.
Two models of double descent for weak features
Belkin, M · 1903
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T · 1903
Earlier work this paper cites.
Zhao, P · 1903
Earlier work this paper cites.
Azizan, N · 1906
Earlier work this paper cites.
A refined primal-dual analysis of the implicit bias
Ji, Z · 1906
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S · 1908
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y · 1910
Earlier work this paper cites.
A model of double descent for high-dimensional binary linear classification
Deng, Z · 1911
Earlier work this paper cites.
Montanari, A · 1911
Earlier work this paper cites.
Exact expressions for double descent and implicit regularization via surrogate random design
Dereziński, M · 1912
Earlier work this paper cites.
Ma, C · 1912
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Stein, C · 1972
Earlier work this paper cites.
A generalized linear model with “Gaussian” regressor variables
Brillinger, D. R · 1982
Earlier work this paper cites.
Non-parametric analysis of a generalized regression model: The maximum rank correlation estimator
Han, A. K · 1987
Earlier work this paper cites.
Regression analysis under link violation
Li, K.-C · 1989
Earlier work this paper cites.
Generalized Linear Models
McCullagh, P · 1989
Earlier work this paper cites.
Slicing regression: A link-free regression method
Duan, N · 1991
Earlier work this paper cites.
Sliced inverse regression for dimension reduction
Li, K.-C · 1991
Earlier work this paper cites.
On principal Hessian directions for data visualization and dimension reduction: Another application of Stein’s lemma
Li, K.-C · 1992
Earlier work this paper cites.
Optimal smoothing in single-index models
Hardle, W · 1993
Earlier work this paper cites.
Generalized partially linear single-index models
Carroll, R. J · 1997
Earlier work this paper cites.
Principal Hessian directions revisited
Cook, R. D · 1998
Earlier work this paper cites.
Dimension reduction in binary response regression
Cook, R. D · 1999
Earlier work this paper cites.
On extended partially linear single-index models
Xia, Y · 1999
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Fan, J · 2001
Earlier work this paper cites.
Analytic study of double descent in binary classification: The impact of loss
Kini, G · 2001
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L · 2002
Earlier work this paper cites.
A distribution-free theory of nonparametric regression
Györfi, L · 2002
Earlier work this paper cites.
Huang, K · 2002
Earlier work this paper cites.
Natural language processing advancements by deep learning: A survey
Torfi, A · 2003
Earlier work this paper cites.
Use of exchangeable pairs in the analysis of simulations
Stein, C · 2004
Earlier work this paper cites.
Sufficient dimension reduction via inverse regression: A minimum discrepancy approach
Cook, R. D · 2005
Earlier work this paper cites.
Compressed sensing
Donoho, D. L · 2006
Earlier work this paper cites.
The restricted isometry property and its implications for compressed sensing
Candés, E. J · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A · 2008
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Tsybakov, A. B · 2008
Earlier work this paper cites.
A multiple-index model and dimension reduction
Xia, Y · 2008
Earlier work this paper cites.
Semiparametric and nonparametric methods in econometrics
Horowitz, J. L · 2009
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Recht, B · 2010
Earlier work this paper cites.
Nearly unbiased variable selection under minimax concave penalty
Zhang, C.-H · 2010
Earlier work this paper cites.
Nuclear-norm penalization and optimal rates for noisy low-rank matrix completion
Koltchinskii, V · 2011
Earlier work this paper cites.
Minimax rates of estimation for high-dimensional linear regression over ℓ q \ell_{q} -balls
Raskutti, G · 2011
Earlier work this paper cites.
Estimation of high-dimensional low-rank matrices
Rohde, A · 2011
Earlier work this paper cites.
Challenging the empirical mean and empirical variance: A deviation study
Catoni, O · 2012
Earlier work this paper cites.
Robust 1-bit compressed sensing and sparse logistic regression: A convex programming approach
Plan, Y · 2012
Cited alongside, same era.
One-bit compressed sensing by linear programming
Plan, Y · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N · 2014
Cited alongside, same era.
Strong oracle optimality of folded concave penalized estimation
Fan, J · 2014
Cited alongside, same era.
Variable selection for general index models via sliced inverse regression
Jiang, B · 2014
Cited alongside, same era.
Empirical risk minimization for heavy-tailed losses
Brownlees, C · 2015
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y · 2018
Later among the works it cites.
Just interpolate: Kernel ”ridgeless” regression can generalize
Liang, T · 2018
Later among the works it cites.
On consistency and sparsity for sliced inverse regression in high dimensions
Lin, Q · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Mei, S · 2018
Later among the works it cites.
Sub-Gaussian estimators of the mean of a random matrix with heavy-tailed entries
Minsker, S · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Phase retrieval via matrix completion
Candés, E. J · 2015
Cited alongside, same era.
Deep learning
LeCun, Y · 2015
Cited alongside, same era.
On model selection consistency of regularized M-estimators
Lee, J. D · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B · 2015
Cited alongside, same era.
Phase retrieval with application to optical imaging: a contemporary overview
Shechtman, Y · 2015
Cited alongside, same era.
Extra: An exact first-order algorithm for decentralized consensus optimization
Shi, W · 2015
Cited alongside, same era.
Oymak, S · 2018
Later among the works it cites.
Rotskoff, G. M · 2018
Later among the works it cites.
Mean field analysis of neural networks
Sirignano, J · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
Later among the works it cites.
A convex formulation for high-dimensional sparse sliced inverse regression
Tan, K. M · 2018
Later among the works it cites.
The generalized Lasso for sub-Gaussian observations with dithered quantization
Thrampoulidis, C · 2018
Later among the works it cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
Vershynin, R · 2018
Later among the works it cites.
Deep learning for computer vision: A brief review
Voulodimos, A · 2018
Later among the works it cites.
Structured recovery with heavy-tailed measurements: A thresholding procedure and optimal rates
Wei, X · 2018
Later among the works it cites.
When will gradient methods converge to max-margin classifier under ReLU models?
Xu, T · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D · 2018
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L · 2019
Later among the works it cites.
Robust covariance estimation for approximate factor models
Fan, J · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gidel, G · 2019
Later among the works it cites.
Non-Gaussian observations in nonlinear compressed sensing via Stein discrepancies
Goldstein, L · 2019
Later among the works it cites.
Communication-efficient distributed statistical inference
Jordan, M. I · 2019
Later among the works it cites.
User-friendly covariance estimation for heavy-tailed distributions
Ke, Y · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 2019
Later among the works it cites.
Sparse sliced inverse regression via Lasso
Lin, Q · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit
Mei, S · 2019
Later among the works it cites.
High-dimensional varying index coefficient models via stein’s identity
Na, S · 2019
Later among the works it cites.
Convergence of gradient descent on separable data
Nacson, M. S · 2019
Later among the works it cites.
Sparse minimum discrepancy approach to sufficient dimension reduction with simultaneous variable selection in ultrahigh dimension
Qian, W · 2019
Later among the works it cites.
Implicit regularization for optimal sparse recovery
Vaškevičius, T · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C · 2019
Later among the works it cites.
A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics
Weinan, E · 2019
Later among the works it cites.
Misspecified nonconvex statistical optimization for sparse phase retrieval
Yang, Z · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Yehudai, G · 2019
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Yun, C · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L · 2020
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K · 2020
Closest in time.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution
Ma, C · 2020
Closest in time.
Robust modifications of U-statistics and applications to covariance estimation problems
Minsker, S · 2020
Closest in time.
Harmless interpolation of noisy data in regression
Muthukumar, V · 2020
Closest in time.
A survey of the usages of deep learning in natural language processing
Otter, D. W · 2020
Closest in time.
Graph-dependent implicit regularisation for distributed stochastic subgradient descent
Richards, D · 2020
Closest in time.
Decentralised learning with distributed gradient descent and random features
Richards, D · 2020
Closest in time.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Li, Z · 2021
Closest in time.
Robust 1-bit compressive sensing via binary stable embeddings of sparse vectors
Jacques, L · 2082
Closest in time.