Fetching the paper…
Reading the bibliography…
A continuing mystery in understanding the empirical success of deep neural networks is their ability to achieve zero training error and generalize well, even when the training data is noisy and there are more parameters than data points.
D. L. Hanson and F. T. Wright, “A bound on tail probabilities for quadratic forms in independent random variables,” The Annals of Mathematical Statistics , vol. 42, no. 3, pp. 1079–1083, 1971
1971
Earlier work this paper cites.
A. Edelman, “Eigenvalues and condition numbers of random matrices,” SIAM Journal on Matrix Analysis and Applications , vol. 9, no. 4, pp. 543–560, 1988
1988
Earlier work this paper cites.
S. J. Szarek, “Condition numbers of random matrices,” Journal of Complexity , vol. 7, no. 2, pp. 131–149, 1991
1991
Earlier work this paper cites.
Y. C. Pati, R. Rezaiifar, and P. S. Krishnaprasad, “Orthogonal matching pursuit: Recursive function approximation with applications to wavelet decomposition,” in Signals, Systems and Computers, 1993. 1993 Conference Record of The Twenty-Seventh Asilomar Conference on . IEEE, 1993, pp. 40–44
1993
Earlier work this paper cites.
C.-K. Li and W. So, “Isometries of lp-norm,” The American Mathematical Monthly , vol. 101, no. 5, pp. 452–453, 1994. [Online]. Available: https://doi.org/10.1080/00029890.1994.11996972
1994
Earlier work this paper cites.
D. Bertsimas and J. Tsitsiklis, Introduction to Linear Optimization , 1st ed. Athena Scientific, 1997
1997
Earlier work this paper cites.
R. E. Schapire, Y. Freund, P. Bartlett, W. S. Lee et al. , “Boosting the margin: A new explanation for the effectiveness of voting methods,” The annals of statistics , vol. 26, no. 5, pp. 1651–1686, 1998
1998
Earlier work this paper cites.
V. N. Vapnik, “An overview of statistical learning theory,” IEEE transactions on Neural Networks , vol. 10, no. 5, pp. 988–999, 1999
1999
Earlier work this paper cites.
S. S. Chen, D. L. Donoho, and M. A. Saunders, “Atomic decomposition by basis pursuit,” SIAM review , vol. 43, no. 1, pp. 129–159, 2001
2001
Earlier work this paper cites.
K. R. Davidson and S. J. Szarek, “Local operator theory, random matrices and Banach spaces,” Handbook of the geometry of Banach spaces , vol. 1, no. 317-366, p. 131, 2001
2001
Earlier work this paper cites.
A. Miller, Subset selection in regression . Chapman and Hall/CRC, 2002
2002
Earlier work this paper cites.
P. L. Bartlett, O. Bousquet, S. Mendelson et al. , “Local rademacher complexities,” The Annals of Statistics , vol. 33, no. 4, pp. 1497–1537, 2005
2005
Earlier work this paper cites.
J. A. Tropp and A. C. Gilbert, “Signal recovery from random measurements via orthogonal matching pursuit,” IEEE Transactions on information theory , vol. 53, no. 12, pp. 4655–4666, 2007
2007
Earlier work this paper cites.
A. Rahimi and B. Recht, “Random features for large-scale kernel machines,” in Advances in neural information processing systems , 2008, pp. 1177–1184
2008
Earlier work this paper cites.
M. Rudelson and R. Vershynin, “The Littlewood–Offord problem and invertibility of random matrices,” Advances in Mathematics , vol. 218, no. 2, pp. 600–633, 2008
2008
Earlier work this paper cites.
M. J. Wainwright, “Information-theoretic limits on sparsity recovery in the high-dimensional and noisy setting,” IEEE Transactions on Information Theory , vol. 55, no. 12, pp. 5728–5741, 2009
2009
Earlier work this paper cites.
P. J. Bickel, Y. Ritov, A. B. Tsybakov et al. , “Simultaneous analysis of Lasso and Dantzig selector,” The Annals of Statistics , vol. 37, no. 4, pp. 1705–1732, 2009
2009
Earlier work this paper cites.
A. K. Fletcher, S. Rangan, and V. K. Goyal, “Necessary and sufficient conditions for sparsity pattern recovery,” IEEE Transactions on Information Theory , vol. 55, no. 12, pp. 5758–5772, 2009
2009
Earlier work this paper cites.
S. Aeron, V. Saligrama, and M. Zhao, “Information theoretic bounds for compressed sensing,” IEEE Transactions on Information Theory , vol. 56, no. 10, pp. 5111–5130, 2010
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
H. Rauhut, “Compressive sensing and structured random matrices,” Theoretical foundations and numerical methods for sparse recovery , vol. 9, pp. 1–92, 2010
2010
Earlier work this paper cites.
G. Raskutti, M. J. Wainwright, and B. Yu, “Restricted eigenvalue properties for correlated Gaussian designs,” Journal of Machine Learning Research , vol. 11, no. Aug, pp. 2241–2259, 2010
2010
Earlier work this paper cites.
V. Saligrama and M. Zhao, “Thresholded basis pursuit: Lp algorithm for order-wise optimal support recovery for sparse and approximately sparse signals from noisy random measurements,” IEEE Transactions on Information Theory , vol. 57, no. 3, pp. 1567–1586, 2011
2011
Earlier work this paper cites.
T. T. Cai and L. Wang, “Orthogonal matching pursuit for sparse signal recovery with noise,” IEEE Transactions on Information theory , vol. 57, no. 7, pp. 4680–4688, 2011
2011
Cited alongside, same era.
A. Belloni, V. Chernozhukov, and L. Wang, “Square-root lasso: pivotal recovery of sparse signals via conic programming,” Biometrika , vol. 98, no. 4, pp. 791–806, 2011
2011
Cited alongside, same era.
D. Hsu, S. M. Kakade, and T. Zhang, “Random design analysis of ridge regression,” in Conference on Learning Theory , 2012, pp. 9–1
2012
Cited alongside, same era.
D. L. Donoho, Y. Tsaig, I. Drori, and J.-L. Starck, “Sparse solution of underdetermined systems of linear equations by stagewise orthogonal matching pursuit,” IEEE Transactions on Information Theory , vol. 58, no. 2, pp. 1094–1121, 2012
2012
Cited alongside, same era.
M. Belkin, S. Ma, and S. Mandal, “To understand deep learning we need to understand kernel learning,” in International Conference on Machine Learning , 2018, pp. 540–548
2018
Later among the works it cites.
M. Belkin, D. J. Hsu, and P. Mitra, “Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate,” in Advances in Neural Information Processing Systems , 2018, pp. 2300–2311
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
B. Neyshabur, R. Tomioka, and N. Srebro, “Norm-based capacity control in neural networks,” in Conference on Learning Theory , 2015, pp. 1376–1401
2015
Cited alongside, same era.
R. F. Barber, E. J. Candès et al. , “Controlling the false discovery rate via knockoffs,” The Annals of Statistics , vol. 43, no. 5, pp. 2055–2085, 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht, “The marginal value of adaptive gradient methods in machine learning,” in Advances in Neural Information Processing Systems , 2017, pp. 4148–4158
2017
Cited alongside, same era.
P. L. Bartlett, D. J. Foster, and M. J. Telgarsky, “Spectrally-normalized margin bounds for neural networks,” in Advances in Neural Information Processing Systems , 2017, pp. 6240–6249
2017
Cited alongside, same era.
2017
Cited alongside, same era.
B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro, “Exploring generalization in deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 5947–5956
2017
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
P. C. Bellec, G. Lecué, A. B. Tsybakov et al. , “Slope meets Lasso: improved oracle bounds and optimality,” The Annals of Statistics , vol. 46, no. 6B, pp. 3603–3642, 2018
2018
Later among the works it cites.
N. Verzelen, E. Gassiat et al. , “Adaptive estimation of high-dimensional signal-to-noise ratios,” Bernoulli , vol. 24, no. 4B, pp. 3683–3710, 2018
2018
Later among the works it cites.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
Z. Allen-Zhu, Y. Li, and Z. Song, “A convergence theory for deep learning via over-parameterization,” in International Conference on Machine Learning , 2019, pp. 242–252
2019
Closest in time.
2019
Closest in time.
M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry, “Convergence of gradient descent on separable data,” in The 22nd International Conference on Artificial Intelligence and Statistics , 2019, pp. 3420–3428
2019
Closest in time.
2019
Closest in time.
M. Belkin, A. Rakhlin, and A. B. Tsybakov, “Does data interpolation contradict statistical optimality?” in The 22nd International Conference on Artificial Intelligence and Statistics , 2019, pp. 1611–1619
2019
Closest in time.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off,” Proceedings of the National Academy of Sciences , vol. 116, no. 32, pp. 15 849–15 854, 2019
2019
Closest in time.
2019
Closest in time.
M. J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint . Cambridge University Press, 2019, vol. 48
2019
Closest in time.