Fetching the paper…
Reading the bibliography…
Recently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning.
J. F. Claerbout and F. Muir, “Robust modeling with erratic data,”
1973
Earlier work this paper cites.
E. J. Candes and T. Tao, “Decoding by linear programming,”
2005
Earlier work this paper cites.
J. Wright, A. Y. Yang, A. Ganesh, S. S. Sastry, and Y. Ma, “Robust face recognition via sparse representation,”
2008
Earlier work this paper cites.
A. Krizhevsky, G. Hinton,
2009
Earlier work this paper cites.
A. Cohen, W. Dahmen, and R. DeVore, “Compressed sensing and best
2009
Earlier work this paper cites.
E. J. Candès, X. Li, Y. Ma, and J. Wright, “Robust principal component analysis?,”
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,”
2012
Earlier work this paper cites.
B. Frénay and M. Verleysen, “Classification in the presence of label noise: a survey,”
2013
Earlier work this paper cites.
A. Domahidi, E. Chu, and S. Boyd, “ECOS: An SOCP solver for embedded systems,” in
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Xiao, T. Xia, Y. Yang, C. Huang, and X. Wang, “Learning from massive noisy labeled data for image classification,” in
2015
Earlier work this paper cites.
R. Ge, F. Huang, C. Jin, and Y. Yuan, “Escaping from saddle points—online stochastic gradient for tensor decomposition,” in
2015
Earlier work this paper cites.
T. Liu and D. Tao, “Classification with noisy labels by importance reweighting,”
2015
Earlier work this paper cites.
X. Chen and A. Gupta, “Webly supervised learning of convolutional networks,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht, “Gradient descent only converges to minimizers,” in
2016
Earlier work this paper cites.
S. Diamond and S. Boyd, “CVXPY: A Python-embedded modeling language for convex optimization,”
2016
Earlier work this paper cites.
M. A. Davenport and J. Romberg, “An overview of low-rank matrix recovery from incomplete observations,”
2016
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts,”
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Patrini, A. Rozza, A. Krishna Menon, R. Nock, and L. Qu, “Making deep neural networks robust to label noise: A loss correction approach,” in
2017
Earlier work this paper cites.
A. Ghosh, H. Kumar, and P. Sastry, “Robust loss functions under label noise for deep neural networks,” in
2017
Earlier work this paper cites.
R. Wang, T. Liu, and D. Tao, “Multiclass learning with partially corrupted labels,”
2017
Earlier work this paper cites.
H.-S. Chang, E. Learned-Miller, and A. McCallum, “Active bias: Training more accurate neural networks by emphasizing high variance samples,”
2017
Earlier work this paper cites.
J. Goldberger and E. Ben-Reuven, “Training deep neural-networks using a noise adaptation layer,” 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. Tsang, and M. Sugiyama, “Co-teaching: Robust training of deep neural networks with extremely noisy labels,” in
2018
Earlier work this paper cites.
Z. Zhang and M. R. Sabuncu, “Generalized cross entropy loss for training deep neural networks with noisy labels,” in
2018
Earlier work this paper cites.
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro, “The implicit bias of gradient descent on separable data,”
2018
Earlier work this paper cites.
S. Gunasekar, B. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro, “Implicit regularization in matrix factorization,” in
2018
Earlier work this paper cites.
Y. Li, T. Ma, and H. Zhang, “Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations,” in
2018
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in
2018
Earlier work this paper cites.
D. Hendrycks, M. Mazeika, D. Wilson, and K. Gimpel, “Using trusted data to train deep networks on labels corrupted by severe noise,”
2018
Earlier work this paper cites.
X. Ma, Y. Wang, M. E. Houle, S. Zhou, S. Erfani, S. Xia, S. Wijewickrema, and J. Bailey, “Dimensionality-driven learning with noisy labels,” in
2018
Cited alongside, same era.
D. Tanaka, D. Ikami, T. Yamasaki, and K. Aizawa, “Joint optimization framework for learning with noisy labels,” in
2018
Cited alongside, same era.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off,”
2019
Cited alongside, same era.
T. Vaskevicius, V. Kanade, and P. Rebeschini, “Implicit regularization for optimal sparse recovery,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Li, M. Soltanolkotabi, and S. Oymak, “Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks,” in
2020
Later among the works it cites.
L. Huang, C. Zhang, and H. Zhang, “Self-adaptive training: beyond empirical risk minimization,”
2020
Later among the works it cites.
S. Zheng, P. Wu, A. Goswami, M. Goswami, D. Metaxas, and C. Chen, “Error-bounded correction of noisy labels,” in
2020
Later among the works it cites.
V. Papyan, X. Han, and D. L. Donoho, “Prevalence of neural collapse during the terminal phase of deep learning training,”
2020
Later among the works it cites.
Q. Xie, Z. Dai, E. H. Hovy, M.-T. Luong, and Q. V. Le, “Unsupervised data augmentation for consistency training,”
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Chizat, E. Oyallon, and F. Bach, “On lazy training in differentiable programming,”
2019
Cited alongside, same era.
Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, and J. Bailey, “Symmetric cross entropy for robust learning with noisy labels,” in
2019
Cited alongside, same era.
D. Kalimeris, G. Kaplun, P. Nakkiran, B. Edelman, T. Yang, B. Barak, and H. Zhang, “Sgd on neural networks learns functions of increasing complexity,”
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Chi, Y. M. Lu, and Y. Chen, “Nonconvex optimization meets low-rank matrix factorization: An overview,”
2019
Cited alongside, same era.
S. Oymak and M. Soltanolkotabi, “Overparameterized nonlinear learning: Gradient descent takes the shortest path?,” in
2019
Cited alongside, same era.
S. Arora, N. Cohen, W. Hu, and Y. Luo, “Implicit regularization in deep matrix factorization,” in
2019
Cited alongside, same era.
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning (still) requires rethinking generalization,”
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Cheng, Z. Zhu, X. Li, Y. Gong, X. Sun, and Y. Liu, “Learning with instance-dependent label noise: A sample sieve approach,” in
2021
Later among the works it cites.
D. Stöger and M. Soltanolkotabi, “Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction,”
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Li, T. Nguyen, C. Hegde, and K. W. Wong, “Implicit sparse regularization: The impact of depth and early stopping,”
2021
Later among the works it cites.
H.-H. Chou, J. Maly, and H. Rauhut, “More is less: Inducing sparsity via overparameterization,”
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Ding, L. Jiang, Y. Chen, Q. Qu, and Z. Zhu, “Rank overspecified robust matrix recovery: Subgradient method and exact recovery,”
2021
Later among the works it cites.
2021
Later among the works it cites.
G. Algan and I. Ulusoy, “Image classification with deep learning in the presence of noisy labels: A survey,”
2021
Later among the works it cites.
J. Wei and Y. Liu, “When optimizing
2021
Later among the works it cites.
H. Zhang, X. Xing, and L. Liu, “Dualgraph: A graph-based method for reasoning about label noise,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Kim, J. Ko, J. Choi, S.-Y. Yun,
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Z. Lin and J. Bradic, “Learning to combat noisy labels via classification margins,”
2021
Later among the works it cites.
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste, “Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,”
2021
Later among the works it cites.
S. Liu, L. Yin, D. C. Mocanu, and M. Pechenizkiy, “Do we actually need dense over-parameterization? in-time over-parameterization in sparse training,” in
2021
Later among the works it cites.
T. Chen, Z. Zhang, S. Balachandra, H. Ma, Z. Wang, Z. Wang,
2021
Later among the works it cites.
2021
Later among the works it cites.
R. J. Tibshirani, “Equivalences between sparse models and neural networks,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Zhu, T. Ding, J. Zhou, X. Li, C. You, J. Sulam, and Q. Qu, “A geometric analysis of neural collapse with unconstrained features,”
2021
Later among the works it cites.
C. Fang, H. He, Q. Long, and W. J. Su, “Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training,”
2021
Later among the works it cites.
2021
Later among the works it cites.