Fetching the paper…
Reading the bibliography…
Large neural networks trained in the overparameterized regime are able to fit noise to zero train error.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 1908
Earlier work this paper cites.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le · 1911
Earlier work this paper cites.
Learning from noisy examples
Dana Angluin and Philip Laird · 1988
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
A nonasymptotic condorcet jury theorem
Ruth Ben-Yashar and Jacob Paroush · 2000
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Avrim Blum, Adam Kalai, and Hal Wasserman · 2003
Earlier work this paper cites.
Hieu Pham, Qizhe Xie, Zihang Dai, and Quoc V. Le · 2003
Earlier work this paper cites.
Monotonicity in condorcet jury theorem
Daniel Berend and Luba Sapir · 2005
Earlier work this paper cites.
Why distillation helps: a statistical perspective
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim, and Sanjiv Kumar · 2005
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E. Hinton · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
The deep bootstrap: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2010
Cited alongside, same era.
Classification in the presence of label noise: a survey
Benoît Frénay and Michel Verleysen · 2013
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Billion-scale semi-supervised learning for image classification, 2019
I. Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri, and Dhruv Mahajan · 2019
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Later among the works it cites.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2020
Later among the works it cites.
Learning from noisy labels with deep neural networks: A survey
Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Born again neural networks, 2018
Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang · 2019
Cited alongside, same era.
A function space view of bounded norm infinite width relu nets: The multivariate case
Greg Ongie, Rebecca Willett, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Colin Wei, Kendrick Shen, Yining Chen, and Tengyu Ma · 2020
Later among the works it cites.
Self-training converts weak learners to strong learners in mixture models
Spencer Frei, Difan Zou, Zixiang Chen, and Quanquan Gu · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Later among the works it cites.
Iterative label improvement: Robust training by confidence based filtering and dataset partitioning
Christian Haase-Schütz, Rainer Stal, Heinz Hertlein, and Bernhard Sick · 2021
Later among the works it cites.
Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen, and Jinjun Xiong · 2022
Closest in time.