Fetching the paper…
Reading the bibliography…
Label smoothing is commonly used in training deep learning models, wherein one-hot training labels are mixed with uniform label vectors.
The well-calibrated Bayesian
A. P. Dawid · 1982
Earlier work this paper cites.
Learning from noisy examples
Dana Angluin and Philip Laird · 1988
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
Convexity, classification, and risk bounds
Peter L Bartlett, Michael I Jordan, and Jon D McAuliffe · 2006
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Random classification noise defeats all convex potential boosters
Philip M. Long and Rocco A. Servedio · 2010
Earlier work this paper cites.
Learning with noisy labels
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep D Ravikumar, and Ambuj Tewari · 2013
Earlier work this paper cites.
Classification with asymmetric label noise: consistency and maximal denoising
Clayton Scott, Gilles Blanchard, and Gregory Handy · 2013
Earlier work this paper cites.
Consistency of losses for learning from weak labels
Jesús Cid-Sueiro, Darío García-García, and Raúl Santos-Rodríguez · 2014
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping, 2014
Scott Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Training convolutional networks with noisy labels
Sainbayar Sukhbaatar, Joan Bruna, Manohar Paluri, Lubomir Bourdev, and Rob Fergus · 2015
Earlier work this paper cites.
Learning with symmetric label noise: the importance of being unhinged
Brendan van Rooyen, Aditya Krishna Menon, and Robert C. Williamson · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Loss factorization, weakly supervised learning and label noise robustness
Giorgio Patrini, Frank Nielsen, Richard Nock, and Marcello Carioni · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Disturblabel: Regularizing CNN on the loss layer
Lingxi Xie, Jingdong Wang, Zhen Wei, Meng Wang, and Qi Tian · 2016
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Regularized Evolution for Image Classifier Architecture Search
E. Real, A. Aggarwal, Y. Huang, and Q. V Le · 2018
Later among the works it cites.
A theory of learning with corrupted labels
Brendan van Rooyen and Robert C. Williamson · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le · 2018
Later among the works it cites.
Robust bi-tempered logistic loss based on Bregman divergences
Ehsan Amid, Manfred K. Warmuth, Rohan Anil, and Tomer Koren · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Positive-unlabeled learning with non-negative risk estimator
Ryuichi Kiryo, Gang Niu, Marthinus C. du Plessis, and Masashi Sugiyama · 2017
Cited alongside, same era.
Making deep neural networks robust to label noise: a loss correction approach
Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Large batch training of convolutional networks, 2017
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Nontawat Charoenphakdee, Jongyeong Lee, and Masashi Sugiyama · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2019
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Later among the works it cites.
Combating label noise in deep learning using abstention
Sunil Thulasidasan, Tanmoy Bhattacharya, Jeff Bilmes, Gopinath Chennupati, and Jamal Mohd-Yusof · 2019
Later among the works it cites.
Are anchor points really indispensable in label-noise learning?
Xiaobo Xia, Tongliang Liu, Nannan Wang, Bo Han, Chen Gong, Gang Niu, and Masashi Sugiyama · 2019
Later among the works it cites.
Regularization via structural label smoothing
Weizhi Li, Gautam Dasarathy, and Visar Berisha · 2020
Closest in time.