Fetching the paper…
Reading the bibliography…
The generalization and learning speed of a multi-class neural network can often be significantly improved by using soft targets that are a weighted average of the hard targets and the uniform distribution over labels.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, Ronald J Williams, et al · 1986
Earlier work this paper cites.
Supervised learning of probability distributions by neural networks
Eric B Baum and Frank Wilczek · 1988
Earlier work this paper cites.
Accelerated learning in layered neural networks
Sara Solla, Esther Levin, and Michael Fleisher · 1988
Earlier work this paper cites.
An empirical study of learning speed in back-propagation networks
Scott E Fahlman · 1988
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Disturblabel: Regularizing cnn on the loss layer
Lingxi Xie, Jingdong Wang, Zhen Wei, Meng Wang, and Qi Tian · 2016
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Introduction to the theory of neural computation
John A Hertz, Anders Krogh, and Richard G. Palmer · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2018
Later among the works it cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, and Zhifeng Chen · 2018
Later among the works it cites.
Analyzing uncertainty in neural machine translation
Myle Ott, Michael Auli, David Grangier, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Simon Kornblith, Jonathon Shlens, and Quoc V Le · 2018
Later among the works it cites.
Calibration of encoder decoder models for neural machine translation
Aviral Kumar and Sunita Sarawagi · 2019
Closest in time.
Adaptive estimators show information compression in deep neural networks
Ivan Chelombiev, Conor Houghton, and Cian O’Donnell · 2019
Closest in time.