Fetching the paper…
Reading the bibliography…
We introduce a temperature into the exponential function and replace the softmax output layer of neural nets by a high temperature generalization.
Relative loss bounds for single neurons
D. P. Helmbold, J. Kivinen, and M. K. Warmuth · 1999
Earlier work this paper cites.
The MNIST database of handwritten digits, 1999
Yann LeCun and Corinna Cortes · 1999
Earlier work this paper cites.
Relative loss bounds for multidimensional regression problems
J. Kivinen and M. K. Warmuth · 2001
Earlier work this paper cites.
Deformed exponentials and logarithms in generalized thermostatistics
Jan Naudts · 2002
Earlier work this paper cites.
Loss functions for binary class probability estimation and classification: Structure and applications
Andreas Buja, Werner Stuetzle, and Yi Shen · 2005
Earlier work this paper cites.
On the consistency of multiclass classification methods
Ambuj Tewari and Peter L Bartlett · 2007
Earlier work this paper cites.
Random classification noise defeats all convex potential boosters
Philip M Long and Rocco A Servedio · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Cited alongside, same era.
Surrogate regret bounds for proper losses
M. D. Reid and R. C. Williamson · 2009
Cited alongside, same era.
Families of alpha-beta-and gamma-divergences: Flexible and robust measures of similarities
Andrzej Cichocki and Shun-ichi Amari · 2010
Cited alongside, same era.
t t -logistic regression
Nan Ding and S. V. N. Vishwanathan · 2010
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz et al · 2012
Cited alongside, same era.
Statistical machine learning in the t-exponential family of distributions
Nan Ding · 2013
Later among the works it cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Composite multiclass losses
Robert C. Williamson, Elodie Vernet, and Mark D. Reid · 2016
Later among the works it cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew D Zeiler · 2012
Cited alongside, same era.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhilu Zhang and Mert Sabuncu · 2018
Later among the works it cites.
Two-temperature logistic regression based on the Tsallis divergence
Ehsan Amid, Manfred K. Warmuth, and Sriram Srinivasan · 2019
Closest in time.