Fetching the paper…
Reading the bibliography…
We introduce a tunable loss function called $\alpha$-loss, parameterized by $\alpha \in (0,\infty]$, which interpolates between the exponential loss ($\alpha = 1/2$), the log-loss ($\alpha = 1$), and the 0-1 loss ($\alpha = \infty$), for the machine learning setting of classification.
A. Rényi, “On measures of entropy and information,” in Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics . Berkeley, Calif.: University of California Press, 1961, pp. 547–561. [Online]. Available: https://projecteuclid.org/euclid.bsmsp/1200512181
1961
Earlier work this paper cites.
S. Arimoto, “Information-theoretical considerations on estimation problems,” Information and Control , vol. 19, no. 3, pp. 181 – 194, 1971. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0019995871900659
1971
Earlier work this paper cites.
S. Arimoto, “Information measures and capacity of order α \alpha for discrete memoryless channels,” Topics in Information Theory , 1977
1977
Earlier work this paper cites.
Y. E. Nesterov, “Minimization methods for nonsmooth convex and quasiconvex functions,” Matekon , vol. 29, pp. 519–531, 1984
1984
Earlier work this paper cites.
T. M. Cover, Elements of information theory . John Wiley & Sons, 1999
1999
Earlier work this paper cites.
J. Mo and J. Walrand, “Fair end-to-end window-based congestion control,” IEEE/ACM Transactions on networking , vol. 8, no. 5, pp. 556–567, 2000
2000
Earlier work this paper cites.
J. Friedman, T. Hastie, and R. Tibshirani, The Elements of Statistical Learning . Springer series in statistics New York, 2001, vol. 1, no. 10
2001
Earlier work this paper cites.
C. E. Shannon, “A mathematical theory of communication,” ACM SIGMOBILE mobile computing and communications review , vol. 5, no. 1, pp. 3–55, 2001
2001
Earlier work this paper cites.
A. Papoulis and S. U. Pillai, Probability, Random Variables, and Stochastic Processes . Tata McGraw-Hill Education, 2002
2002
Earlier work this paper cites.
S. Ben-David, N. Eiron, and P. M. Long, “On the difficulty of approximately maximizing agreements,” Journal of Computer and System Sciences , vol. 66, no. 3, pp. 496–514, 2003
2003
Earlier work this paper cites.
Y. Lin, “A note on margin-based loss functions in classification,” Statistical & Probability Letters , vol. 68, no. 1, pp. 73–82, 2004
2004
Earlier work this paper cites.
L. Rosasco, E. D. Vito, A. Caponnetto, M. Piana, and A. Verri, “Are loss functions all the same?” Neural Computation , vol. 16, no. 5, pp. 1063–1076, 2004
2004
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe, “Convexity, classification, and risk bounds,” Journal of the American Statistical Association , vol. 101, no. 473, pp. 138–156, 2006
2006
Earlier work this paper cites.
A. Tewari and P. L. Bartlett, “On the consistency of multiclass classification methods,” Journal of Machine Learning Research , vol. 8, no. May, pp. 1007–1025, 2007
2007
Earlier work this paper cites.
Y. Wu and Y. Liu, “Robust truncated hinge loss support vector machines,” Journal of the American Statistical Association , vol. 102, no. 479, pp. 974–983, 2007
2007
Earlier work this paper cites.
J.-Y. Audibert, A. B. Tsybakov et al. , “Fast learning rates for plug-in classifiers,” The Annals of statistics , vol. 35, no. 2, pp. 608–633, 2007
2007
Earlier work this paper cites.
Y. Sasaki, “The truth of the f-measure,” Teach Tutor Mater , 01 2007
2007
Earlier work this paper cites.
H. Masnadi-Shirazi and N. Vasconcelos, “On the design of loss functions for classification: theory, robustness to outliers, and SavageBoost,” in Advances in Neural Information Processing Systems , 2009, pp. 1049–1056
2009
Earlier work this paper cites.
X. Nguyen, M. J. Wainwright, and M. I. Jordan, “On surrogate loss functions and f f -divergences,” The Annals of Statistics , vol. 37, no. 2, pp. 876–904, 04 2009
2009
Earlier work this paper cites.
O. Chapelle, C. B. Do, C. H. Teo, Q. V. Le, and A. J. Smola, “Tighter bounds for structured estimation,” in Advances in neural information processing systems , 2009, pp. 281–288
2009
Earlier work this paper cites.
A. Singh and J. C. Principe, “A loss function for classification based on a robust similarity metric,” in The 2010 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2010, pp. 1–6
2010
Earlier work this paper cites.
L. Zhao, M. Mammadov, and J. Yearwood, “From convex to nonconvex: a loss function analysis for binary classification,” in 2010 IEEE International Conference on Data Mining Workshops . IEEE, 2010, pp. 1281–1288
2010
Earlier work this paper cites.
M. D. Reid and R. C. Williamson, “Composite binary losses,” The Journal of Machine Learning Research , vol. 11, pp. 2387–2422, 2010
2010
Earlier work this paper cites.
P. M. Long and R. A. Servedio, “Random classification noise defeats all convex potential boosters,” Machine learning , vol. 78, no. 3, pp. 287–304, 2010
2010
Earlier work this paper cites.
R. A. Horn and C. R. Johnson, Matrix Analysis . Cambridge university press, 2012
2012
Earlier work this paper cites.
T. Nguyen and S. Sanner, “Algorithms for direct 0–1 loss optimization in binary classification,” in International Conference on Machine Learning , 2013, pp. 1085–1093
2013
Earlier work this paper cites.
R. E. Schapire and Y. Freund, “Boosting: Foundations and algorithms,” Kybernetes , 2013
2013
Cited alongside, same era.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . Cambridge university press, 2014
2014
Cited alongside, same era.
A. Krizhevsky, V. Nair, and G. Hinton, “The CIFAR-10 dataset,” online: http://www. cs. toronto. edu/kriz/cifar. html , vol. 55, 2014
2014
Cited alongside, same era.
S. Verdú, “ α \alpha -mutual information,” in 2015 Information Theory and Applications Workshop (ITA) , 2015, pp. 1–6
2015
Cited alongside, same era.
E. Hazan, K. Levy, and S. Shalev-Shwartz, “Beyond convexity: Stochastic quasi-convex optimization,” in Advances in Neural Information Processing Systems , 2015, pp. 1594–1602
2015
Cited alongside, same era.
J. Liao, L. Sankar, O. Kosut, and F. P. Calmon, “Robustness of maximal α \alpha -leakage to side information,” in 2019 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2019, pp. 642–646
2019
Closest in time.
I. Issa, A. B. Wagner, and S. Kamath, “An operational approach to information leakage,” IEEE Transactions on Information Theory , vol. 66, no. 3, pp. 1625–1657, 2019
2019
Closest in time.
2019
Closest in time.
L. Engstrom, B. Tran, D. Tsipras, L. Schmidt, and A. Madry, “Exploring the landscape of spatial robustness,” in International Conference on Machine Learning , 2019, pp. 1802–1811
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT press, 2016
2016
Cited alongside, same era.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Cited alongside, same era.
I. Sason and S. Verdú, “Arimoto–rényi conditional entropy and bayesian m m -ary hypothesis testing,” IEEE Transactions on Information theory , vol. 64, no. 1, pp. 4–25, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Q. Nguyen and M. Hein, “The loss surface of deep and wide neural networks,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 2603–2612
2017
Cited alongside, same era.
J. Pennington and P. Worah, “Nonlinear random matrix theory for deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 2637–2646
2017
Cited alongside, same era.
A. Xu and M. Raginsky, “Information-theoretic analysis of generalization capability of learning algorithms,” in Advances in Neural Information Processing Systems , 2017, pp. 2524–2533
2017
Cited alongside, same era.
2019
Closest in time.
H. Wang, M. Diaz, J. C. S. Santos Filho, and F. P. Calmon, “An information-theoretic view of generalization via wasserstein distance,” in 2019 IEEE International Symposium on Information Theory (ISIT) . IEEE, 2019, pp. 577–581
2019
Closest in time.
Y. Wang, X. Ma, Z. Chen, Y. Luo, J. Yi, and J. Bailey, “Symmetric cross entropy for robust learning with noisy labels,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 322–330
2019
Closest in time.
E. Amid, M. K. Warmuth, R. Anil, and T. Koren, “Robust bi-tempered logistic loss based on bregman divergences,” in Advances in Neural Information Processing Systems , 2019, pp. 15 013–15 022
2019
Closest in time.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” in Advances in neural information processing systems , 2019, pp. 8026–8037
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
P. Kairouz, J. Liao, C. Huang, and L. Sankar, “Censored and fair universal representations using generative adversarial models,” arXiv , pp. arXiv–1910, 2019
2019
Closest in time.
T. Sypherd, M. Diaz, L. Sankar, and G. Dasarathy, “On the α \alpha -loss landscape in the logistic model,” in 2020 IEEE International Symposium on Information Theory (ISIT) , 2020, pp. 2700–2705
2020
Closest in time.
——, “Maximal α \alpha -leakage and its properties,” in 2020 IEEE Conference on Communications and Network Security (CNS) . IEEE, 2020, pp. 1–6
2020
Closest in time.
2020
Closest in time.
C. Walder and R. Nock, “All your loss are belong to bayes,” arXiv preprint arXiv:2006.04633 , 2020
2020
Closest in time.
D. Russo and J. Zou, “How much does your data exploration overfit? controlling bias via information usage,” IEEE Transactions on Information Theory , vol. 66, no. 1, pp. 302–323, 2020
2020
Closest in time.
Y. Bu, S. Zou, and V. V. Veeravalli, “Tightening mutual information-based bounds on generalization error,” IEEE Journal on Selected Areas in Information Theory , vol. 1, no. 1, pp. 121–130, 2020
2020
Closest in time.
T. Steinke and L. Zakynthinou, “Reasoning about generalization via conditional mutual information,” in Conference on Learning Theory . PMLR, 2020, pp. 3437–3452
2020
Closest in time.
A. R. Esposito, M. Gastpar, and I. Issa, “Generalization error bounds via rényi-, f-divergences and maximal leakage,” IEEE Transactions on Information Theory , 2021
2021
Closest in time.
B. R. Gálvez, G. Bassi, R. Thobaben, and M. Skoglund, “Tighter expected generalization error bounds via wasserstein distance,” in Thirty-Fifth Conference on Neural Information Processing Systems , 2021
2021
Closest in time.
G. Neu, G. K. Dziugaite, M. Haghifam, and D. M. Roy, “Information-theoretic generalization bounds for stochastic gradient descent,” in COLT , 2021
2021
Closest in time.
2021
Closest in time.
G. R. Kurri, T. Sypherd, and L. Sankar, “Realizing gans via a tunable loss function,” 2021
2021
Closest in time.
R. Nock, T. Sypherd, and L. Sankar, “Being properly improper,” 2021
2021
Closest in time.
J. K. Cava, “Code for a tunable loss function for robust classification,” 2021. [Online]. Available: https://github.com/SankarLab/AlphaLoss
2021
Closest in time.