Fetching the paper…
Reading the bibliography…
A loss function measures the discrepancy between the true values (observations) and their estimated fits, for a given instance of data.
E. H. Shuford, A. Albert, and H. E. Massengill, “Admissible probability measurement procedures,” Psychometrika , vol. 31, no. 2, pp. 125–145, 1966
1966
Earlier work this paper cites.
L. J. Savage, “Elicitation of personal probabilities and expectations,” J. Am. Stat. Assoc. , vol. 66, no. 336, pp. 783–801, 1971
1971
Earlier work this paper cites.
N. Merhav and M. Feder, “Universal schemes for sequential decision from individual data sequences,” IEEE Trans. Inform. Theory , vol. 39, no. 4, pp. 1280–1292, 1993
1993
Earlier work this paper cites.
F. Pereira, N. Tishby, and L. Lee, “Distributional clustering of english words,” in Proceedings of the 31st annual meeting on Association for Computational Linguistics . Association for Computational Linguistics, 1993, pp. 183–190
1993
Earlier work this paper cites.
N. Merhav and M. Feder, “Universal prediction,” IEEE Trans. Inform. Theory , vol. 44, no. 6, pp. 2124–2147, 1998
1998
Earlier work this paper cites.
H. H. Bauschke and J. M. Borwein, “Joint and separate convexity of the bregman distance,” Studies in Computational Mathematics , vol. 8, pp. 23–36, 2001
2001
Earlier work this paper cites.
A. Buja, W. Stuetzle, and Y. Shen, “Loss functions for binary class probability estimation and classification: Structure and applications,” Working draft, November , 2005
2005
Cited alongside, same era.
J. Langford, “Tutorial on practical prediction theory for classification,” J. Machine Learning Res. , vol. 6, no. Mar, pp. 273–306, 2005
2005
Cited alongside, same era.
A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh, “Clustering with bregman divergences,” J. Machine Learning Res. , vol. 6, no. Oct, pp. 1705–1749, 2005
2005
Cited alongside, same era.
T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation,” J. Am. Stat. Assoc. , vol. 102, no. 477, pp. 359–378, 2007
2007
Cited alongside, same era.
P. Harremoës and N. Tishby, “The information bottleneck revisited or how to choose a good distortion measure,” in Proc. Int. Symp. Inform. Theory . IEEE, 2007, pp. 566–570
T. M. Cover and J. A. Thomas, Elements of Information Theory . John Wiley & Sons, 2012
2012
Later among the works it cites.
J. Jiao, T. A. Courtade, A. No, K. Venkat, and T. Weissman, “Information measures: the curious case of the binary alphabet,” IEEE Trans. Inform. Theory , vol. 60, no. 12, pp. 7616–7626, 2014
2014
Later among the works it cites.
C. L. Byrne, Iterative Optimization in Inverse Problems . CRC Press, 2014
2014
Later among the works it cites.
I. Sason and S. Verdú, “ f f -divergence inequalities,” IEEE Trans. Inform. Theory , vol. 62, no. 11, pp. 5973–6006, 2016
2016
Later among the works it cites.
A. Painsky and S. Rosset, “Cross-validated variable selection in tree-based methods improves predictive performance,” IEEE Trans. Pattern Analysis and Machine Intelligence , vol. 39, no. 11, pp. 2142–2153, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2007
Cited alongside, same era.
M. D. Reid and R. C. Williamson, “Composite binary losses,” J. Machine Learning Res. , vol. 11, no. Sep, pp. 2387–2422, 2010
2010
Cited alongside, same era.