Fetching the paper…
Reading the bibliography…
A loss function measures the discrepancy between the true values and their estimated fits, for a given instance of data.
E. H. Shuford, A. Albert, and H. E. Massengill, “Admissible probability measurement procedures,” Psychometrika , vol. 31, no. 2, pp. 125–145, June 1966
1966
Earlier work this paper cites.
I. Csiszár, “Information-type measures of difference of probability distributions and indirect observations,” Studia Sci. Math. Hungar. , vol. 2, no. Jan., pp. 299–318, 1967
1967
Earlier work this paper cites.
J. MacQueen, “Some methods for classification and analysis of multivariate observations,” in Proc. Berkeley Symp. Math. Stat., Prob. , vol. 1, 1967, pp. 281–297
1967
Earlier work this paper cites.
F. Itakura and S. Saito, “Analysis synthesis telephony based on the maximum likelihood method,” in Proc. Int. Congress Acoustics , Aug. 1968, pp. C17–C20
1968
Earlier work this paper cites.
L. J. Savage, “Elicitation of personal probabilities and expectations,” J. Am. Stat. Assoc. , vol. 66, no. 336, pp. 783–801, Dec. 1971
1971
Earlier work this paper cites.
M. Zakai and J. Ziv, “A generalization of the rate-distortion theory and applications,” in Information Theory New Trends and Open Problems , G. Longo, Ed. Vienna, Austria: Springer, 1975, pp. 87–123
1975
Earlier work this paper cites.
Y. Linde, A. Buzo, and R. Gray, “An algorithm for vector quantizer design,” IEEE Trans. Commun. , vol. 28, no. 1, pp. 84–95, Jan. 1980
1980
Earlier work this paper cites.
A. Buzo, A. Gray, R. Gray, and J. Markel, “Speech coding based upon vector quantization,” IEEE Trans. Acoustics, Speech, and Signal Processing , vol. 28, no. 5, pp. 562–574, 1980
1980
Earlier work this paper cites.
L. Breiman, J. Friedman, C. J. Stone, and R. A. Olshen, Classification and Regression Trees . Boca Raton, FL: CRC Press, 1984
1984
Earlier work this paper cites.
R. Linsker, “Self-organization in a perceptual network,” Computer , vol. 21, no. 3, pp. 105–117, Mar. 1988
1988
Earlier work this paper cites.
D. Haussler, “Decision theoretic generalizations of the PAC model for neural net and other learning applications,” Inf. Comput. , vol. 100, no. 1, pp. 78–150, Sep. 1992
1992
Earlier work this paper cites.
N. Merhav and M. Feder, “Universal schemes for sequential decision from individual data sequences,” IEEE Trans. Inform. Theory , vol. 39, no. 4, pp. 1280–1292, Apr. 1993
1993
Earlier work this paper cites.
F. Pereira, N. Tishby, and L. Lee, “Distributional clustering of English words,” in Proc. Meet. Assoc. Comput. Ling. Columbus, OH: Association for Computational Linguistics, June 1993, pp. 183–190
1993
Earlier work this paper cites.
I. Csiszár, “Generalized projections for non-negative functions,” Acta Math. Hungar. , vol. 68, no. 1–2, pp. 161–185, 1995
1995
Earlier work this paper cites.
R. L. Winkler, J. Munoz, J. L. Cervera, J. M. Bernardo, G. Blattenberger, J. B. Kadane, D. V. Lindley, A. H. Murphy, R. M. Oliver, and D. Ríos-Insua, “Scoring rules and the evaluation of probabilities,” Test , vol. 5, no. 1, pp. 1–60, June 1996
1996
Earlier work this paper cites.
N. Merhav and M. Feder, “Universal prediction,” IEEE Trans. Inform. Theory , vol. 44, no. 6, pp. 2124–2147, June 1998
1998
Earlier work this paper cites.
L. D. Baker and A. K. McCallum, “Distributional clustering of words for text classification,” in Proc. Int. Conf. Res., Dev. Inform. Retrieval (SIGIR) , Melbourne, Australia, Aug. 1998, pp. 96–103
1998
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” in Proc. Allerton Conf. Commun., Contr., Computing , Monticello, IL, Sep. 1999, pp. 368–377
1999
Earlier work this paper cites.
D. A. McAllester, “PAC-Bayesian model averaging,” in Proc. Conf. Comput. Learn. Theory (COLT) , Santa Cruz, CA, July 1999, pp. 164–170
1999
Earlier work this paper cites.
J. M. Bernardo and A. F. M. Smith, Bayesian Theory . New York, NY: Wiley, 2000
2000
Earlier work this paper cites.
N. Slonim and N. Tishby, “Document clustering using word clusters via the information bottleneck method,” in Proc. Int. Conf. Res., Dev. Inform. Retrieval (SIGIR) , Athens, Greece, July 2000, pp. 208–215
2000
Earlier work this paper cites.
H. H. Bauschke and J. M. Borwein, “Joint and separate convexity of the Bregman distance,” in Inherently Parallel Algorithms in Feasibility and Optimization and their Applications , D. Butnariu, Y. Censor, and S. Reich, Eds. Elsevier, 2001, vol. 8, pp. 23–36
2001
Earlier work this paper cites.
A. Clark, “Unsupervised induction of stochastic context-free grammars using distributional clustering,” in Proc. Workshop Comput. Nat. Lang. Learn. (ConLL) , vol. 7, Toulouse, France, July 2001
2001
Earlier work this paper cites.
R. Bekkerman, R. El-Yaniv, N. Tishby, and Y. Winter, “On feature distributional clustering for text categorization,” in Proc. Int. Conf. Res., Dev. Inform. Retrieval (SIGIR) , New Orleans, LA, 2001, pp. 146–153
2001
Earlier work this paper cites.
N. Friedman, O. Mosenzon, N. Slonim, and N. Tishby, “Multivariate information bottleneck,” in Proc. Conf. Uncertainty in Artificial Intelligence (UAI) , San Francisco, CA, Aug. 2001, pp. 152–161
2001
Earlier work this paper cites.
N. Abe, J.-i. Takeuchi, and M. K. Warmuth, “Polynomial learnability of stochastic rules with respect to the KL-divergence and quadratic distance,” IEICE Trans. Inform., Syst. , vol. 84, no. 3, pp. 299–316, Mar. 2001
2001
Earlier work this paper cites.
J. Keith and D. P. Kroese, “Sequence alignment by rare event simulation,” in Proc. Winter Simulation Conf. (WSC) , vol. 1. IEEE, 2002, pp. 320–327
2002
Earlier work this paper cites.
J. Sinkkonen and S. Kaski, “Clustering based on conditional distributions in an auxiliary space,” Neural Comput. , vol. 14, no. 1, pp. 217–239, 2002
2002
Cited alongside, same era.
E. Schneidman, N. Slonim, N. Tishby, R. de Ruyter van Steveninck, and W. Bialek, “Analyzing neural codes using the information bottleneck method,” Unpublished manuscript, 2001. [Online]. Available: ftp://ftp.cis.upenn.edu/pub/cse140/public_html/2002/schneidman.pdf
2002
Cited alongside, same era.
M. Seeger, “PAC-Bayesian generalisation error bounds for Gaussian process classification,” J. Mach. Learn. Res. , vol. 3, no. Oct., pp. 233–269, 2002
2002
Cited alongside, same era.
I. S. Dhillon, S. Mallela, and R. Kumar, “A divisive information-theoretic feature clustering algorithm for text classification,” J. Mach. Learn. Res. , vol. 3, no. Mar., pp. 1265–1287, 2003
2003
Cited alongside, same era.
A. Buja, W. Stuetzle, and Y. Shen, “Loss functions for binary class probability estimation and classification: Structure and applications,” Statistics Dept., Wharton School, University of Pennsylvania, Philadelphia, PA, Tech. Rep., Nov. 2005. [Online]. Available: https://faculty.wharton.upenn.edu/wp-content/uploads/2012/04/Paper-proper-scoring.pdf
2012
Later among the works it cites.
M. Parry, A. P. Dawid, and S. Lauritzen, “Proper local scoring rules,” Ann. Stat. , vol. 40, no. 1, pp. 561–592, 2012
2012
Later among the works it cites.
M. Liu, B. C. Vemuri, S.-i. Amari, and F. Nielsen, “Shape retrieval using hierarchical total Bregman soft clustering,” IEEE Trans. Pattern Anal., Mach. Intell. , vol. 34, no. 12, pp. 2407–2419, Dec. 2012
2012
Later among the works it cites.
E. C. Merkle and M. Steyvers, “Choosing a strictly proper scoring rule,” Decision Analysis , vol. 10, no. 4, pp. 292–304, Dec. 2013
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Csiszár and P. C. Shields, “Information theory and statistics: A tutorial,” Foundations and Trends in Communications and Information Theory , vol. 1, no. 4, pp. 417–528, 2004
2004
Cited alongside, same era.
T. Zhang, “Statistical behavior and consistency of classification methods based on convex risk minimization,” Ann. Stat. , vol. 32, no. 1, pp. 56–134, 2004
2004
Cited alongside, same era.
A. Maurer, “A note on the PAC Bayesian theorem,” CoRR , vol. cs.LG/0411099, 2004. [Online]. Available: http://arxiv.org/abs/cs.LG/0411099
2004
Cited alongside, same era.
A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh, “Clustering with Bregman divergences,” J. Mach. Learn. Res. , vol. 6, pp. 1705–1749, Oct. 2005
2005
Cited alongside, same era.
N. Slonim, G. S. Atwal, G. Tkačik, and W. Bialek, “Information-based clustering,” Proc. Nat. Acad. Sci. (PNAS) , vol. 102, no. 51, pp. 18 297–18 302, 2005
2005
Cited alongside, same era.
J. Langford, “Tutorial on practical prediction theory for classification,” J. Mach. Learn. Res. , vol. 6, no. Mar., pp. 273–306, 2005
2005
Cited alongside, same era.
T. M. Cover and J. A. Thomas, Elements of Information Theory , 2nd ed. New York, NY: John Wiley & Sons, 2006
2006
Cited alongside, same era.
T. Gneiting and A. E. Raftery, “Strictly proper scoring rules, prediction, and estimation,” J. Am. Stat. Assoc. , vol. 102, no. 477, pp. 359–378, Mar. 2007
2007
Cited alongside, same era.
2013
Later among the works it cites.
A. P. Dawid and M. Musio, “Theory and applications of proper scoring rules,” METRON , vol. 72, no. 2, pp. 169–183, Aug. 2014
2014
Later among the works it cites.
C. L. Byrne, Iterative Optimization in Inverse Problems . Boca Raton, FL: CRC Press, 2014
2014
Later among the works it cites.
J. Jiao, T. A. Courtade, A. No, K. Venkat, and T. Weissman, “Information measures: the curious case of the binary alphabet,” IEEE Trans. Inform. Theory , vol. 60, no. 12, pp. 7616–7626, Dec. 2014
2014
Later among the works it cites.
S. Bharadwaj and M. Hasegawa-Johnson, “A PAC-Bayesian approach to minimum perplexity language modeling,” in Proc. Int. Conf. Comput. Ling. (COLING) , Dublin, Ireland, Aug. 2014, pp. 130–140
2014
Later among the works it cites.
A. No and T. Weissman, “Universality of logarithmic loss in lossy compression,” in Proc. Int. Symp. Inform. Theory (ISIT) . Hong Kong, China: IEEE, June 2015, pp. 2166–2170
2015
Later among the works it cites.
J. Jiao, T. A. Courtade, K. Venkat, and T. Weissman, “Justification of logarithmic loss via the benefit of side information,” IEEE Trans. Inform. Theory , vol. 61, no. 10, pp. 5357–5365, 2015
2015
Later among the works it cites.
R. Nock, F. Nielsen, and S.-i. Amari, “On conformal divergences and their population minimizers,” IEEE Trans. Inform. Theory , vol. 62, no. 1, pp. 527–538, Jan. 2015
2015
Later among the works it cites.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in Proc. Inform. Theory Workshop (ITW) , Apr. 2015
2015
Later among the works it cites.
I. Sason and S. Verdú, “ f f -divergence inequalities,” IEEE Trans. Inform. Theory , vol. 62, no. 11, pp. 5973–6006, Nov. 2016
2016
Later among the works it cites.
A. Painsky, S. Rosset, and M. Feder, “Generalized independent component analysis over finite alphabets,” IEEE Trans. Inform. Theory , vol. 62, no. 2, pp. 1038–1053, Feb. 2016
2016
Later among the works it cites.
Australian Bureau of Meteorology, “Australian data archive for meteorology,” http://www.bom.gov.au/climate/data/ , retrieved 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Y. Shkel and S. Verdú, “A single-shot approach to lossy source coding under logarithmic loss,” IEEE Trans. Inform. Theory , vol. 64, no. 1, pp. 129–147, Jan. 2017
2017
Later among the works it cites.
A. Painsky and S. Rosset, “Cross-validated variable selection in tree-based methods improves predictive performance,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 11, pp. 2142–2153, Nov. 2017
2017
Later among the works it cites.
A. Painsky and G. Wornell, “On the universality of the logistic loss function,” in Proc. Int. Symp. Inform. Theory (ISIT) , Vail, Colorado, June 2018, pp. 936–940
2018
Closest in time.
I. Sason, “On f f -divergences: Integral representations, local behavior, and inequalities,” Entropy , vol. 20, no. 5, 383, 2018
2018
Closest in time.
——, “Linear independent component analysis over finite fields: Algorithms and bounds,” IEEE Trans. Signal Processing , vol. 66, no. 22, Nov. 15, 2018
2018
Closest in time.
R. Shwartz-Ziv, A. Painsky, and N. Tishby, “Representation compression and generalization in deep neural networks,” Unpublished Manuscript, 2018. [Online]. Available: https://openreview.net/pdf?id=SkeL6sCqK7
2018
Closest in time.
M. Broniatowski and W. Stummer, “Some universal insights on divergences for statistics, machine learning and artificial intelligence,” in Geometric Structures of Information , F. Nielsen, Ed. Cham, Switzerland: Springer International Publishing, 2019, pp. 149–211
2019
Closest in time.
Z. Goldfeld, E. Van Den Berg, K. Greenewald, I. Melnyk, N. Nguyen, B. Kingsbury, and Y. Polyanskiy, “Estimating information flow in deep neural networks,” in Proc. Int. Conf. Mach. Learn. (ICML) , vol. 97, Long Beach, CA, June 2019, pp. 2299–2308
2019
Closest in time.
R. A. Horn and C. R. Johnson, Matrix Analysis , 2nd ed. Cambridge, UK: Cambridge Uiversity Press, 2012
2019
Closest in time.