Fetching the paper…
Reading the bibliography…
This paper studies Fenchel-Young losses, a generic way to construct convex loss functions from a regularization function.
Variabilità e mutabilità
Corrado Gini · 1912
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier · 1950
Earlier work this paper cites.
The perceptron: A probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
Uncertainty, information, and sequential experiments
Morris H DeGroot · 1962
Earlier work this paper cites.
Pseudo-convex functions
Olvi L Mangasarian · 1965
Earlier work this paper cites.
The theory of max-min, with applications
John M Danskin · 1966
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Lev M Bregman · 1967
Earlier work this paper cites.
Diversity of planktonic foraminifera in deep-sea sediments
Wolfgang H Berger and Frances L Parker · 1970
Earlier work this paper cites.
Convex Analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Leonard J Savage · 1971
Earlier work this paper cites.
Generalized Linear Models
John Ashworth Nelder and R Jacob Baker · 1972
Earlier work this paper cites.
An O ( n ) O(n) algorithm for quadratic knapsack problems
Peter Brucker · 1984
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Sharp uniform convexity and smoothness inequalities for trace norms
Keith Ball, Eric A Carlen, and Elliott H Lieb · 1994
Earlier work this paper cites.
Statistical Learning Theory
Vladimir N Vapnik · 1998
Earlier work this paper cites.
Nonlinear Programming
Dimitri P Bertsekas · 1999
Cited alongside, same era.
On the algorithmic implementation of multiclass kernel-based vector machines
Koby Crammer and Yoram Singer · 2001
Cited alongside, same era.
Learning with Kernels
Bernhard Schölkopf and Alexander J Smola · 2002
Cited alongside, same era.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Nonextensive Entropy: Interdisciplinary Applications
Murray Gell-Mann and Constantino Tsallis · 2004
Cited alongside, same era.
Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory
Peter D Grünwald and A Philip Dawid · 2004
Cited alongside, same era.
Composite binary losses
Mark D Reid and Robert C Williamson · 2010
Later among the works it cites.
The design of bayes consistent loss functions for classification
Hamed Masnadi-Shirazi · 2011
Later among the works it cites.
Smoothing and first order methods: A unified framework
Amir Beck and Marc Teboulle · 2012
Later among the works it cites.
Efficient Bregman projections onto the simplex
Walid Krichene, Syrine Krichene, and Alexandre Bayen · 2015
Later among the works it cites.
Information Geometry and Its Applications
Shun-ichi Amari · 2016
Later among the works it cites.
Fast projection onto the simplex and the ℓ 1 \ell_{1} ball
Laurent Condat · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization of Shannon-Khinchin axioms to nonextensive systems and the uniqueness theorem for the nonextensive entropy
Hiroki Suyari · 2004
Cited alongside, same era.
Clustering with bregman divergences
Arindam Banerjee, Srujana Merugu, Inderjit S Dhillon, and Joydeep Ghosh · 2005
Cited alongside, same era.
Loss functions for binary class probability estimation and classification: Structure and applications
Andreas Buja, Werner Stuetzle, and Yi Shen · 2005
Cited alongside, same era.
Smooth minimization of non-smooth functions
Yurii Nesterov · 2005
Cited alongside, same era.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Cited alongside, same era.
VC theory of large margin multi-category classifiers
Yann Guermeur · 2007
Cited alongside, same era.
André FT Martins and Ramón Fernandez Astudillo · 2016
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2016
Later among the works it cites.
Robert C Williamson, Elodie Vernet, and Mark D Reid · 2016
Later among the works it cites.
Two-temperature logistic regression based on the Tsallis divergence
Ehsan Amid and Manfred K Warmuth · 2017
Later among the works it cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Heinz H Bauschke and Patrick L Combettes · 2017
Later among the works it cites.
A regularized framework for sparse and structured neural attention
Vlad Niculae and Mathieu Blondel · 2017
Later among the works it cites.
Multiclass classification, information, divergence, and surrogate risk
John C Duchi, Khashayar Khosravi, and Feng Ruan · 2018
Closest in time.
Differentiable dynamic programming for structured prediction and attention
Arthur Mensch and Mathieu Blondel · 2018
Closest in time.
SparseMAP: Differentiable sparse structured inference
Vlad Niculae, André FT Martins, Mathieu Blondel, and Claire Cardie · 2018
Closest in time.
Learning with fenchel-young losses
Mathieu Blondel, André FT Martins, and Vlad Niculae · 2019
Closest in time.