Fetching the paper…
Reading the bibliography…
Over the past decades, numerous loss functions have been been proposed for a variety of supervised learning tasks, including regression, classification, ranking, and more generally structured prediction.
Variabilità e mutabilità
Corrado Gini · 1912
Earlier work this paper cites.
Tres observaciones sobre el algebra lineal
Garrett Birkhoff · 1946
Earlier work this paper cites.
The Mathematical Theory of Communication
Claude E Shannon and Warren Weaver · 1949
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier · 1950
Earlier work this paper cites.
The generalized simplex method for minimizing a linear form under linear inequality restraints
George B Dantzig, Alex Orden, and Philip Wolfe · 1955
Earlier work this paper cites.
The Hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
Uncertainty, information, and sequential experiments
Morris H DeGroot · 1962
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1964
Earlier work this paper cites.
On the shortest arborescence of a directed graph
Yoeng-Jin Chu and Tseng-Hong Liu · 1965
Earlier work this paper cites.
Pseudo-convex functions
Olvi L Mangasarian · 1965
Earlier work this paper cites.
Proximité et dualité dans un espace hilbertien
Jean-Jacques Moreau · 1965
Earlier work this paper cites.
Statistical inference for probabilistic functions of finite state Markov chains
Leonard E. Baum and Ted Petrie · 1966
Earlier work this paper cites.
The theory of max-min, with applications
John M Danskin · 1966
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Lev M Bregman · 1967
Earlier work this paper cites.
Optimum branchings
Jack Edmonds · 1967
Earlier work this paper cites.
Concerning nonnegative matrices and doubly stochastic matrices
Richard Sinkhorn and Paul Knopp · 1967
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
Andrew Viterbi · 1967
Earlier work this paper cites.
Diversity of planktonic foraminifera in deep-sea sediments
Wolfgang H Berger and Frances L Parker · 1970
Earlier work this paper cites.
Convex Analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Leonard J Savage · 1971
Earlier work this paper cites.
Generalized Linear Models
John Ashworth Nelder and R Jacob Baker · 1972
Earlier work this paper cites.
I-divergence geometry of probability distributions and minimization problems
Imre Csiszár · 1975
Earlier work this paper cites.
Finding the nearest point in a polytope
Philip Wolfe · 1976
Earlier work this paper cites.
Finding optimum branchings
Robert E Tarjan · 1977
Earlier work this paper cites.
Information and Exponential Families: In Statistical Theory
Ole Barndorff-Nielsen · 1978
Earlier work this paper cites.
Conditional gradient algorithms with open loop step size rules
Joseph C Dunn and S Harshbarger · 1978
Earlier work this paper cites.
Dynamic programming algorithm optimization for spoken word recognition
Hiroaki Sakoe and Seibi Chiba · 1978
Earlier work this paper cites.
The complexity of computing the permanent
Leslie G Valiant · 1979
Earlier work this paper cites.
An O ( n ) O(n) algorithm for quadratic knapsack problems
Peter Brucker · 1984
Earlier work this paper cites.
A shortest augmenting path algorithm for dense and sparse linear assignment problems
Roy Jonker and Anton Volgenant · 1987
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
On ordered weighted averaging aggregation operators in multicriteria decisionmaking
Ronald R Yager · 1988
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
Generalized Linear Models , volume 37
Peter McCullagh and John A Nelder · 1989
Earlier work this paper cites.
Sharp uniform convexity and smoothness inequalities for trace norms
Keith Ball, Eric A Carlen, and Elliott H Lieb · 1994
Earlier work this paper cites.
Statistical Learning Theory
Vladimir Vapnik · 1998
Earlier work this paper cites.
Nonlinear Programming
Dimitri P Bertsekas · 1999
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen Wright · 1999
Earlier work this paper cites.
On the algorithmic implementation of multiclass kernel-based vector machines
Koby Crammer and Yoram Singer · 2001
Earlier work this paper cites.
Conditional Random Fields: Probabilistic models for segmenting and labeling sequence data
John D Lafferty, Andrew McCallum, and Fernando CN Pereira · 2001
Earlier work this paper cites.
Discriminative training methods for Hidden Markov Models: Theory and experiments with perceptron algorithms
Michael Collins · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims · 2002
Earlier work this paper cites.
Learning With Kernels
Bernhard Schölkopf and Alexander J Smola · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Nonextensive Entropy: Interdisciplinary Applications
Murray Gell-Mann and Constantino Tsallis · 2004
Earlier work this paper cites.
Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory
Peter D Grünwald and A Philip Dawid · 2004
Earlier work this paper cites.
In defense of one-vs-all classification
Ryan M Rifkin and Aldebaro Klautau · 2004
Cited alongside, same era.
Generalization of Shannon-Khinchin axioms to nonextensive systems and the uniqueness theorem for the nonextensive entropy
Hiroki Suyari · 2004
Cited alongside, same era.
Learning Structured Prediction Models: A Large Margin Approach
Ben Taskar · 2004
Cited alongside, same era.
Statistical behavior and consistency of classification methods based on convex risk minimization
Tong Zhang · 2004
Cited alongside, same era.
Clustering with Bregman divergences
Arindam Banerjee, Srujana Merugu, Inderjit S Dhillon, and Joydeep Ghosh · 2005
Cited alongside, same era.
Loss functions for binary class probability estimation and classification: Structure and applications
Andreas Buja, Werner Stuetzle, and Yi Shen · 2005
A primal-dual convergence analysis of boosting
Matus Telgarsky · 2012
Later among the works it cites.
Marginal inference in MRFs using Frank-Wolfe
David Belanger, Dan Sheldon, and Andrew McCallum · 2013
Later among the works it cites.
API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Gaël Varoquaux · 2013
Later among the works it cites.
Sinkhorn distances: Lightspeed computation of optimal transportation distances
Marco Cuturi · 2013
Later among the works it cites.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
Martin Jaggi · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Non-projective dependency parsing using spanning tree algorithms
Ryan T McDonald, Fernando CN Pereira, Kiril Ribarov, and Jan Hajič · 2005
Cited alongside, same era.
Smooth minimization of non-smooth functions
Yurii Nesterov · 2005
Cited alongside, same era.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun · 2005
Cited alongside, same era.
Convexity, classification, and risk bounds
Peter L Bartlett, Michael I Jordan, and Jon D McAuliffe · 2006
Cited alongside, same era.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Cited alongside, same era.
VC theory of large margin multi-category classifiers
Yann Guermeur · 2007
Cited alongside, same era.
Later among the works it cites.
Sparse projections onto the simplex
Anastasios Kyrillidis, Stephen Becker, Volkan Cevher, and Christoph Koch · 2013
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Later among the works it cites.
Convex foundations for generalized MaxEnt models
Rafael Frongillo and Mark D Reid · 2014
Later among the works it cites.
Metric learning for temporal sequence alignment
Damien Garreau, Rémi Lajugie, Sylvain Arlot, and Francis Bach · 2014
Later among the works it cites.
Orbit regularization
Renato Negrinho and Andre Martins · 2014
Later among the works it cites.
On learning to localize objects with minimal supervision
Hyun Oh Song, Ross Girshick, Stefanie Jegelka, Julien Mairal, Zaid Harchaoui, and Trevor Darrell · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Later among the works it cites.
Efficient Bregman projections onto the simplex
Walid Krichene, Syrine Krichene, and Alexandre Bayen · 2015
Later among the works it cites.
Barrier Frank-Wolfe for marginal inference
Rahul G Krishnan, Simon Lacoste-Julien, and David Sontag · 2015
Later among the works it cites.
On the global linear convergence of Frank-Wolfe optimization variants
Simon Lacoste-Julien and Martin Jaggi · 2015
Later among the works it cites.
Sequential kernel herding: Frank-Wolfe optimization for particle filtering
Simon Lacoste-Julien, Fredrik Lindsten, and Francis Bach · 2015
Later among the works it cites.
Maksim Lapin, Matthias Hein, and Bernt Schiele · 2015
Later among the works it cites.
The ordered weighted ℓ 1 \ell_{1} norm: Atomic formulation and conditional gradient algorithm
Xiangrong Zeng and Mário AT Figueiredo · 2015
Later among the works it cites.
Bandit online optimization over the permutahedron
Nir Ailon, Kohei Hatano, and Eiji Takimoto · 2016
Later among the works it cites.
Information Geometry and Its Applications
Shun-ichi Amari · 2016
Later among the works it cites.
Fast projection onto the simplex and the ℓ 1 \ell_{1} ball
Laurent Condat · 2016
Later among the works it cites.
Regularized optimal transport and the Rot Mover’s Distance
Arnaud Dessein, Nicolas Papadakis, and Jean-Luc Rouas · 2016
Later among the works it cites.
Inside-outside and forward-backward algorithms are just backprop (tutorial paper)
Jason Eisner · 2016
Later among the works it cites.
Simple and accurate dependency parsing using bidirectional LSTM feature representations
Eliyahu Kiperwasser and Yoav Goldberg · 2016
Later among the works it cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer · 2016
Later among the works it cites.
Efficient bregman projections onto the permutahedron and related polytopes
Cong Han Lim and Stephen J Wright · 2016
Later among the works it cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André FT Martins and Ramón Fernandez Astudillo · 2016
Later among the works it cites.
Universal Dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D Manning, Ryan T McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, et al · 2016
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2016
Later among the works it cites.
Composite multiclass losses
Robert C Williamson, Elodie Vernet, and Mark D Reid · 2016
Later among the works it cites.
Two-temperature logistic regression based on the Tsallis divergence
Ehsan Amid and Manfred K Warmuth · 2017
Later among the works it cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Heinz H Bauschke and Patrick L Combettes · 2017
Later among the works it cites.
Soft-DTW: A differentiable loss function for time-series
Marco Cuturi and Mathieu Blondel · 2017
Later among the works it cites.
Greedy algorithms for cone constrained optimization with convergence guarantees
Francesco Locatello, Michael Tschannen, Gunnar Rätsch, and Martin Jaggi · 2017
Later among the works it cites.
A regularized framework for sparse and structured neural attention
Vlad Niculae and Mathieu Blondel · 2017
Later among the works it cites.
Computational Optimal Transport
Gabriel Peyré and Marco Cuturi · 2017
Later among the works it cites.
Tokenizing, POS tagging, lemmatizing and parsing UD 2.0 with UDPipe
Milan Straka and Jana Straková · 2017
Later among the works it cites.
Fast column generation for atomic norm regularization
Marina Vinyes and Guillaume Obozinski · 2017
Later among the works it cites.
CoNLL 2017 shared task: Multilingual parsing from raw text to universal dependencies
Daniel Zeman, Martin Popel, Milan Straka, Jan Hajic, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, et al · 2017
Later among the works it cites.
Smooth and sparse optimal transport
Mathieu Blondel, Vivien Seguy, and Antoine Rolet · 2018
Later among the works it cites.
Multiclass classification, information, divergence, and surrogate risk
John C Duchi, Khashayar Khosravi, and Feng Ruan · 2018
Later among the works it cites.
Speech and Language Processing (3rd ed.)
Dan Jurafsky and James H Martin · 2018
Later among the works it cites.
Differentiable dynamic programming for structured prediction and attention
Arthur Mensch and Mathieu Blondel · 2018
Later among the works it cites.
SparseMAP: Differentiable sparse structured inference
Vlad Niculae, André FT Martins, Mathieu Blondel, and Claire Cardie · 2018
Later among the works it cites.
A Smoother Way to Train Structured Prediction Models
Venkata Krishna Pillutla, Vincent Roulet, Sham M Kakade, and Zaid Harchaoui · 2018
Later among the works it cites.
Structured prediction with projection oracles
Mathieu Blondel · 2019
Closest in time.
Learning classifiers with Fenchel-Young losses: Generalized entropies, margins, and algorithms
Mathieu Blondel, André FT Martins, and Vlad Niculae · 2019
Closest in time.
A general theory for structured prediction with smooth convex surrogates
Alex Nowak-Vila, Francis Bach, and Alessandro Rudi · 2019
Closest in time.