Fetching the paper…
Reading the bibliography…
We present a framework to train a structured prediction model by performing smoothing on the inference algorithm it builds upon.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
A. J. Viterbi · 1967
Earlier work this paper cites.
Validation of subgradient optimization
M. Held, P. Wolfe, and H. P. Crowder · 1974
Earlier work this paper cites.
Syntactic analysis of two-dimensional visual signals in noisy conditions
M. I. Schlesinger · 1976
Earlier work this paper cites.
Submodular functions and convexity
L. Lovász · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Descent methods for composite nondifferentiable optimization problems
J. V. Burke · 1985
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems: networks of plausible inference
J. Pearl · 1988
Earlier work this paper cites.
Exact maximum a posteriori estimation for binary images
D. M. Greig, B. T. Porteous, and A. H. Seheult · 1989
Earlier work this paper cites.
A Framework for the Cooperation of Learning Algorithms
L. Bottou and P. Gallinari · 1990
Earlier work this paper cites.
The computational complexity of probabilistic inference using bayesian belief networks
G. F. Cooper · 1990
Earlier work this paper cites.
Applications of a general propagation algorithm for probabilistic expert systems
A. P. Dawid · 1992
Earlier work this paper cites.
Polynomial-time approximation algorithms for the Ising model
M. Jerrum and A. Sinclair · 1993
Earlier work this paper cites.
An algorithm directly finding the K K most probable configurations in Bayesian networks
B. Seroussi and J. Golmard · 1994
Earlier work this paper cites.
LeRec: A NN/HMM Hybrid for On-Line Handwriting Recognition
Y. Bengio, Y. LeCun, C. Nohl, and C. Burges · 1995
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
D. P. Bertsekas · 1995
Earlier work this paper cites.
Maximum-weight bipartite matching technique and its application in image feature matching
Y.-Q. Cheng, V. Wu, R. Collins, A. R. Hanson, and E. M. Riseman · 1996
Earlier work this paper cites.
Global Training of Document Processing Systems Using Graph Transformer Networks
L. Bottou, Y. Bengio, and Y. LeCun · 1997
Earlier work this paper cites.
Segmentation by grouping junctions
H. Ishikawa and D. Geiger · 1998
Earlier work this paper cites.
Turbo Decoding as an Instance of Pearl’s ”Belief Propagation” Algorithm
R. J. McEliece, D. J. C. MacKay, and J. Cheng · 1998
Earlier work this paper cites.
An efficient algorithm for finding the M M most probable configurations in probabilistic expert systems
D. Nilsson · 1998
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Loopy belief propagation for approximate inference: An empirical study
K. P. Murphy, Y. Weiss, and M. I. Jordan · 1999
Earlier work this paper cites.
The OpenCV Library
G. Bradski · 2000
Earlier work this paper cites.
On the algorithmic implementation of multiclass kernel-based vector machines
K. Crammer and Y. Singer · 2001
Earlier work this paper cites.
Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data
J. Lafferty, A. McCallum, and F. C. Pereira · 2001
Earlier work this paper cites.
Hidden Markov Support Vector Machines
Y. Altun, I. Tsochantaridis, and T. Hofmann · 2003
Earlier work this paper cites.
Combinatorial Optimization - Polyhedra and Efficiency
A. Schrijver · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
E. F. Tjong Kim Sang and F. De Meulder · 2003
Earlier work this paper cites.
What energy functions can be minimized via graph cuts?
V. Kolmogorov and R. Zabin · 2004
Earlier work this paper cites.
Max-margin Markov networks
B. Taskar, C. Guestrin, and D. Koller · 2004
Earlier work this paper cites.
Support vector machine learning for interdependent and structured output spaces
I. Tsochantaridis, T. Hofmann, T. Joachims, and Y. Altun · 2004
Earlier work this paper cites.
Finding the M M most probable configurations using loopy belief propagation
C. Yanover and Y. Weiss · 2004
Earlier work this paper cites.
Learning as search optimization: approximate large margin methods for structured prediction
H. Daumé III and D. Marcu · 2005
Earlier work this paper cites.
A discriminative matching approach to word alignment
B. Taskar, S. Lacoste-Julien, and D. Klein · 2005
Cited alongside, same era.
MAP estimation via agreement on trees: message-passing and linear programming
M. J. Wainwright, T. S. Jaakkola, and A. S. Willsky · 2005
Cited alongside, same era.
Using Combinatorial Optimization within Max-Product Belief Propagation
J. C. Duchi, D. Tarlow, G. Elidan, and D. Koller · 2006
Cited alongside, same era.
Structured prediction, dual extragradient and Bregman projections
B. Taskar, S. Lacoste-Julien, and M. I. Jordan · 2006
Cited alongside, same era.
(Approximate) Subgradient Methods for Structured Prediction
N. D. Ratliff, J. A. Bagnell, and M. Zinkevich · 2007
Cited alongside, same era.
Exponentiated gradient algorithms for conditional random fields and max-margin markov networks
M. Collins, A. Globerson, T. Koo, X. Carreras, and P. L. Bartlett · 2008
Block-Coordinate Frank-Wolfe Optimization for Structural SVMs
S. Lacoste-Julien, M. Jaggi, M. Schmidt, and P. Pletscher · 2013
Later among the works it cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Later among the works it cites.
Dual subgradient algorithms for large-scale nonsmooth learning problems
B. Cox, A. Juditsky, and A. Nemirovski · 2014
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convex relaxation methods for graphical models: Lagrangian and maximum entropy approaches
J. K. Johnson · 2008
Cited alongside, same era.
Beyond sliding windows: Object localization by efficient subwindow search
C. H. Lampert, M. B. Blaschko, and T. Hofmann · 2008
Cited alongside, same era.
Graphical models, exponential families, and variational inference
M. J. Wainwright and M. I. Jordan · 2008
Cited alongside, same era.
An LP view of the M M -best MAP problem
M. Fromer and A. Globerson · 2009
Cited alongside, same era.
Cutting-plane training of structural SVMs
T. Joachims, T. Finley, and C.-N. J. Yu · 2009
Cited alongside, same era.
Probabilistic Graphical Models - Principles and Techniques
D. Koller and N. Friedman · 2009
Cited alongside, same era.
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Later among the works it cites.
Speech and Language Processing
D. Jurafsky, J. H. Martin, P. Norvig, and S. Russell · 2014
Later among the works it cites.
A* CCG parsing with a supertag-factored model
M. Lewis and M. Steedman · 2014
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2014
Later among the works it cites.
On learning to localize objects with minimal supervision
H. O. Song, R. B. Girshick, S. Jegelka, J. Mairal, Z. Harchaoui, and T. Darrell · 2014
Later among the works it cites.
Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization
R. Frostig, R. Ge, S. Kakade, and A. Sidford · 2015
Later among the works it cites.
Semi-Proximal Mirror-Prox for Nonsmooth Composite Minimization
N. He and Z. Harchaoui · 2015
Later among the works it cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
Incremental majorization-minimization optimization with application to large-scale machine learning
J. Mairal · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
Non-uniform stochastic average gradient method for training conditional random fields
M. Schmidt, R. Babanezhad, M. Ahmed, A. Defazio, A. Clifton, and A. Sarkar · 2015
Later among the works it cites.
Structured prediction energy networks
D. Belanger and A. McCallum · 2016
Later among the works it cites.
A simple practical accelerated method for finite sums
A. Defazio · 2016
Later among the works it cites.
Searching for the M M Best Solutions in Graphical Models
N. Flerova, R. Marinescu, and R. Dechter · 2016
Later among the works it cites.
Blending Learning and Inference in Conditional Random Fields
T. Hazan, A. G. Schwing, and R. Urtasun · 2016
Later among the works it cites.
From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
A. F. T. Martins and R. F. Astudillo · 2016
Later among the works it cites.
Minding the gaps for block Frank-Wolfe optimization of structured SVMs
A. Osokin, J.-B. Alayrac, I. Lukasewitz, P. Dokania, and S. Lacoste-Julien · 2016
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
B. Palaniappan and F. Bach · 2016
Later among the works it cites.
Tight complexity bounds for optimizing composite objectives
B. E. Woodworth and N. Srebro · 2016
Later among the works it cites.
Katyusha: The First Direct Acceleration of Stochastic Gradient Methods
Z. Allen-Zhu · 2017
Later among the works it cites.
Deep Semantic Role Labeling: What Works and What’s Next
L. He, K. Lee, M. Lewis, and L. Zettlemoyer · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
D. Davis and D. Drusvyatskiy · 2018
Later among the works it cites.
Efficiency of minimizing compositions of convex functions and smooth maps
D. Drusvyatskiy and C. Paquette · 2018
Later among the works it cites.
Catalyst Acceleration for First-order Convex Optimization: from Theory to Practice
H. Lin, J. Mairal, and Z. Harchaoui · 2018
Later among the works it cites.
Differentiable dynamic programming for structured prediction and attention
A. Mensch and M. Blondel · 2018
Later among the works it cites.
SparseMAP: Differentiable Sparse Structured Inference
V. Niculae, A. F. Martins, M. Blondel, and C. Cardie · 2018
Later among the works it cites.
Catalyst for gradient-based nonconvex optimization
C. Paquette, H. Lin, D. Drusvyatskiy, J. Mairal, and Z. Harchaoui · 2018
Later among the works it cites.