Fetching the paper…
Reading the bibliography…
We apply stochastic average gradient (SAG) algorithms for training conditional random fields (CRFs).
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
A maximum entropy model for part-of-speech tagging
A. Ratnaparkhi · 1996
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
A. Nedic and D. Bertsekas · 2000
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
J. Lafferty, A. McCallum, and F. Pereira · 2001
Earlier work this paper cites.
Discriminative training methods for hidden Markov models: theory and experiments with perceptron algorithms
M. Collins · 2002
Earlier work this paper cites.
Efficient training of conditional random fields
H. Wallach · 2002
Earlier work this paper cites.
Dynamic conditional random fields for jointly labeling multiple sequences
A. McCallum, K. Rohanimanesh, and C. Sutton · 2003
Earlier work this paper cites.
Shallow parsing with conditional random fields
F. Sha and F. Pereira · 2003
Earlier work this paper cites.
Max-margin Markov networks
B. Taskar, C. Guestrin, and D. Koller · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2004
Earlier work this paper cites.
Biomedical named entity recognition using conditional random fields and rich feature sets
B. Settles · 2004
Earlier work this paper cites.
Semantic role labelling with tree conditional random fields
T. Cohn and P. Blunsom · 2005
Earlier work this paper cites.
minFunc: unconstrained multivariate differentiable optimization in Matlab, 2005
M. Schmidt · 2005
Earlier work this paper cites.
Information extraction from research papers using conditional random fields
F. Peng and A. McCallum · 2006
Earlier work this paper cites.
Accelerated training of conditional random fields with stochastic gradient methods
S. Vishwanathan, N. N. Schraudolph, M. W. Schmidt, and K. P. Murphy · 2006
Cited alongside, same era.
Exponentiated gradient algorithms for conditional random fields and max-margin Markov networks
M. Collins, A. Globerson, T. Koo, X. Carreras, and P. Bartlett · 2008
Cited alongside, same era.
Efficient, feature-based, conditional random field parsing
J. R. Finkel, A. Kleeman, and C. D. Manning · 2008
Cited alongside, same era.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
A. Beck and M. Teboulle · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
A randomized Kaczmarz algorithm with exponential convergence
Pegasos: primal estimated sub-gradient solver for svm
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter · 2011
Later among the works it cites.
A fast accurate two-stage training algorithm for L1-regularized CRFs with heuristic line search strategy
J. Zhou, X. Qiu, and X. Huang · 2011
Later among the works it cites.
Hybrid deterministic-stochastic methods for data fitting
M. P. Friedlander and M. Schmidt · 2012
Later among the works it cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
S. Ghadimi and G. Lan · 2012
Later among the works it cites.
A stochastic gradient method with an exponential convergence rate for strongly-convex optimization with finite training sets
N. Le Roux, M. Schmidt, and F. Bach · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Strohmer and R. Vershynin · 2009
Cited alongside, same era.
Stochastic gradient descent training for L1-regularized log-linear models with cumulative penalty
Y. Tsuruoka, J. Tsujii, and S. Ananiadou · 2009
Cited alongside, same era.
Practical very large scale CRFs
T. Lavergne, O. Cappé, and F. Yvon · 2010
Cited alongside, same era.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
W. Xu · 2010
Cited alongside, same era.
Conditional Topic Random Fields
J. Zhu and E. Xing · 2010
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Later among the works it cites.
Block-coordinate frank-wolfe optimization for structural svms
S. Lacoste-Julien, M. Jaggi, M. Schmidt, and P. Pletscher · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2013
Later among the works it cites.
Linear convergence with condition number independent access of full gradients
L. Zhang, M. Mahdavi, and R. Jin · 2013
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2014
Later among the works it cites.
A proximal stochastic gradient method with progressive variance reduction
L. Xiao and T. Zhang · 2014
Later among the works it cites.