Fetching the paper…
Reading the bibliography…
A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Conditional Random Fields: Probabilistic models for segmenting and labeling sequence data
J. D. Lafferty, A. McCallum, and F. C. N. Pereira · 2001
Earlier work this paper cites.
Max-margin markov networks
B. Taskar, C. Guestrin, and D. Koller · 2004
Earlier work this paper cites.
Clustering with Bregman divergences
A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh · 2005
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
I. Tsochantaridis, T. Joachims, T. Hofmann, and Y. Altun · 2005
Earlier work this paper cites.
Linearly-solvable markov decision problems
E. Todorov · 2006
Earlier work this paper cites.
Search-based structured prediction
H. Daumé, III, J. Langford, and D. Marcu · 2009
Earlier work this paper cites.
Learning model-free robot control by a Monte Carlo EM algorithm
N. Vlassis, M. Toussaint, G. Kontes, and S. Piperidis · 2009
Earlier work this paper cites.
Softmax-margin crfs: Training log-linear models with cost functions
K. Gimpel and N. A. Smith · 2010
Earlier work this paper cites.
Relative entropy policy search
J. Peters, K. Mülling, and Y. Altün · 2010
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
D. Povey, A. Ghoshal, G. Boulianne, et al · 2011
Cited alongside, same era.
Empirical risk minimization of graphical model parameters given approximate inference, decoding, and model structure
V. Stoyanov, A. Ropson, and J. Eisner · 2011
Cited alongside, same era.
Loss-sensitive training of probabilistic conditional random fields
M. Volkovs, H. Larochelle, and R. Zemel · 2011
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice
T. Degris, P. M. Pilarski, and R. S. Sutton · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
J. Domke · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
H. J. Kappen, V. Gómez, and M. Opper · 2012
Effective approaches to attention-based neural machine translation
M.-T. Luong, H. Pham, and C. D. Manning · 2015
Later among the works it cites.
Addressing the rare word problem in neural machine translation
M.-T. Luong, I. Sutskever, Q. V. Le, O. Vinyals, and W. Zaremba · 2015
Later among the works it cites.
Learning dynamic feature selection for fast sequential prediction
E. Strubell, L. Vilnis, K. Silverstein, and A. McCallum · 2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2015
Later among the works it cites.
Learning using privileged information: Similarity control and knowledge transfer
V. Vapnik and R. Izmailov · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning from limited demonstrations
B. Kim, A. M. Farahmand, J. Pineau, and D. Precup · 2013
Cited alongside, same era.
Guided policy search
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Variational policy search via trajectory optimization
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
End-to-end continuous speech recognition using attention-based recurrent nn: first results
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
D. Andor, C. Alberti, D. Weiss, A. Severyn, A. Presta, K. Ganchev, S. Petrov, and M. Collins · 2016
Closest in time.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Closest in time.
Listen, attend and spell
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2016
Closest in time.
Ask me anything: Dynamic memory networks for natural language processing
A. Kumar, O. Irsoy, J. Su, J. Bradbury, R. English, B. Pierce, P. Ondruska, I. Gulrajani, and R. Socher · 2016
Closest in time.
Unifying distillation and privileged information
D. Lopez-Paz, B. Schölkopf, L. Bottou, and V. Vapnik · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Closest in time.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2016
Closest in time.
Minimum risk training for neural machine translation
S. Shen, Y. Cheng, Z. He, W. He, H. Wu, M. Sun, and Y. Liu · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
D. Silver et al · 2016
Closest in time.
Sequence-to-sequence learning as beam-search optimization
S. Wiseman and A. M. Rush · 2016
Closest in time.