Fetching the paper…
Reading the bibliography…
Reward augmented maximum likelihood (RAML), a simple and effective learning framework to directly optimize towards the reward function in structured prediction tasks, has led to a number of impressive empirical successes.
Building a large annotated corpus of English: the Penn Treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Rich Caruana, Steve Lawrence, and Giles Lee · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, Fernando Pereira, et al · 2001
Earlier work this paper cites.
Cubic-time parsing and learning algorithms for grammatical bigram models
Mark A Paskin · 2001
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Sang Tjong Kim, Erik F., and Fien De Meulder · 2003
Earlier work this paper cites.
Minimum bayes-risk decoding for statistical machine translation
Shankar Kumar and William Byrne · 2004
Earlier work this paper cites.
Max-margin markov networks
Ben Taskar, Carlos Guestrin, and Daphne Koller · 2004
Earlier work this paper cites.
Conditional random fields: An introduction
Hanna M Wallach · 2004
Earlier work this paper cites.
Incorporating non-local information into information extraction systems by gibbs sampling
Jenny Rose Finkel, Trond Grenager, and Christopher Manning · 2005
Earlier work this paper cites.
Online large-margin training of dependency parsers
Ryan McDonald, Koby Crammer, and Fernando Pereira · 2005
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
Marie-Catherine De Marneffe, Bill MacCartney, Christopher D Manning, et al · 2006
Earlier work this paper cites.
Softmax-margin CRFs: Training log-linear models with cost functions
Kevin Gimpel and Noah A. Smith · 2010
Cited alongside, same era.
Loss-sensitive training of probabilistic conditional random fields
Maksims N Volkovs, Hugo Larochelle, and Richard S Zemel · 2011
Cited alongside, same era.
Discriminative Feature-Rich Modeling for Syntax-Based Machine Translation
K. Gimpel · 2012
Cited alongside, same era.
Probabilistic models for high-order projective dependency parsing
Xuezhe Ma and Hai Zhao · 2012
Cited alongside, same era.
All of statistics: a concise course in statistical inference
Larry Wasserman · 2013
Cited alongside, same era.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Later among the works it cites.
Optimization of image description metrics using policy gradient methods
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2016
Later among the works it cites.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy · 2016
Later among the works it cites.
Improving policy gradient by exploring under-appreciated rewards
Ofir Nachum, Mohammad Norouzi, and Dale Schuurmans · 2016
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mauro Cettolo, Jan Niehues, Sebastian Stuker, Luisa Bentivogli, and Marcello Federico · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Transition-based dependency parsing with stack long short-term memory
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A. Smith · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Later among the works it cites.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu · 2016
Later among the works it cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Later among the works it cites.
Review networks for caption generation
Zhilin Yang, Ye Yuan, Yuexin Wu, William W Cohen, and Ruslan R Salakhutdinov · 2016
Later among the works it cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Closest in time.
Learning to decode for future success
Jiwei Li, Will Monroe, and Dan Jurafsky · 2017
Closest in time.