Fetching the paper…
Reading the bibliography…
We propose SEARNN, a novel training algorithm for recurrent neural networks (RNNs) inspired by the "learning to search" (L2S) approach to structured prediction.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Max-margin Markov networks
B. Taskar, C. Guestrin, and D. Koller · 2003
Earlier work this paper cites.
Multicategory support vector machines: Theory and application to the classification of microarray data and satellite radiance data
Yoonkyung Lee, Yi Lin, and Grace Wahba · 2004
Earlier work this paper cites.
Learning as search optimization: approximate large margin methods for structured prediction
Hal Daumé, III and Daniel Marcu · 2005
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun · 2005
Earlier work this paper cites.
Lower bounds for reductions
Matti Kääriäinen · 2006
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, III, John Langford, and Daniel Marcu · 2009
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M Rush · 2009
Earlier work this paper cites.
Softmax-margin CRFs: Training loglinear models with cost functions
Kevin Gimpel and Noah A Smith · 2010
Earlier work this paper cites.
A primal-dual message-passing algorithm for approximated large scale structured prediction
Tamir Hazan and Raquel Urtasun · 2010
Earlier work this paper cites.
Entropy and margin maximization for structured output learning
Patrick Pletscher, Cheng Soon Ong, and Joachim M. Buhmann · 2010
Cited alongside, same era.
A dynamic oracle for arc-eager dependency parsing
Yoav Golberg and Joakim Nivre · 2012
Cited alongside, same era.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stuker, Luisa Bentivogli, and Marcello Federico · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J. Andrew Bagnell · 2014
Cited alongside, same era.
Training with exploration improves a greedy stack-LSTM parser
Miguel Ballesteros, Yoav Goldberg, Chris Dyer, and Noah A Smith · 2016
Later among the works it cites.
Learning reductions that really work
Alina Beygelzimer, Hal Daumé, III, John Langford, and Paul Mineiro · 2016
Later among the works it cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Later among the works it cites.
Noise reduction and targeted exploration in imitation learning for abstract meaning representation parsing
James Goodman, Andreas Vlachos, and Jason Naradowsky · 2016
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Sammy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Later among the works it cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Cited alongside, same era.
Learning to search better than your teacher
Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, Hal Daumé, III, and John Langford · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2015
Cited alongside, same era.
A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Later among the works it cites.
Self-critical sequence training for image captioning
Steven Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel · 2016
Later among the works it cites.
Minimum risk training for neural machine translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Later among the works it cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Closest in time.
Deeply AggreVaTeD: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J. Gordon, Byron Boots, and J. Andrew Bagnell · 2017
Closest in time.