Fetching the paper…
Reading the bibliography…
There has been much recent work on training neural attention models at the sequence-level using either reinforcement learning-style methods or by optimizing the beam.
Dropout: a simple way to prevent Neural Networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams. 1992 · 1992
Earlier work this paper cites.
Support vector learning for ordinal regression
Ralf Herbrich, Thore Graepel, and Klaus Obermayer. 1999 · 1999
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003 · 2003
Earlier work this paper cites.
Minimum Error Rate Training in Statistical Machine Translation
Franz Josef Och. 2003 · 2003
Earlier work this paper cites.
Max-margin markov networks
Ben Taskar, Carlos Guestrin, and Daphne Koller. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Orange: a method for evaluating automatic evaluation metrics for machine translation
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun. 2005 · 2005
Earlier work this paper cites.
Minimum Risk Annealing for Training Log-Linear Models
David A. Smith and Jason Eisner. 2006 · 2006
Earlier work this paper cites.
Minimum error rate training by sampling the translation lattice
Samidh Chatterjee and Nicola Cancedda. 2010 · 2010
Earlier work this paper cites.
Softmax-margin crfs: Training log-linear models with cost functions
Kevin Gimpel and Noah Smith. 2010 · 2010
Earlier work this paper cites.
BBN System Description for WMT10 System Combination Task
Antti-Veikko I Rosti, Bing Zhang, Spyros Matsoukas, and Richard Schwartz. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton. 2013 · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Cited alongside, same era.
An Empirical Comparison of Features and Tuning for Phrase-based Machine Translation
Spence Green, Daniel Cer, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Later among the works it cites.
Minimum Risk Training for Neural Machine Translation
Shiqi Shen, Yong Cheng, Zhongjun He, Wei He, Hua Wu, Maosong Sun, and Yang Liu. 2016 · 2016
Later among the works it cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Later among the works it cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Later among the works it cites.
Language Modeling with Gated Convolutional Networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Sequence level Training with Recurrent Neural Networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2015 · 2015
Cited alongside, same era.
Neural headline generation with sentence-wise optimization
Ayana, Shiqi Shen, Yu Zhao, Zhiyuan Liu, Maosong Sun, et al. 2016 · 2016
Cited alongside, same era.
Po-Sen Huang, Chong Wang, Dengyong Zhou, and Li Deng. 2017 · 2017
Closest in time.
Deep recurrent generative decoder for abstractive text summarization
Piji Li, Wai Lam, Lidong Bing, and Zihao Wang. 2017 · 2017
Closest in time.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2017 · 2017
Closest in time.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. 2017 · 2017
Closest in time.
Cutting-off redundant repeating generations for neural abstractive summarization
Jun Suzuki and Masaaki Nagata. 2017 · 2017
Closest in time.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Closest in time.
Deliberation networks: Sequence generation beyond one-pass decoding
Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu. 2017 · 2017
Closest in time.
Selective encoding for abstractive sentence summarization
Qingyu Zhou, Nan Yang, Furu Wei, and Ming Zhou. 2017 · 2017
Closest in time.