Maximum mutual information estimation of hidden Markov model parameters for speech recognition. In Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP’86
Lalit Bahl, Peter Brown, Peter De Souza, and Robert Mercer. 1986 · 1986
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos. 1990 · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
A maximum entropy approach to natural language processing
Adam L Berger, Vincent J Della Pietra, and Stephen A Della Pietra. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
Advances in automatic text summarization
Inderjeet Mani and Mark T Maybury. 1999 · 1999
Earlier work this paper cites.
The PageRank citation ranking: Bringing order to the web
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999 · 1999
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning. In Proceedings of the Seventeenth conference on Uncertainty in artificial intelligence
Lex Weaver and Nigel Tao. 2001 · 2001
Earlier work this paper cites.
Topic-sensitive pagerank. In Proceedings of the 11th international conference on World Wide Web
Taher H Haveliwala. 2002 · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Introduction to the special issue on summarization
Dragomir R Radev, Eduard Hovy, and Kathleen McKeown. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Minimum error rate training in statistical machine translation. In Proceedings of the 41st Annual Meeting on Association for Computational Linguistics-Volume 1
Franz Josef Och. 2003 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Textrank: Bringing order into text. In Proceedings of the 2004 conference on empirical methods in natural language processing
Rada Mihalcea and Paul Tarau. 2004 · 2004
Earlier work this paper cites.
A survey on automatic text summarization
Dipanjan Das and André FT Martins. 2007 · 2007
Earlier work this paper cites.
Probabilistic latent maximal marginal relevance. In Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval
Shengbo Guo and Scott Sanner. 2010 · 2010
Earlier work this paper cites.
Automatic summarization
Ani Nenkova, Kathleen McKeown, et al · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Text summarisation in progress: a literature review
Elena Lloret and Manuel Palomar. 2012 · 2012
Earlier work this paper cites.
Annotated gigaword. In Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction
Courtney Napoles, Matthew Gormley, and Benjamin Van Durme. 2012 · 2012
Earlier work this paper cites.
A survey of extractive and abstractive text summarization techniques. In Emerging Trends in Engineering and Technology (ICETET), 2013 6th International Conference on
Vipul Dalal and Latesh G Malik. 2013 · 2013
Earlier work this paper cites.
A systematic exploration of diversity in machine translation. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing
Kevin Gimpel, Dhruv Batra, Chris Dyer, and Gregory Shakhnarovich. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Original
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks. In International Conference on Machine Learning
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Automatic text summarization: Past, present and future
Horacio Saggion and Thierry Poibeau. 2013 · 2013
Earlier work this paper cites.
Training recurrent neural networks
Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Original
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS 2014 Workshop on Deep Learning, December 2014
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks. In International Conference on Machine Learning
Alex Graves and Navdeep Jaitly. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models. In International Conference on Machine Learning
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks. In Advances in neural information processing systems
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks. In Advances in Neural Information Processing Systems
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing
Samuel Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data. In Advances in neural information processing systems
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Generating news headlines with recurrent neural networks
Original
Konstantin Lopyrev. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing
Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding. In Automatic Speech Recognition and Understanding (ASRU), 2015 IEEE Workshop on
Yajie Miao, Mohammad Gowayyed, and Florian Metze. 2015 · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Original
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015 · 2015
Earlier work this paper cites.
A Neural Attention Model for Abstractive Sentence Summarization. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Improving Multi-Step Prediction of Learned Time Series Models. In Twenty-Ninth AAAI Conference on Artificial Intelligence
Arun Venkatraman, Martial Hebert, and J Andrew Bagnell. 2015 · 2015
Earlier work this paper cites.
Pointer networks. In Advances in Neural Information Processing Systems
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
Reinforcement learning neural turing machines-revised
Original
Wojciech Zaremba and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Original
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2016a · 2016
Earlier work this paper cites.