Fetching the paper…
Reading the bibliography…
Deep learning methods have recently achieved great empirical success on machine translation, dialogue response generation, summarization, and other text generation tasks.
Eliza—a computer program for the study of natural language communication between man and machine
J. Weizenbaum · 1966
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
S. F. Chen and J. Goodman · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Evaluation metrics for language models
S. F. Chen, D. Beeferman, and R. Rosenfeld · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Statistical phrase-based translation
P. Koehn, F. J. Och, and D. Marcu · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Advice for applying machine learning
A. Ng · 2007
Earlier work this paper cites.
Mind the gap: Dangers of divorcing evaluations of summary content from linguistic quality
J. M. Conroy and H. T. Dang · 2008
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Addressing the rare word problem in neural machine translation
M.-T. Luong, I. Sutskever, Q. V. Le, O. Vinyals, and W. Zaremba · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
C.-W. Liu, R. Lowe, I. V. Serban, M. Noseworthy, L. Charlin, and J. Pineau · 2016
Later among the works it cites.
Attention and augmented recurrent neural networks
C. Olah and S. Carter · 2016
Later among the works it cites.
Modeling coverage for neural machine translation
Z. Tu, Z. Lu, Y. Liu, X. Liu, and H. Li · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2015
Cited alongside, same era.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2016
Cited alongside, same era.
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Closest in time.
Massive exploration of neural machine translation architectures
D. Britz, A. Goldie, T. Luong, and Q. Le · 2017
Closest in time.
Snapshot ensembles: Train 1, get m for free
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger · 2017
Closest in time.
Six challenges for neural machine translation
P. Koehn and R. Knowles · 2017
Closest in time.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, A. Ritter, and D. Jurafsky · 2017
Closest in time.
On the state of the art of evaluation in neural language models
G. Melis, C. Dyer, and P. Blunsom · 2017
Closest in time.
Why we need new evaluation metrics for NLG
J. Novikova, O. Dušek, A. C. Curry, and V. Rieser · 2017
Closest in time.
Get to the point: Summarization with pointer-generator networks
A. See, P. J. Liu, and C. D. Manning · 2017
Closest in time.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Closest in time.
Data noising as smoothing in neural network language models
Z. Xie, S. I. Wang, J. Li, D. Lévy, A. Nie, D. Jurafsky, and A. Y. Ng · 2017
Closest in time.