Fetching the paper…
Reading the bibliography…
The Softmax function is used in the final layer of nearly all existing sequence-to-sequence models for language generation.
The psycho-biology of language
George Kingsley Zipf · 1935
Earlier work this paper cites.
Vector quantization
Robert M. Gray · 1990
Earlier work this paper cites.
Mixture density networks
Christopher M. Bishop · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Theory of Point Estimation
E.L. Lehmann and G. Casella · 1998
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Yoshua Bengio and Jean-Sébastien Senecal · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Frederic Morin and Yoshua Bengio · 2005
Earlier work this paper cites.
A simple, fast, and effective reparameterization of IBM Model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
In Proc. ACL , 2014
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, and John Makhoul · 2014
Earlier work this paper cites.
Dependency-based word embeddings
Omer Levy and Yoav Goldberg · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Cited alongside, same era.
On the accuracy of self-normalized log-linear models
Jacob Andreas, Maxim Rabinovich, Michael I. Jordan, and Dan Klein · 2015
Cited alongside, same era.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Shihao Ji, S. V. N. Vishwanathan, Nadathur Satish, Michael J. Anderson, and Pradeep Dubey · 2015
Cited alongside, same era.
Hubness and pollution: Delving into cross-space mapping for zero-shot learning
Angeliki Lazaridou, Georgiana Dinu, and Marco Baroni · 2015
Cited alongside, same era.
Two/too simple adaptations of word2vec for syntax problems
Exploring the limits of language modeling, 2016
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Later among the works it cites.
A new type of sharp bounds for ratios of modified bessel functions
Diego Ruiz-Antolín and Javier Segura · 2016
Later among the works it cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Later among the works it cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang Ling, Chris Dyer, Alan Black, and Isabel Trancoso · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Cited alongside, same era.
Oriol Vinyals and Quoc V. Le · 2015
Cited alongside, same era.
Normalized word embedding and orthogonal transform for bilingual word translation
Chao Xing, Dong Wang, Chao Liu, and Yiye Lin · 2015
Cited alongside, same era.
Neural generative question answering
Jun Yin, Xin Jiang, Zhengdong Lu, Lifeng Shang, Hang Li, and Xiaoming Li · 2015
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, et al · 2016
Cited alongside, same era.
Michael J. Denkowski and Graham Neubig · 2017
Later among the works it cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2017
Later among the works it cites.
Opennmt: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush · 2017
Later among the works it cites.
Neural machine translation via binary code prediction
Yusuke Oda, Philip Arthur, Graham Neubig, Koichiro Yoshino, and Satoshi Nakamura · 2017
Later among the works it cites.
Mimicking word embeddings using subword rnns
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein · 2017
Later among the works it cites.
A survey of cross-lingual embedding models
Sebastian Ruder · 2017
Later among the works it cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Later among the works it cites.
Toward human parity in conversational speech recognition
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Michael L Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig · 2017
Later among the works it cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Closest in time.