Fetching the paper…
Reading the bibliography…
Current language models have a significant limitation in the ability to encode and decode factual knowledge.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 1922
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P, Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 1993
Earlier work this paper cites.
Wordnet: a lexical database for english
Miller, George A · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Topic-based language models using em
Gildea, Daniel and Hofmann, Thomas · 1999
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Yoshua, Ducharme, Réjean, Vincent, Pascal, and Jauvin, Christian · 2003
Earlier work this paper cites.
Latent dirichlet allocation
Blei, David M, Ng, Andrew Y, and Jordan, Michael I · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
A scalable hierarchical distributed language model
Mnih, Andriy and Hinton, Geoffrey E · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian, Bergeron, Arnaud, Bouchard, Nicolas, Warde-Farley, David, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Context dependent recurrent neural network language model
Mikolov, Tomas and Zweig, Geoffrey · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, Andriy and Teh, Yee Whye · 2012
Cited alongside, same era.
Translating embeddings for modeling multi-relational data
Bordes, Antoine, Usunier, Nicolas, Garcia-Duran, Alberto, Weston, Jason, and Yakhnenko, Oksana · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, Ciprian, Mikolov, Tomas, Schuster, Mike, Ge, Qi, Brants, Thorsten, Koehn, Phillipp, and Robinson, Tony · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Vinyals, Oriol and Le, Quoc · 2015
Later among the works it cites.
Pointer networks
Vinyals, Oriol, Fortunato, Meire, and Jaitly, Navdeep · 2015
Later among the works it cites.
Memory networks
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2015
Later among the works it cites.
Incorporating copying mechanism in sequence-to-sequence learning
Gu, Jiatao, Lu, Zhengdong, Li, Hang, and Li, Victor O. K · 2016
Closest in time.
Pointing the unknown words
Gulcehre, Caglar, Ahn, Sungjin, Nallapati, Ramesh, Zhou, Bowen, and Bengio, Yoshua · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A neural network for factoid question answering over paragraphs
Iyyer, Mohit, Boyd-Graber, Jordan L, Claudino, Leonardo Max Batista, Socher, Richard, and Daumé III, Hal · 2014
Cited alongside, same era.
Convolutional neural networks for sentence classification
Kim, Yoon · 2014
Cited alongside, same era.
Large-scale simple question answering with memory networks
Bordes, Antoine, Usunier, Nicolas, Chopra, Sumit, and Weston, Jason · 2015
Cited alongside, same era.
Enriching word embeddings using knowledge graph for semantic tagging in conversational dialog systems
Celikyilmaz, Asli, Hakkani-Tur, Dilek, Pasupat, Panupong, and Sarikaya, Ruhi · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Jean, Sebastien, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2015
Cited alongside, same era.
Nickel, Maximilian, Murphy, Kevin, Tresp, Volker, and Gabrilovich, Evgeniy · 2015
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural networks
Serban, Iulian V, Sordoni, Alessandro, Bengio, Yoshua, Courville, Aaron, and Pineau, Joelle · 2015
Cited alongside, same era.
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Closest in time.
Neural text generation from structured data with application to the biography domain
Lebret, Rémi, Grangier, David, and Auli, Michael · 2016
Closest in time.
Leveraging lexical resources for learning entity embeddings in multi-relational data
Long, Teng, Lowe, Ryan, Cheung, Jackie Chi Kit, and Precup, Doina · 2016
Closest in time.
Pointer sentinel mixture models
Merity, Stephen, Xiong, Caiming, Bradbury, James, and Socher, Richard · 2016
Closest in time.
Neural programmer-interpreters
Reed, Scott and de Freitas, Nando · 2016
Closest in time.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, Jason, Bordes, Antoine, Chopra, Sumit, and Mikolov, Tomas · 2016
Closest in time.