Fetching the paper…
Reading the bibliography…
Neural language models predict the next token using a latent representation of the immediate token history.
Pattern recognition by means of automatic analogue apparatus
WK Taylor · 1959
Earlier work this paper cites.
Learning matrices and their applications
Karl Steinbuch and UAW Piske · 1963
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Empirical evaluation and combination of advanced language modeling techniques
Tomas Mikolov, Anoop Deoras, Stefan Kombrink, Lukas Burget, and Jan Cernockỳ · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Neural programmer-interpreters
Scott Reed and Nando de Freitas · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Cited alongside, same era.
Sparse non-negative matrix language modeling for skip-grams
Noam Shazeer, Joris Pelemans, and Ciprian Chelba · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, and Rob Fergus · 2015
Cited alongside, same era.
Memory networks
Jason Weston, Sumit Chopra, and Antoine Bordes · 2015
Cited alongside, same era.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston · 2016
Later among the works it cites.
Blackout: Speeding up recurrent neural network language models with very large vocabularies
Shihao Ji, SVN Vishwanathan, Nadathur Satish, Michael J Anderson, and Pradeep Dubey · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Text understanding with the attention sum reader network
Rudolf Kadlec, Martin Schmid, Ondrej Bajgar, and Jan Kleindienst · 2016
Later among the works it cites.
Key-value memory networks for directly reading documents
Alexander Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling recurrent neural network language models
Will Williams, Niranjani Prasad, David Mrva, Tom Ash, and Tony Robinson · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Cited alongside, same era.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Cited alongside, same era.
Gated-attention readers for text comprehension
Bhuwan Dhingra, Hanxiao Liu, William W Cohen, and Ruslan Salakhutdinov · 2016
Cited alongside, same era.
Dynamic neural turing machine with soft and hard addressing schemes
Caglar Gulcehre, Sarath Chandar, Kyunghyun Cho, and Yoshua Bengio · 2016
Cited alongside, same era.
Later among the works it cites.
Reasoning about entailment with neural attention
Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomas Kocisky, and Phil Blunsom · 2016
Later among the works it cites.
Higher order recurrent neural networks
Rohollah Soltani and Hui Jiang · 2016
Later among the works it cites.
Recurrent memory networks for language modeling
Ke Tran, Arianna Bisazza, and Christof Monz · 2016
Later among the works it cites.
Natural language comprehension with the epireader
Adam Trischler, Zheng Ye, Xingdi Yuan, and Kaheer Suleman · 2016
Later among the works it cites.
Separating answers from queries for neural reading comprehension
Dirk Weissenborn · 2016
Later among the works it cites.
Reference-aware language models
Zichao Yang, Phil Blunsom, Chris Dyer, and Wang Ling · 2016
Later among the works it cites.