Fetching the paper…
Reading the bibliography…
Memory-based neural networks model temporal data by leveraging an ability to remember information for long periods.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Memory systems 1994
Daniel L Schacter and Endel Tulving · 1994
Earlier work this paper cites.
Long short term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
A system for relational reasoning in human prefrontal cortex
James A Waltz, Barbara J Knowlton, Keith J Holyoak, Kyle B Boone, Fred S Mishkin, Marcia de Menezes Santos, Carmen R Thomas, and Bruce L Miller · 1999
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Using language models for information retrieval
Djoerd Hiemstra · 2001
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2009
Earlier work this paper cites.
English gigaword fifth edition ldc2011t07. dvd
Robert Parker, David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda · 2011
Earlier work this paper cites.
A neurocomputational system for relational reasoning
Barbara J Knowlton, Robert G Morrison, John E Hummel, and Keith J Holyoak · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Wojciech Zaremba and Ilya Sutskever · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Navdeep Jaitly Noam Shazeer Samy Bengio, Oriol Vinyals · 2015
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Discovering objects and their relations from entangled scene representations
David Raposo, Adam Santoro, David Barrett, Razvan Pascanu, Timothy Lillicrap, and Peter Battaglia · 2017
Later among the works it cites.
Relation networks for object detection
Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei · 2017
Later among the works it cites.
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Cited alongside, same era.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard S. Zemel · 2016
Cited alongside, same era.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2017
Later among the works it cites.
Tracking the world state with recurrent entity networks
Mikael Henaff, Jason Weston, Arthur Szlam, Antoine Bordes, and Yann LeCun · 2017
Later among the works it cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2018
Closest in time.
Non-local recurrent network for image restoration
Ding Liu, Bihan Wen, Yuchen Fan, Chen Change Loy, and Thomas S Huang · 2018
Closest in time.
Working memory networks: Augmenting memory networks with a relational reasoning module
Juan Pavez, Héctor Allende, and Héctor Allende-Cid · 2018
Closest in time.
Fast parametric learning with activation memorization
Jack W Rae, Chris Dyer, Peter Dayan, and Timothy P Lillicrap · 2018
Closest in time.
Convolutional sequence modeling revisited
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Closest in time.
Scalable language modeling: Wikitext-103 on a single gpu in 12 hours
Stephen Merity, Nitish Shirish Keskar, James Bradbury, and Richard Socher · 2018
Closest in time.
Relational deep reinforcement learning
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, Murray Shanahan, Victoria Langston, Razvan Pascanu, Matthew Botvinick, Oriol Vinyals, and Peter Battaglia · 2018
Closest in time.
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Closest in time.