Fetching the paper…
Reading the bibliography…
Memory-augmented neural networks consisting of a neural controller and an external memory have shown potentials in long-term sequential learning.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Visuo-spatial working memory
Robert H Logie · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Very deep convolutional networks for natural language processing
Alexis Conneau, Holger Schwenk, Loıc Barrault, and Yann Lecun · 2016
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Cited alongside, same era.
Scaling memory-augmented neural networks with sparse reads and writes
Jack Rae, Jonathan J Hunt, Ivo Danihelka, Timothy Harley, Andrew W Senior, Gregory Wayne, Alex Graves, and Tim Lillicrap · 2016
Cited alongside, same era.
Generative and discriminative text classification with recurrent neural networks
Dani Yogatama, Chris Dyer, Wang Ling, and Phil Blunsom · 2017
Later among the works it cites.
Learning to skim text
Adams Wei Yu, Hongrae Lee, and Quoc Le · 2017
Later among the works it cites.
Robust and scalable differentiable neural computer for question answering
Jörg Franke, Jan Niehues, and Alex Waibel · 2018
Later among the works it cites.
When recurrent models don’t need to be recurrent
John Miller and Moritz Hardt · 2018
Later among the works it cites.
A new method of region embedding for text classification
Chao Qui, Bo Huang, Guocheng Niu, Daren Li, Daxiang Dong, Wei He, Dianhai Yu, and Hua Wu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Cited alongside, same era.
Dilated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Unbounded cache model for online language modeling with open vocabulary
Edouard Grave, Moustapha M Cisse, and Armand Joulin
Cited in the paper.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier
Cited in the paper.
Minjoon Seo, Sewon Min, Ali Farhadi, and Hannaneh Hajishirzi · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu H Trinh, Andrew M Dai, Thang Luong, and Quoc V Le · 2018
Later among the works it cites.
Memory architectures in recurrent neural network language models
Dani Yogatama, Yishu Miao, Gabor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, and Phil Blunsom · 2018
Later among the works it cites.
Fast and accurate text classification: Skimming, rereading and early stopping
Keyi Yu, Yang Liu, Alexander G Schwing, and Jian Peng · 2018
Later among the works it cites.