Fetching the paper…
Reading the bibliography…
Transformers are unable to model long-term memories effectively, since the amount of computation they need to perform grows with the context length.
Adaptive multivariate ridge regression
Philip J Brown, James V Zidek, et al. 1980 · 1980
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter, J urgen Schmidhuber, and Corso Elvezia. 1997 · 1997
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Cadherins: actin with the cytoskeleton to form synapses
Shernaz X Bamji. 2005 · 2005
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Psychology of Language
D.W. Carroll. 2007 · 2007
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder–Decoder Approaches
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Cognitive flexibility and long-term depression (LTD) are impaired following β \beta -catenin stabilization in vivo
Fergil Mills, Thomas E Bartlett, Lasse Dissing-Olesen, Marta B Wisniewska, Jacek Kuznicki, Brian A Macvicar, Yu Tian Wang, and Shernaz X Bamji. 2014 · 2014
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2014 · 2014
Earlier work this paper cites.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Sarath Chandar, Sungjin Ahn, Hugo Larochelle, Pascal Vincent, Gerald Tesauro, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Improving Neural Language Models with a Continuous Cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Scaling memory-augmented neural networks with sparse reads and writes
Jack W Rae, Jonathan J Hunt, Tim Harley, Ivo Danihelka, Andrew Senior, Greg Wayne, Alex Graves, and Timothy P Lillicrap. 2016 · 2016
Cited alongside, same era.
Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
A Dataset for Document Grounded Conversations
Kangyan Zhou, Shrimai Prabhumoye, and Alan W Black. 2018 · 2018
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Long-Lasting Verbatim Memory for the Words of Books After a Single Reading Without Any Learning Intention
Christof Kuhbandner. 2020 · 2020
Later among the works it cites.
Sparse and Continuous Attention Mechanisms
André FT Martins, Marcos Treviso, António Farinhas, Vlad Niculae, Mário AT Figueiredo, and Pedro MQ Aguiar. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Gaussian Transformer: A Lightweight Approach for Natural Language Inference
Maosheng Guo, Yu Zhang, and Ting Liu. 2019 · 2019
Cited alongside, same era.
Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Compressive Transformers for Long-Range Sequence Modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap. 2019 · 2019
Cited alongside, same era.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020 · 2020
Later among the works it cites.
Hard-Coded Gaussian Attention for Neural Machine Translation
Weiqiu You, Simeng Sun, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2021 · 2021
Closest in time.
Augmenting Transformers with KNN-Based Composite Memory for Dialog
Angela Fan, Claire Gardent, Chloé Braud, and Antoine Bordes. 2021 · 2021
Closest in time.
Multimodal Continuous Visual Attention Mechanisms
António Farinhas, André F. T. Martins, and P. Aguiar. 2021 · 2021
Closest in time.
Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. 2021 · 2021
Closest in time.
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis. 2021 · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Closest in time.
Cluster-Former: Clustering-based Sparse Transformer for Question Answering
Shuohang Wang, Luowei Zhou, Zhe Gan, Yen-Chun Chen, Yuwei Fang, Siqi Sun, Yu Cheng, and Jingjing Liu. 2021 · 2021
Closest in time.
Adaptive Semiparametric Language Models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Closest in time.