Fetching the paper…
Reading the bibliography…
Transformer-based language models (LMs) track contextual information through large, hard-coded input windows.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter. 1998 · 1998
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. 2020 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. 2015 · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwinska, Sergio Gomez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity Stephen, Xiong Caiming, Bradbury James, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
T-rex: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
A large-scale corpus for conversation disentanglement
Jonathan K. Kummerfeld, Sai R. Gouravajhala, Joseph J. Peper, Vignesh Athreya, Chulaka Gunasekara, Jatin Ganhotra, Siva Sankalp Patel, Lazaros C Polymenakos, and Walter Lasecki. 2019 · 2019
Cited alongside, same era.
Beyond goldfish memory: Long-term open-domain conversation
Jing Xu, Arthur Szlam, and Jason Weston. 2021 · 2021
Later among the works it cites.
Factual probing is [MASK]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Recurrent memory transformer
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022 · 2022
Later among the works it cites.
DeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Later among the works it cites.
Memformer: A memory-augmented transformer for sequence modeling
Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Later among the works it cites.
Scaling Transformer to 1M tokens and beyond with RMT
Aydar Bulatov, Yuri Kuratov, and Mikhail Burtsev. 2023 · 2023
Later among the works it cites.
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023 · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional
Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2023 · 2023
Later among the works it cites.