Fetching the paper…
Reading the bibliography…
Dynamic evaluation of language models (LMs) adapts model parameters at test time using gradient information from previous tokens and substantially improves LM performance.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Dynamic evaluation of transformer language models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. 2019 · 1904
Earlier work this paper cites.
Using fast weights to deblur old memories
Geoffrey E Hinton and David C Plaut. 1987 · 1987
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber. 1992 · 1992
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Efficient softmax approximation for GPUs
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Meta networks
Tsendsuren Munkhdalai and Hong Yu. 2017 · 2017
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Metalearned neural memory
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang, and Adam Trischler. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Rapid learning or feature reuse? Towards understanding the effectiveness of MAML
Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals. 2020 · 2020
Later among the works it cites.
Going beyond linear transformers with recurrent fast weight programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber. 2021 · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Later among the works it cites.
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Later among the works it cites.
Reconsidering the past: Optimizing hidden states in language models
Davis Yoshida and Kevin Gimpel. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu. 2016a
Cited in the paper.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016b
Cited in the paper.
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc V Le. 2022 · 2022
Closest in time.
Fast model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2022 · 2022
Closest in time.