Fetching the paper…
Reading the bibliography…
Transformers with linearised attention (''linear Transformers'') have demonstrated the practical scalability and effectiveness of outer product-based Fast Weight Programmers (FWPs) from the '90s.
Adaptive switching circuits
Bernard Widrow and Marcian E Hoff · 1960
Earlier work this paper cites.
The correlation theory of brain function
Christoph von der Malsburg · 1981
Earlier work this paper cites.
Dynamic connections in neural networks
Jerome A Feldman · 1982
Earlier work this paper cites.
Putting knowledge in its place: A scheme for programming parallel processing structures on the fly
James L McClelland · 1985
Earlier work this paper cites.
Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in non-stationary environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
A stochastic version of the delta rule
Stephen José Hanson · 1990
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Reducing the ratio between learning complexity and number of time varying variables in fully recurrent nets
Jürgen Schmidhuber · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 2000
Earlier work this paper cites.
Solving deep memory POMDPs with recurrent policy gradients
Daan Wierstra, Alexander Förster, Jan Peters, and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
Recurrent policy gradients
Daan Wierstra, Alexander Förster, Jan Peters, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Lecture 6.5- RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Wojciech Zaremba and Ilya Sutskever · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A dynamic convolutional layer for short rangeweather prediction
Benjamin Klein, Lior Wolf, and Yehuda Afek · 2015
Earlier work this paper cites.
Highway networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Increasing the action gap: New operators for reinforcement learning
Marc G. Bellemare, Georg Ostrovski, Arthur Guez, Philip S. Thomas, and Rémi Munos · 2016
Earlier work this paper cites.
Image question answering using convolutional neural network with dynamic parameter prediction
Hyeonwoo Noh, Paul Hongsuck Seo, and Bohyung Han · 2016
Earlier work this paper cites.
Dynamic filter networks
Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc V Gool · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Hypernetworks
David Ha, Andrew Dai, and Quoc V Le · 2017
Towards interpretable reinforcement learning using attention augmented agents
Alexander Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, and Danilo Jimenez Rezende · 2019
Later among the works it cites.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Juergen Schmidhuber · 2019
Later among the works it cites.
Training agents using upside-down reinforcement learning
Rupesh Kumar Srivastava, Pranav Shyam, Filipe Mutz, Wojciech Jaśkowski, and Jürgen Schmidhuber · 2019
Later among the works it cites.
Language modeling with deep Transformers
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gated fast weights for on-the-fly neural program generation
Imanol Schlag and Jürgen Schmidhuber · 2017
Cited alongside, same era.
ListOps: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel Bowman · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
Thomas Miconi, Kenneth Stanley, and Jeff Clune · 2018
Cited alongside, same era.
Fast weight long short-term memory
T Anderson Keller, Sharath Nittur Sridhar, and Xin Wang · 2018
Cited alongside, same era.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2018
Cited alongside, same era.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz · 2018
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Later among the works it cites.
Stabilizing Transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matthew M. Botvinick, Nicolas Heess, and Raia Hadsell · 2020
Later among the works it cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Later among the works it cites.
Addressing some limitations of Transformers with feedback memory
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar · 2020
Later among the works it cites.
Agent57: Outperforming the Atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Deformable DETR: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Closest in time.
Learning advanced mathematical computations from examples
Francois Charton, Amaury Hayat, and Guillaume Lample · 2021
Closest in time.
Efficient transformers in reinforcement learning using actor-learner distillation
Emilio Parisotto and Ruslan Salakhutdinov · 2021
Closest in time.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong · 2021
Closest in time.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
Training and generating neural networks in compressed weight space
Kazuki Irie and Jürgen Schmidhuber · 2021
Closest in time.
LambdaNetworks: Modeling long-range interactions without attention
Irwan Bello · 2021
Closest in time.
Pretrained transformers as universal computation engines
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch · 2021
Closest in time.
Decision Transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Closest in time.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Closest in time.
The Neural Data Router: Adaptive control flow in Transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Closest in time.
What matters for on-policy deep actor-critic methods? A large scale study
Marcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, et al · 2021
Closest in time.