Fetching the paper…
Reading the bibliography…
We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
James L McClelland, Bruce L McNaughton, and Randall C O’Reilly · 1995
Earlier work this paper cites.
Experiments with reinforcement learning in problems with continuous state and action spaces
Juan C Santamaría, Richard S Sutton, and Ashwin Ram · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Barycentric interpolators for continuous space and time reinforcement learning
Remi Munos and Andrew W Moore · 1998
Earlier work this paper cites.
Kernel-based reinforcement learning
Dirk Ormoneit and Śaunak Sen · 2002
Earlier work this paper cites.
A robot that reinforcement-learns to identify and memorize important previous observations
Bram Bakker, Viktor Zhumatiy, Gabriel Gruener, and Jürgen Schmidhuber · 2003
Earlier work this paper cites.
Cbr for state value function approximation in reinforcement learning
Thomas Gabel and Martin Riedmiller · 2005
Earlier work this paper cites.
Decision theory, reinforcement learning, and the brain
Peter Dayan and Nathaniel D Daw · 2008
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
David Silver, Richard S Sutton, and Martin Müller · 2008
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al · 2016
Cited alongside, same era.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Honglak Lee, et al · 2016
Reinforcement learning and episodic memory in humans and animals: an integrative framework
Samuel J Gershman and Nathaniel D Daw · 2017
Later among the works it cites.
Is prioritized sweeping the better episodic control?
Johanni Brea · 2017
Later among the works it cites.
The pycolab game engine
Tom Stepleton · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Closest in time.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Learning to remember rare events
Lukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Cited alongside, same era.
Neural episodic control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adrià Puigdomènech, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Cited alongside, same era.
Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis
Cited in the paper.
Closest in time.
Unsupervised predictive memory in a goal-directed agent
Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja, Agnieszka Grabska-Barwinska, Jack Rae, Piotr Mirowski, Joel Z Leibo, Adam Santoro, et al · 2018
Closest in time.
Faster deep q-learning using neural episodic control
Daichi Nishio and Satoshi Yamane · 2018
Closest in time.
Semiparametric reinforcement learning
Mika Sarkin Jain and Jack Lindsey · 2018
Closest in time.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Closest in time.
Memory-augmented monte carlo tree search
Chenjun Xiao, Jincheng Mei, and Martin Müller · 2018
Closest in time.
Memory-based parameter adaptation
Pablo Sprechmann, Siddhant M Jayakumar, Jack W Rae, Alexander Pritzel, Adrià Puigdomènech Badia, Benigno Uria, Oriol Vinyals, Demis Hassabis, Razvan Pascanu, and Charles Blundell · 2018
Closest in time.