Fetching the paper…
Reading the bibliography…
In this work we present a new agent architecture, called Reactor, which combines multiple algorithmic and architectural contributions to produce an agent with higher sample-efficiency than Prioritized Dueling DQN (Wang et al., 2016) and Categorical DQN (Bellemare et al., 2017), while giving better run-time performance than A3C (Mnih et al., 2016).
An algorithm for the organisation of information
Adel’son G Velskii and E Landis · 1976
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-H Lin · 1992
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Explorations in efficient reinforcement learning
Marco A Wiering · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S Sutton, and Satinder Singh · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Neural fitted q iteration-first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Deep reinforcement learning with experience replay based on sarsa
Dongbin Zhao, Haitao Wang, Kun Shao, and Yuanheng Zhu · 2016
Later among the works it cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Closest in time.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Closest in time.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2015
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
Frank S He, Yang Liu, Alexander G Schwing, and Jian Peng · 2017
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2017
Closest in time.
Combining policy gradient and q-learning
Brendan O’Donoghue, Remi Munos, Koray Kavukcuoglu, and Volodymyr Mnih · 2017
Closest in time.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Closest in time.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Closest in time.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Closest in time.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2017
Closest in time.