Fetching the paper…
Reading the bibliography…
Despite significant advances in the field of deep Reinforcement Learning (RL), today's algorithms still fail to learn human-level policies consistently over a set of diverse tasks such as Atari 2600 games.
Robust estimation of a location parameter
Peter J. Huber · 1964
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Marc L. Puterman · 1994
Earlier work this paper cites.
Robot learning from demonstration
Christopher Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1997
Earlier work this paper cites.
Learning from limited demonstrations
Beomjoon Kim, Amir-massoud Farahmand, Joelle Pineau, and Doina Precup · 2013
Earlier work this paper cites.
Boosted bellman residual minimization handling expert demonstrations
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Marc Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2015
Earlier work this paper cites.
Direct policy iteration with demonstrations
Jessica Chemali and Alessandro Lazaric · 2015
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Večerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Later among the works it cites.
The reactor: A fast and sample-efficient actor-critic agent for reinforcement learning
Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc Bellemare, and Remi Munos · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad G. Azar, and David Silver · 2018
Closest in time.
Deep Q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Andrew Sendonaris, Gabriel Dulac-Arnold, Ian Osband, and John Agapiou · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Cited alongside, same era.
TD learning with constrained gradients
Ishan Durugkar and Peter Stone · 2017
Cited alongside, same era.
Learning values across many orders of magnitude
Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, and David Silver
Cited in the paper.
Deep reinforcement learning with double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver
Cited in the paper.
Closest in time.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Closest in time.