Fetching the paper…
Reading the bibliography…
The exploration mechanism used by a Deep Reinforcement Learning (RL) agent plays a key role in determining its sample efficiency.
Introduction to reinforcement learning , volume 135
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Reinforcement learning with supervision by a stable controller
M. T. Rosenstein and A. G. Barto · 2004
Earlier work this paper cites.
Probabilistic policy reuse in a reinforcement learning agent
F. Fernández and M. Veloso · 2006
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Thompson sampling: An asymptotically optimal finite-time analysis
E. Kaufmann, N. Korda, and R. Munos · 2012
Earlier work this paper cites.
Regret bounds for reinforcement learning with policy advice
M. G. Azar, A. Lazaric, and E. Brunskill · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
J. Zhang and K. Cho · 2016
Earlier work this paper cites.
Y. Zhan, H. B. Ammar, et al · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Večerík, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Cited alongside, same era.
Bayesian policy gradients via alpha divergence dropout inference
P. Henderson, T. Doan, R. Islam, and D. Meger · 2017
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Later among the works it cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Later among the works it cites.
Reinforcement learning from imperfect demonstrations
Y. Gao, J. Lin, F. Yu, S. Levine, T. Darrell, et al · 2018
Later among the works it cites.
Z. Wang and M. E. Taylor · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving reinforcement learning with confidence-based demonstrations
Z. Wang and M. E. Taylor · 2017
Cited alongside, same era.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Cited alongside, same era.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Prm-rl: Long-range robotic navigation tasks by combining reinforcement learning and sampling-based planning
A. Faust, K. Oslund, O. Ramirez, A. Francis, L. Tapia, M. Fiser, and J. Davidson · 2018
Cited alongside, same era.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jowl, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Cited alongside, same era.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Cited alongside, same era.
Residual reinforcement learning for robot control
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine · 2018
Later among the works it cites.
T. Silver, K. Allen, J. Tenenbaum, and L. Kaelbling · 2018
Later among the works it cites.
Learning with training wheels: Speeding up training with a simple controller for deep reinforcement learning
L. Xie, S. Wang, S. Rosa, A. Markham, and N. Trigoni · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, et al · 2018
Later among the works it cites.
Crossnorm: Normalization for off-policy td reinforcement learning
A. Bhatt, M. Argus, A. Amiranashvili, and T. Brox · 2019
Closest in time.
Towards characterizing divergence in deep q-learning
J. Achiam, E. Knight, and P. Abbeel · 2019
Closest in time.