Fetching the paper…
Reading the bibliography…
Supervised regression to demonstrations has been demonstrated to be a stable way to train deep policy networks.
Lin, K.; and Zhou, J. 2019 · 1906
Earlier work this paper cites.
RLCard: A Toolkit for Reinforcement Learning in Card Games
Zha, D.; Lai, K.-H.; Cao, Y.; Huang, S.; Wei, R.; Guo, J.; and Hu, X. 2019a · 1910
Earlier work this paper cites.
Training Agents using Upside-Down Reinforcement Learning
Srivastava, R. K.; Shyam, P.; Mutz, F.; Jaśkowski, W.; and Schmidhuber, J. 2019 · 1912
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. 1992 · 1992
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Thrun, S. B. 1992 · 1992
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Lin, L.-J. 1993 · 1993
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Dayan, P.; and Hinton, G. E. 1997 · 1997
Earlier work this paper cites.
The cross entropy method for fast policy search
Mannor, S.; Rubinstein, R. Y.; and Gat, Y. 2003 · 2003
Earlier work this paper cites.
AutoOD: Automated Outlier Detection via Curiosity-guided Search and Self-imitation Learning
Li, Y.; Chen, Z.; Zha, D.; Zhou, K.; Jin, H.; Chen, H.; and Hu, X. 2020 · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J.; and Schaal, S. 2007 · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D.; Maas, A.; Bagnell, J. A.; and Dey, A. K. 2008 · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, J.; and Peters, J. R. 2009 · 2009
Earlier work this paper cites.
Double Q-learning
Hasselt, H. V. 2010 · 2010
Earlier work this paper cites.
Variational policy search via trajectory optimization
Levine, S.; and Koltun, V. 2013 · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Schaul, T.; Quan, J.; Antonoglou, I.; and Silver, D. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M.; Srinivasan, S.; Ostrovski, G.; Schaul, T.; Saxton, D.; and Munos, R. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2016 · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 2016
Maximum a posteriori policy optimisation
Abdolmaleki, A.; Springenberg, J. T.; Tassa, Y.; Munos, R.; Heess, N.; and Riedmiller, M. 2018 · 2018
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S.; Moitra, S.; Gelada, C.; Kumar, S.; and Bellemare, M. G. 2018 · 2018
Later among the works it cites.
Minimalistic Gridworld Environment for OpenAI Gym
Chevalier-Boisvert, M.; Willems, L.; and Pal, S. 2018 · 2018
Later among the works it cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W.; Ostrovski, G.; Silver, D.; and Munos, R. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Policy distillation
Rusu, A. A.; Colmenarejo, S. G.; Gulcehre, C.; Desjardins, G.; Kirkpatrick, J.; Pascanu, R.; Mnih, V.; Kavukcuoglu, K.; and Hadsell, R. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H.; Guez, A.; and Silver, D. 2016 · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Abbeel, O. P.; and Zaremba, W. 2017 · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G.; Dabney, W.; and Munos, R. 2017 · 2017
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R.; Henderson, P.; Gomrokchi, M.; and Precup, D. 2017 · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Cited alongside, same era.
Gangwani, T.; Liu, Q.; and Peng, J. 2018 · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P.; Islam, R.; Bachman, P.; Pineau, J.; Precup, D.; and Meger, D. 2018 · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M.; Modayil, J.; Van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; and Silver, D. 2018 · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
Mania, H.; Guy, A.; and Recht, B. 2018 · 2018
Later among the works it cites.
Remember and Forget for Experience Replay
Novati, G.; and Koumoutsakos, P. 2018 · 2018
Later among the works it cites.
Self-imitation learning
Oh, J.; Guo, Y.; Singh, S.; and Lee, H. 2018 · 2018
Later among the works it cites.
Organizing experience: a deeper look at replay mechanisms for sample-based planning in continuous state domains
Pan, Y.; Zaheer, M.; White, A.; Patterson, A.; and White, M. 2018 · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
Dual Policy Distillation
Lai, K.-H.; Zha, D.; Li, Y.; and Hu, X. 2020 · 2020
Later among the works it cites.
Random search and reproducibility for neural architecture search
Li, L.; and Talwalkar, A. 2020 · 2020
Later among the works it cites.