Fetching the paper…
Reading the bibliography…
Deep reinforcement learning algorithms can learn complex behavioral skills, but real-world application of these methods requires a large amount of experience to be collected by the agent.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Learning compound multi-step controllers under unknown dynamics
Weiqiao Han, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Brandon Amos, Lei Xu, and J Zico Kolter · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Learning from demonstrations through the use of non-rigid registration
John Schulman, Jonathan Ho, Cameron Lee, and Pieter Abbeel · 2016
Cited alongside, same era.
Path integral guided policy search
Yevgen Chebotar, Mrinal Kalakrishnan, Ali Yahya, Adrian Li, Stefan Schaal, and Sergey Levine · 2017
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
David Held, Xinyang Geng, Carlos Florensa, and Pieter Abbeel · 2017
Closest in time.
Uncertainty-aware reinforcement learning for collision avoidance
Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Teacher-student curriculum learning
Tambet Matiisen, Avital Oliver, Taco Cohen, and John Schulman · 2017
Closest in time.
Discrete sequential prediction of continuous actions for deep rl
Luke Metz, Julian Ibarz, Navdeep Jaitly, and James Davidson · 2017
Closest in time.
Learning to push by grasping: Using multiple tasks for effective learning
Lerrel Pinto and Abhinav Gupta · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carlos Florensa, David Held, Markus Wulfmeier, and Pieter Abbeel · 2017
Cited alongside, same era.
Dhiraj Gandhi, Lerrel Pinto, and Abhinav Gupta · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Cited alongside, same era.
Risk aversion in markov decision processes via near optimal chernoff bounds
Teodor M Moldovan and Pieter Abbeel
Cited in the paper.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel
Cited in the paper.
Closest in time.
Safe visual navigation via deep learning and novelty detection
Charles Richter and Nicholas Roy · 2017
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2017
Closest in time.