Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) enables robots to learn skills from interactions with the real world.
On the theory of the brownian motion
G. E. Uhlenbeck and L. S. Ornstein · 1930
Earlier work this paper cites.
Shakey the robot
N. J. Nilsson · 1984
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
State-dependent exploration for policy gradient methods
T. Rückstieß, M. Felder, and J. Schmidhuber · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
J. Kober and J. R. Peters · 2009
Earlier work this paper cites.
Exploring parameter space in reinforcement learning
T. Rückstiess, F. Sehnke, T. Schaul, D. Wierstra, Y. Sun, and J. Schmidhuber · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
F. Sehnke, C. Osendorfer, T. Rückstieß, A. Graves, J. Peters, and J. Schmidhuber · 2010
Earlier work this paper cites.
Robot skill learning: From reinforcement learning to evolution strategies
F. Stulp and O. Sigaud · 2013
Earlier work this paper cites.
A survey on policy search for robotics
M. P. Deisenroth, G. Neumann, J. Peters, et al · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
A structurally flexible humanoid spine based on a tendon-driven elastic continuum
J. Reinecke, B. Deutschmann, and D. Fehrenbach · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Stable baselines
A. Hill, A. Raffin, M. Ernestus, A. Gleave, A. Kanervisto, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2018
Later among the works it cites.
Rl baselines zoo
A. Raffin · 2018
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Generalized exploration in policy search
H. van Hoof, D. Tanneberg, and J. Peters · 2017
Cited alongside, same era.
Position control of an underactuated continuum mechanism using a reduced nonlinear model
B. Deutschmann, A. Dietrich, and C. Ott · 2017
Cited alongside, same era.
F. Such, V. Madhavan, E. Conti, J. Lehman, K. Stanley, and J. Clune · 2017
Cited alongside, same era.
Time limits in reinforcement learning
F. Pardo, A. Tavakoli, V. Levdik, and P. Kormushev · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
A. Rajeswaran, K. Lowrey, E. V. Todorov, and S. M. Kakade · 2017
Cited alongside, same era.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2018
Cited alongside, same era.
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah · 2019
Later among the works it cites.
Autoregressive policies for continuous control deep reinforcement learning
D. Korenkevych, A. R. Mahmood, G. Vasan, and J. Bergstra · 2019
Later among the works it cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2019
Later among the works it cites.
Stable baselines3
A. Raffin, A. Hill, M. Ernestus, A. Gleave, A. Kanervisto, and N. Dormann · 2019
Later among the works it cites.
Six-dof pose estimation for a tendon-driven continuum mechanism without a deformation model
B. Deutschmann, M. Chalon, J. Reinecke, M. Maier, and C. Ott · 2019
Later among the works it cites.
Policy search in continuous action domains: an overview
O. Sigaud and F. Stulp · 2019
Later among the works it cites.
Learning to drive smoothly in minutes
A. Raffin and R. Sokolkov · 2019
Later among the works it cites.
Optuna: A next-generation hyperparameter optimization framework
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama · 2019
Later among the works it cites.
Learning to walk in the real world with minimal human effort
S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan · 2020
Closest in time.
Continuous-discrete reinforcement learning for hybrid control in robotics
M. Neunert, A. Abdolmaleki, M. Wulfmeier, T. Lampe, J. T. Springenberg, R. Hafner, F. Romano, J. Buchli, N. Heess, and M. Riedmiller · 2020
Closest in time.
Rl baselines3 zoo
A. Raffin · 2020
Closest in time.
Implementation matters in deep {rl}: A case study on {ppo} and {trpo}
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry · 2020
Closest in time.
Regularizing action policies for smooth control with reinforcement learning
S. Mysore, B. Mabsout, R. Mancuso, and K. Saenko · 2021
Closest in time.