Fetching the paper…
Reading the bibliography…
Proximal Policy Optimization (PPO) is a highly popular model-free reinforcement learning (RL) approach.
Estimation of distribution algorithms: A new tool for evolutionary computation
Pedro Larrañaga and Jose A Lozano, · 2001
Earlier work this paper cites.
“A tutorial on the cross-entropy method,”
Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein, · 2005
Earlier work this paper cites.
“The cma evolution strategy: a comparing review,”
Nikolaus Hansen, · 2006
Earlier work this paper cites.
“Reinforcement learning in continuous action spaces,”
Hado van Hasselt and Marco A Wiering, · 2007
Earlier work this paper cites.
“Benchmarking a weighted negative covariance matrix update on the bbob-2010 noiseless testbed,”
Nikolaus Hansen and Raymond Ros, · 2010
Earlier work this paper cites.
“Guided policy search,”
Sergey Levine and Vladlen Koltun, · 2013
Earlier work this paper cites.
“High-dimensional continuous control using generalized advantage estimation,”
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel, · 2015
Earlier work this paper cites.
“Trust region policy optimization,”
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz, · 2015
Earlier work this paper cites.
“The cma evolution strategy: A tutorial,”
Nikolaus Hansen, · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba, · 2016
Cited alongside, same era.
“Proximal policy optimization algorithms,”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, · 2017
Cited alongside, same era.
“Limited-memory matrix adaptation for large scale black-box optimization,”
Ilya Loshchilov, Tobias Glasmachers, and Hans-Georg Beyer, · 2017
Cited alongside, same era.
“Roboschool,”
OpenAI, · 2017
Cited alongside, same era.
“Regularizing sampled differential dynamic programming,”
Joose Rajamäki and Perttu Hämäläinen, · 2018
Closest in time.
“Maximum a posteriori policy optimisation,”
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller, · 2018
Closest in time.
“Relative entropy regularized policy iteration,”
Abbas Abdolmaleki, Jost Tobias Springenberg, Jonas Degrave, Steven Bohez, Yuval Tassa, Dan Belov, Nicolas Heess, and Martin Riedmiller, · 2018
Closest in time.
“Deep reinforcement learning in a handful of trials using probabilistic dynamics models,”
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine, · 2018
Closest in time.
“Drecon: data-driven responsive control of physics-based characters,”
Kevin Bergamin, Simon Clavet, Daniel Holden, and James Richard Forbes, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,”
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne, · 2018
Cited alongside, same era.
“Intelligent middle-level game control,”
Amin Babadi, Kourosh Naderi, and Perttu Hämäläinen, · 2018
Cited alongside, same era.
Closest in time.
“Advantage-weighted regression: Simple and scalable off-policy reinforcement learning,”
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine, · 2019
Closest in time.
“Implementation matters in deep rl: A case study on ppo and trpo,”
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry, · 2019
Closest in time.