Fetching the paper…
Reading the bibliography…
The practical usage of reinforcement learning agents is often bottlenecked by the duration of training time.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
On a Connection between Importance Sampling and the Likelihood Ratio Policy Gradient
Tang Jie and Pieter Abbeel · 2010
Earlier work this paper cites.
Continuous Control with Deep Reinforcement Learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Massively Parallel Methods for Deep Reinforcement Learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Earlier work this paper cites.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Sample Efficient Actor-Critic with Experience Replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Cited alongside, same era.
Reinforcement Learning Coach, December 2017
Itai Caspi, Gal Leibovich, Gal Novik, and Shadi Endrawis · 2017
Cited alongside, same era.
AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control
Seungyul Han and Youngchul Sung · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Openai Spinning Up
Joshua Achiam · 2018
Cited alongside, same era.
SURREAL: Open-Source Reinforcement Learning Framework and Robot Manipulation Benchmark
Linxi Fan, Yuke Zhu, Jiren Zhu, Zihua Liu, Orien Zeng, Anchit Gupta, Joan Creus-Costa, Silvio Savarese, and Li Fei-Fei · 2018
Later among the works it cites.
Distributed Prioritized Experience Replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Later among the works it cites.
RLlib: Abstractions for Distributed Reinforcement Learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph E. Gonzalez, Michael I. Jordan, and Ion Stoica · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Trust Region Policy Optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz
Cited in the paper.
High-Dimensional Continuous Control using Generalized Advantage Estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel
Cited in the paper.
Seungyul Han and Youngchul Sung · 2019
Closest in time.