Efficient sample reuse in policy gradients with parameter-based exploration
Tingting Zhao, Hirotaka Hachiya, Voot Tangkaratt, Jun Morimoto, and Masashi Sugiyama · 2013
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Rényi divergence and kullback-leibler divergence
Tim Van Erven and Peter Harremos · 2014
Later among the works it cites.
Concentration inequalities for sums
Bernard Bercu, Bernard Delyon, and Emmanuel Rio · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Original
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Original
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Later among the works it cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Later among the works it cites.
High-confidence off-policy evaluation
Philip S Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Original
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Importance sampling for fair policy selection
Shayan Doroudi, Philip S Thomas, and Emma Brunskill · 2017
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Zhaohan Guo, Philip S Thomas, and Emma Brunskill · 2017
Later among the works it cites.
Effective sample size for importance sampling based on discrepancy measures
Luca Martino, Víctor Elvira, and Francisco Louzada · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Original
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
Original
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.