Fetching the paper…
Reading the bibliography…
Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), as the widely employed policy based reinforcement learning (RL) methods, are prone to converge to a sub-optimal solution as they limit the policy representation to a particular parametric distribution class.
Methods of conjugate gradients for solving linear systems
Magnus R. Hestenes and Eduard Stiefel · 1952
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G. Barto, Richard S. Sutton, and Charles W. Anderson · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Global optimization by basin-hopping and the lowest energy structures of Lennard-Jones clusters containing up to 110 atoms
David Wales and Jonathan Doye · 1998
Earlier work this paper cites.
The MAXQ method for hierarchical reinforcement learning
Thomas G. Dietterich · 1998
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Global optimization on funneling landscapes
Robert Leary · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
On choosing and bounding probability metrics
Alison L. Gibbs and Francis Edward Su · 2002
Earlier work this paper cites.
Two-phase methods for global optimization
Fabio Schoen · 2002
Earlier work this paper cites.
Using policy gradient reinforcement learning on autonomous robot controllers
Gregory Z. Grudic, Vijay Kumar, and Lyle H. Ungar · 2003
Earlier work this paper cites.
Support vector classification with input data uncertainty
Jinbo Bi and Tong Zhang · 2005
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Monte Carlo basin paving: An improved global optimization method
Lixin Zhan, Jeff Chen, and Wing-Ki Liu · 2006
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M. Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
Erick Delage and Yinyu Ye · 2010
Earlier work this paper cites.
Distributionally robust optimization and its tractable approximations
Joel Goh and Melvyn Sim · 2010
Cited alongside, same era.
Distributionally robust joint chance constraints with second-order moment information
Steve Zymler, Daniel Kuhn, and Berç Rustem · 2011
Cited alongside, same era.
Kullback-Leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L. Jeff Hong · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Robust solutions of optimization problems affected by uncertain probabilities
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Po-Wei Chou, Daniel Maturana, and Sebastian Scherer · 2017
Later among the works it cites.
Ambiguous joint chance constraints under mean and dispersion information
Grani A. Hanasusanto, Vladimir Roitch, Daniel Kuhn, and Wolfram Wiesemann · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aharon Ben-Tal, Dick den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
RLPy: A value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H. Klein, William Dabney, and Jonathan P. How · 2015
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Later among the works it cites.
Multi-step reinforcement learning: A unifying algorithm
Kristopher De Asis, J. Fernando Hernandez-Garcia, G. Zacharias Holland, and Richard S. Sutton · 2017
Later among the works it cites.
OpenAI baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Later among the works it cites.
Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations
Peyman Mohajerin Esfahani and Daniel Kuhn · 2018
Later among the works it cites.
Data-driven risk-averse stochastic optimization with Wasserstein metric
Chaoyue Zhao and Yongpei Guan · 2018
Later among the works it cites.
Stable baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Learning to score behaviors for guided policy optimization
Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Anna Choromanska, Krzysztof Choromanski, and Michael I. Jordan · 2019
Later among the works it cites.
Distributional policy optimization: An alternative approach for continuous control
Chen Tessler, Guy Tennenholtz, and Shie Mannor · 2019
Later among the works it cites.
Optimistic distributionally robust optimization for nonparametric likelihood approximation
Viet Anh Nguyen, Soroosh Shafieezadeh Abadeh, Man-Chung Yue, Daniel Kuhn, and Wolfram Wiesemann · 2019
Later among the works it cites.