Fetching the paper…
Reading the bibliography…
Dealing with high variance is a significant challenge in model-free reinforcement learning (RL).
State-space solutions to standard H/sub 2/ and H/sub infinity / control problems
Doyle, J., Glover, K., Khargonekar, P., and Francis, B · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Robust and Optimal Control
Doyle, J · 1996
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R., McAllester, D., Singh, S. P., and Mansour, Y · 1999
Earlier work this paper cites.
Reinforcement learning in POMDP’s via direct gradient ascent
Baxter, J. and Bartlett, P · 2000
Earlier work this paper cites.
Nonlinear Systems (Third Edition)
Khalil, H. K · 2000
Earlier work this paper cites.
The Optimal Reward Baseline for Gradient-Based Reinforcement Learning
Weaver, L. and Tao, N · 2001
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
Perkins, T. J. and Barto, A. G · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, E., Bartlett, P., and Baxter, J · 2004
Earlier work this paper cites.
Multiple model adaptive control with mixing
Kuipers, M. and Ioannou, P · 2010
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Zhao, T., Hachiya, H., Niu, G., and Sugiyama, M · 2012
Earlier work this paper cites.
Learning of closed-loop motion control
Farshidian, F., Neunert, M., and Buchli, J · 2014
Earlier work this paper cites.
Deterministic Policy Gradient Algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
TORCS, The Open Racing Car Simulator
Wymann, B., Espié, E., Guionneau, C., Dimitrakakis, C., Coulom, R., and Sumner, A · 2014
Cited alongside, same era.
A Comprehensive Survey on Safe Reinforcement Learning
García, J. and Fernández, F · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
Regularized Policy Gradients : Direct Variance Reduction in Policy Gradient Estimation
Zhao, T., Niu, G., Xie, N., Yang, J., and Sugiyama, M · 2015
Cited alongside, same era.
Benchmarking Deep Reinforcement Learning for Continuous Control
Duan, Y., Chen, X., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Smooth Imitation Learning for Online Sequence Prediction
Le, H., Kang, A., Yue, Y., and Carr, P · 2016
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Measuring and Regularizing Networks in Function Space
Benjamin, A. S., Rolnick, D., and Kording, K · 2018
Later among the works it cites.
A Lyapunov-based Approach to Safe Reinforcement Learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Later among the works it cites.
Experimental validation of connected automated vehicle design among human-driven vehicles
Ge, J. I., Avedisov, S. S., He, C. R., Qin, W. B., Sadeghpour, M., and Orosz, G · 2018
Later among the works it cites.
Divide-and-conquer reinforcement learning
Ghosh, D., Singh, A., Rajeswaran, A., Kumar, V., and Levine, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Cited alongside, same era.
Constrained Policy Optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Cited alongside, same era.
Safe Model-based Reinforcement Learning with Stability Guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A. P., and Krause, A · 2017
Cited alongside, same era.
Reproducibility of Benchmarked Deep Reinforcement Learning of Tasks for Continuous Control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D · 2017
Cited alongside, same era.
Later among the works it cites.
Deep Reinforcement Learning that Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
Residual Reinforcement Learning for Robot Control
Johannink, T., Bahl, S., Nair, A., Luo, J., Kumar, A., Loskyll, M., Aparicio Ojea, J., Solowjow, E., and Levine, S · 2018
Later among the works it cites.
Temporal Regularization for Markov Decision Process
Thodoroff, P., Durand, A., Pineau, J., and Precup, D · 2018
Later among the works it cites.
Programmatically interpretable reinforcement learning
Verma, A., Murali, V., Singh, R., Kohli, P., and Chaudhuri, S · 2018
Later among the works it cites.
Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Wu, C., Rajeswaran, A., Duan, Y., Kumar, V., Bayen, A. M., Kakade, S., Mordatch, I., and Abbeel, P · 2018
Later among the works it cites.
Batch policy learning under constraints
Le, H. M., Voloshin, C., and Yue, Y · 2019
Closest in time.
A Tour of Reinforcement Learning: The View from Continuous Control
Recht, B · 2019
Closest in time.