Fetching the paper…
Reading the bibliography…
Mirror descent (MD), a well-known first-order method in constrained convex optimization, has recently been shown as an important tool to analyze trust-region algorithms in reinforcement learning (RL).
Problem complexity and method efficiency in optimization (A. Nemirovsky and D. Yudin)
Blair, C · 1985
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Kappen, H · 2005
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Todorov, E · 2006
Earlier work this paper cites.
Dynamic policy programming
Azar, M., Gómez, V., and Kappen, H · 2012
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Earlier work this paper cites.
First-order methods in optimization
Beck, A · 2017
Cited alongside, same era.
OpenAI baselines, 2017
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
A theory of regularized Markov decision processes
Geist, M., Scherrer, B., and Pietquin, O · 2019
Later among the works it cites.
Introduction to online convex optimization
Hazan, E · 2019
Later among the works it cites.
Tsallis reinforcement learning: A unified framework for maximum entropy reinforcement learning
Lee, K., Kim, S., Lim, S., Choi, S., and Oh, S · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z · 2019
Later among the works it cites.
On principled entropy exploration in policy optimization
Mei, J., Xiao, C., Huang, R., Schuurmans, D., and Muller, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Sparse Markov decision processes with causal sparse Tsallis entropy regularization for reinforcement learning
Lee, K., Choi, S., and Oh, S · 2018
Cited alongside, same era.
Path consistency learning in Tsallis entropy regularized mdps
Nachum, O., Chow, Y., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Trust-PCL: An off-policy trust region method for continuous control
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D
Cited in the paper.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D
Cited in the paper.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P
Cited in the paper.
Wang, Y., He, H., and Tan, X · 2019
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Shani, L., Efroni, Y., and Mannor, S · 2020
Closest in time.