Fetching the paper…
Reading the bibliography…
We make three contributions toward better understanding policy gradient methods in the tabular setting.
Une propriété topologique des sous-ensembles analytiques réels
Łojasiewicz, S · 1963
Earlier work this paper cites.
Some modified matrix eigenvalue problems
Golub, G. H · 1973
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Nesterov, Y. and Polyak, B. T · 2006
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Rate of convergence to equilibrium and Łojasiewicz-type estimates
Bárta, T · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Optimality and approximation with policy gradient methods in Markov decision processes, 2019
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2019
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods, 2019
Bhandari, J. and Russo, D · 2019
Later among the works it cites.
On principled entropy exploration in policy optimization
Mei, J., Xiao, C., Huang, R., Schuurmans, D., and Müller, M · 2019
Later among the works it cites.
Maximum entropy monte-carlo planning
Xiao, C., Huang, R., Mei, J., Schuurmans, D., and Müller, M · 2019
Later among the works it cites.
Making sense of reinforcement learning and probabilistic inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Nesterov, Y · 2018
Cited alongside, same era.
Convergence of cubic regularization for nonconvex optimization under KL property
Zhou, Y., Wang, Z., and Liang, Y · 2018
Cited alongside, same era.
O’Donoghue, B., Osband, I., and Ionescu, C · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Shani, L., Efroni, Y., and Mannor, S · 2020
Closest in time.
A short note on soft-max and policy gradients in bandits problems
Walton, N · 2020
Closest in time.