Fetching the paper…
Reading the bibliography…
Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO) are among the most successful policy gradient approaches in deep reinforcement learning (RL).
Truly proximal policy optimization
Wang, Y., He, H., and Tan, X. (2019b) · 1903
Earlier work this paper cites.
Deep conservative policy iteration
Vieillard, N., Pietquin, O., and Geist, M. (2019) · 1906
Earlier work this paper cites.
Imitation learning via off-policy distribution matching
Kostrikov, I., Nachum, O., and Tompson, J. (2019) · 1912
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. (2019b) · 1912
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
Donsker, M. D. and Varadhan, S. S. (1983) · 1983
Earlier work this paper cites.
Markov decision processes
Puterman, M. L. (1990) · 1990
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
On surrogate loss functions and f-divergences
Nguyen, X., Wainwright, M. J., Jordan, M. I., et al. (2009) · 2009
Earlier work this paper cites.
On integral probability metrics, ϕ \phi -divergences and binary classification
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Schölkopf, B., and Lanckriet, G. R. (2009) · 2009
Cited alongside, same era.
Dynamic policy programming
Azar, M. G., Gómez, V., and Kappen, H. J. (2012) · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Safe policy iteration
Pirotta, M., Restelli, M., Pecorino, A., and Calandriello, D. (2013) · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Parametric adversarial divergences are good task losses for generative modeling
Huang, G., Berard, H., Touati, A., Gidel, G., Vincent, P., and Lacoste-Julien, S. (2017) · 2017
Later among the works it cites.
On convergence and stability of gans
Kodali, N., Abernethy, J., Hays, J., and Kira, Z. (2017) · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M. A. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L. (2019a)
Cited in the paper.
Divergence-augmented policy optimization
Wang, Q., Li, Y., Xiong, J., and Zhang, T. (2019a)
Cited in the paper.
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., and Song, L. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Pytorch implementations of reinforcement learning algorithms
Kostrikov, I. (2018) · 2018
Later among the works it cites.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.