Fetching the paper…
Reading the bibliography…
Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate.
The Fokker-Planck equation
Risken, H · 1989
Earlier work this paper cites.
Deterministic diffusion of particles
Russo, G · 1990
Earlier work this paper cites.
Q Q -learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Error analysis for implicit approximations to solutions to Cauchy problems
Rulla, J · 1996
Earlier work this paper cites.
The variational formulation of the Fokker-Planck equation
Jordan, R., Kinderlehrer, D., and Otto, F · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Minimizing movements and evolution problems in Euclidean spaces
Gobbino, M · 1999
Earlier work this paper cites.
Vortex methods
Cottet, G. H. and Koumoutsakos, P. D · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Gradient Flows in Metric Spaces and in the Space of Probability Measures
Ambrosio, L., Gigli, N., and Savaré, G · 2005
Earlier work this paper cites.
Optimal transport: old and new
Villani, C · 2008
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
Hinton, G. E., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
Hernández-Lobato, J. M. and Adams, R. P · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A., Veness, J., Bellemare, M., Graves, A., Riedmiller, M., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Later among the works it cites.
Carrillo, J. A., Craig, K., and Patacchini, F. S · 2017
Later among the works it cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Deep reinforcement learning: An overview
Li, Y · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Dropout as a Bayesian approximation: representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
VIME: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Liu, Y., Ramachandran, P., Liu, Q., and Peng, J · 2017
Later among the works it cites.
Trust-PCL: An off-policy trust region method for continuous control
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
On wasserstein reinforcement learning and the fokker-planck equation
Richemond, P. H. and Maginnis, B · 2017
Later among the works it cites.
A unified particle-optimization framework for scalable Bayesian sampling
Chen, C., Zhang, R., Wang, W., Li, B., and Chen, L · 2018
Closest in time.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and D. Meger, D · 2018
Closest in time.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Closest in time.
Variational inference and model selection with generalized evidence bounds
Tao, C., Chen, L., Zhang, R., Henao, R., and Carin, L · 2018
Closest in time.
Learning structural weight uncertainty for sequential decision-making
Zhang, R., Li, C., Chen, C., and Carin, L · 2018
Closest in time.