Fetching the paper…
Reading the bibliography…
Policy gradient methods have been successfully applied to many complex reinforcement learning problems.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Reinforcement driven information acquisition in non-deterministic environments
Storck, J., Hochreiter, S., and Schmidhuber, J. (1995) · 1995
Earlier work this paper cites.
An overview of the simultaneous perturbation method for efficient optimization
Spall, J. C. (1998) · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2002) · 2002
Earlier work this paper cites.
The cross entropy method for fast policy search
Mannor, S., Rubinstein, R. Y., and Gat, Y. (2003) · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S. P., Barto, A. G., and Chentanez, N. (2004) · 2004
Earlier work this paper cites.
The cma evolution strategy: a comparing review
Hansen, N. (2006) · 2006
Earlier work this paper cites.
Information theory―the bridge connecting bounded rational game theory and statistical physics
Wolpert, D. H. (2006) · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. v. d. and Hinton, G. (2008) · 2008
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
Information, utility and bounded rationality
Ortega, D. A. and Braun, P. A. (2011) · 2011
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Cited alongside, same era.
Accelerating t-sne using tree-based algorithms
Van Der Maaten, L. (2014) · 2014
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J. (2015) · 2015
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Later among the works it cites.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Liu, Q. and Wang, D. (2016) · 2016
Later among the works it cites.
Learning to navigate in complex environments
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al. (2016) · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Later among the works it cites.
Improving policy gradient by exploring under-appreciated rewards
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De Maria, A., Panneershelvam, V., Suleyman, M., Beattie, C., Petersen, S., et al. (2015) · 2015
Cited alongside, same era.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al. (2016) · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., and Levine, S. (2016) · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P. (2015a)
Cited in the paper.
Nachum, O., Norouzi, M., and Schuurmans, D. (2016) · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Later among the works it cites.
Muticore-tsne
Ulyanov, D. (2016) · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N. (2016) · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Closest in time.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R. (2017) · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., and Sutskever, I. (2017) · 2017
Closest in time.