Fetching the paper…
Reading the bibliography…
The performance of policy gradient methods is sensitive to hyperparameter settings that must be tuned for any new application.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H. (2019) · 1902
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Sutton, R. S. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S. P. (2000) · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A. (2001) · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2001) · 2001
Earlier work this paper cites.
Brochu, E., Cora, V. M., and de Freitas, N. (2010) · 2010
Earlier work this paper cites.
Temporal difference bayesian model averaging: A bayesian perspective on adapting lambda
Downey, C., Sanner, S., et al. (2010) · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. (2010) · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J. S., Bardenet, R., Bengio, Y., and Kégl, B. (2011) · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H. H., and Leyton-Brown, K. (2011) · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. (2012) · 2012
Earlier work this paper cites.
Parallel gaussian process optimization with upper confidence bound and pure exploration
Contal, E., Buffoni, D., Robicquet, A., and Vayatis, N. (2013) · 2013
Cited alongside, same era.
Multi-task bayesian optimization
Swersky, K., Snoek, J., and Adams, R. P. (2013) · 2013
Cited alongside, same era.
Parallelizing exploration-exploitation tradeoffs in gaussian process bandit optimization
Desautels, T., Krause, A., and Burdick, J. W. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
Input warping for bayesian optimization of non-stationary functions
Snoek, J., Swersky, K., Zemel, R., and Adams, R. (2014) · 2014
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Later among the works it cites.
Hyperparameter optimization with approximate gradient
Pedregosa, F. (2016) · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Later among the works it cites.
Parallel bayesian global optimization of expensive functions
Wang, J., Clark, S. C., Liu, E., and Frazier, P. I. (2016) · 2016
Later among the works it cites.
A greedy approach to adapting the trace parameter for temporal difference learning
White, M. and White, A. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Swersky, K., Snoek, J., and Adams, R. P. (2014) · 2014
Cited alongside, same era.
How to discount deep reinforcement learning: Towards new dynamic strategies
François-Lavet, V., Fonteneau, R., and Ernst, D. (2015) · 2015
Cited alongside, same era.
Interactive control of diverse complex characters with neural networks
Mordatch, I., Lowrey, K., Andrew, G., Popovic, Z., and Todorov, E. V. (2015) · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Parallel predictive entropy search for batch global optimization of expensive objective functions
Shah, A. and Ghahramani, Z. (2015) · 2015
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P. (2017) · 2017
Later among the works it cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al. (2017) · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M. (2017) · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
White, M. (2017) · 2017
Later among the works it cites.
Bayesian optimization in alphago
Chen, Y., Huang, A., Wang, Z., Antonoglou, I., Schrittwieser, J., Silver, D., and de Freitas, N. (2018) · 2018
Later among the works it cites.
Treeqn and atreec: Differentiable tree-structured models for deep reinforcement learning
Farquhar, G., Rocktaschel, T., Igl, M., and Whiteson, S. (2018) · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
Igl, M., Zintgraf, L. M., Le, T. A., Wood, F., and Whiteson, S. (2018) · 2018
Later among the works it cites.
Parallelised bayesian optimisation via thompson sampling
Kandasamy, K., Krishnamurthy, A., Schneider, J., and Póczos, B. (2018) · 2018
Later among the works it cites.
Benchmarking reinforcement learning algorithms on real-world robots
Mahmood, A. R., Korenkevych, D., Vasan, G., Ma, W., and Bergstra, J. (2018) · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D. (2018) · 2018
Later among the works it cites.