Fetching the paper…
Reading the bibliography…
Policy gradient methods ignore the potential value of adjusting environment variables: unobservable state features that are randomly determined by the environment in a physical setting, but are controllable in a simulator.
Likelihood ratio gradient estimation for stochastic systems
Glynn, P. W · 1990
Earlier work this paper cites.
A statistical method for global optimization
Cox, D. D. and John, S · 1992
Earlier work this paper cites.
SDO: A statistical method for global optimization
Cox, D. D. and John, S · 1997
Earlier work this paper cites.
Reinforcement Learning : An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Sequential design of computer experiments to minimize integrated response functions
Williams, B. J., Santner, T. J., and Notz, W. I · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 2002
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Rasmussen, C. E. and Williams, C. K. I · 2005
Earlier work this paper cites.
Policy gradient methods for robotics
Peters, J. and Schaal, S · 2006
Earlier work this paper cites.
Reinforcement learning in the presence of rare events
Frank, J., Mannor, S., and Precup, D · 2008
Earlier work this paper cites.
Contextual gaussian process bandit optimization
Krause, A. and Ong, C. S · 2011
Earlier work this paper cites.
Bayesian Optimization in High Dimensions via Random Embeddings
Wang, Z., Zoghi, M., Hutter, F., Matheson, D., and De Freitas, N · 2013
Cited alongside, same era.
Interactive control of diverse complex characters with neural networks
Mordatch, I., Lowrey, K., Andrew, G., Popovic, Z., and Todorov, E. V · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gómez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Learning to learn without gradient descent by gradient descent
Chen, Y., Hoffman, M. W., Colmenarejo, S. G., Denil, M., Lillicrap, T. P., Botvinick, M., and de Freitas, N · 2017
Later among the works it cites.
Offer: Off-environment reinforcement learning
Ciosek, K. and Whiteson, S · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Foerster, J., Nardelli, N., Farquhar, G., Torr, P., Kohli, P., and Whiteson, S · 2017
Later among the works it cites.
Automated curriculum learning for neural networks
Graves, A., Bellemare, M. G., Menick, J., Munos, R., and Kavukcuoglu, K · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Bayesian optimization for automated model selection
Malkomes, G., Schaff, C., and Garnett, R · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Cited alongside, same era.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 2016
Cited alongside, same era.
EPOpt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Levine, S., and Ravindran, B
Cited in the paper.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M
Cited in the paper.
Later among the works it cites.
Fast Information-theoretic Bayesian Optimisation
Ru, B., McLeod, M., Granziol, D., and Osborne, M. A · 2017
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Closest in time.
Alternating optimisation and quadrature for robust control
Paul, S., Chatzilygeroudis, K., Ciosek, K., Mouret, J.-B., Osborne, M., and Whiteson, S · 2018
Closest in time.
Bayesian Optimization with Expensive Integrands
Toscano-Palmerin, S. and Frazier, P. I · 2018
Closest in time.