Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) has achieved breakthrough results on many tasks, but agents often fail to generalize beyond the environment they were trained in.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A. W · 1990
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1995
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K · 2001
Earlier work this paper cites.
Robustness in Markov decision problems with uncertain transition matrices
Nilim, A. and Ghaoui, L. E · 2004
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Reinforcement learning in robust Markov decision processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Earlier work this paper cites.
50 years of data science
Donoho, D · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
Optimizing the CVaR via sampling
Tamar, A., Glassner, Y., and Mannor, S · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., et al · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., and Jaśkowski, W · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Learning to learn: Meta-critic networks for sample efficient learning
Sung, F., Zhang, L., Xiang, T., Hospedales, T., and Yang, Y · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2017
Later among the works it cites.
Preparing for the unknown: Learning a universal policy with online system identification
Yu, W., Tan, J., Liu, C. K., and Turk, G · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Cited alongside, same era.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Cited alongside, same era.
OpenAI Baselines
Dhariwal, P., Hesse, C., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Cited alongside, same era.
Steps toward robust artificial intelligence
Dietterich, T. G · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, D. S., and George, D · 2017
Cited alongside, same era.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P · 2018
Closest in time.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Bai, S., Kolter, J. Z., and Koltun, V · 2018
Closest in time.
Learning to adapt: Meta-learning for model-based control
Clavera, I., Nagabandi, A., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C · 2018
Closest in time.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Closest in time.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Closest in time.
Procedural level generation improves generality of deep reinforcement learning
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Closest in time.
Deep learning: A critical appraisal
Marcus, G · 2018
Closest in time.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Closest in time.
Gotta learn fast: A new benchmark for generalization in RL
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Closest in time.
Meta reinforcement learning with latent variable Gaussian processes
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P · 2018
Closest in time.
A dissection of overfitting and generalization in continuous reinforcement learning
Zhang, A., Ballas, N., and Pineau, J · 2018
Closest in time.