Fetching the paper…
Reading the bibliography…
Due to the realization that deep reinforcement learning algorithms trained on high-dimensional tasks can strongly overfit to their training environments, there have been several studies that investigated the generalization performance of these algorithms.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
Schmidhuber, J · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
Graves, A., Bellemare, M. G., Menick, J., Munos, R., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Rllib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Moritz, P., Nishihara, R., Fox, R., Goldberg, K., Gonzalez, J. E., Jordan, M. I., and Stoica, I · 2017
Earlier work this paper cites.
A simple neural attentive meta-learner
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Starcraft ii: A new challenge for reinforcement learning
Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A. S., Yeo, M., Makhzani, A., Küttler, H., Agapiou, J., Schrittwieser, J., et al · 2017
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2019
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Farebrother, J., Machado, M. C., and Bowling, M · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Evolved policy gradients
Houthooft, R., Chen, Y., Isola, P., Stadie, B., Wolski, F., Ho, O. J., and Abbeel, P · 2018
Cited alongside, same era.
Illuminating generalization in deep reinforcement learning through procedural level generation
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Cited alongside, same era.
Promp: Proximal meta-policy search
Rothfuss, J., Lee, D., Clavera, I., Asfour, T., and Abbeel, P · 2018
Cited alongside, same era.
A dissection of overfitting and generalization in continuous reinforcement learning
Zhang, A., Ballas, N., and Pineau, J
Cited in the paper.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S
Cited in the paper.
Kirsch, L., van Steenkiste, S., and Schmidhuber, J · 2019
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Later among the works it cites.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Later among the works it cites.