Fetching the paper…
Reading the bibliography…
Environments with procedurally generated content serve as important benchmarks for testing systematic generalization in deep reinforcement learning.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L · 1910
Earlier work this paper cites.
Improving generalization with active learning
Cohn, D., Atlas, L., and Ladner, R · 1994
Earlier work this paper cites.
PCGRL: procedural content generation via reinforcement learning
Khalifa, A., Bontrager, P., Earle, S., and Togelius, J · 2001
Earlier work this paper cites.
Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2004
Earlier work this paper cites.
Learning with AMIGo: Adversarially Motivated Intrinsic Goals
Campero, A., Raileanu, R., Küttler, H., Tenenbaum, J. B., Rocktäschel, T., and Grefenstette, E · 2006
Earlier work this paper cites.
The impact of non-stationarity on generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S · 2006
Earlier work this paper cites.
Cobbe, K., Hilton, J., Klimov, O., and Schulman, J · 2009
Earlier work this paper cites.
Active learning literature survey
Settles, B · 2009
Earlier work this paper cites.
Bebold: Exploration beyond the boundary of explored regions
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Earlier work this paper cites.
Automated curriculum learning for neural networks
Graves, A., Bellemare, M. G., Menick, J., Munos, R., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Kostrikov, I., Szlam, A., and Fergus, R · 2017
Cited alongside, same era.
BabyAI: First steps towards grounded language learning with a human in the loop
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y · 2018
Cited alongside, same era.
Minimalistic gridworld environment for OpenAI Gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Cited alongside, same era.
Teacher–student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2019
Later among the works it cites.
POET: open-ended coevolution of environments and their optimized solutions
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Later among the works it cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S · 2020
Closest in time.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Closest in time.
RIDE: rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R. and Rocktäschel, T · 2020
Closest in time.
Increasing generality in machine learning through procedural content generation
Risi, S. and Togelius, J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Procedural level generation improves generality of deep reinforcement learning
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Cited alongside, same era.
Learning to multi-task by active sampling
Sharma, S., Jha, A. K., Hegde, P., and Ravindran, B · 2018
Cited alongside, same era.
The hanabi challenge: A new frontier for AI research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., Dunning, I., Mourad, S., Larochelle, H., Bellemare, M. G., and Bowling, M · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Cited alongside, same era.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Igl, M., Ciosek, K., Li, Y., Tschiatschek, S., Zhang, C., Devlin, S., and Hofmann, K · 2019
Cited alongside, same era.
Obstacle tower: A generalization challenge in vision, control, and planning
Juliani, A., Khalifa, A., Berges, V.-P., Harper, J., Teng, E., Henry, H., Crespi, A., Togelius, J., and Lange, D · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Cited alongside, same era.
Closest in time.
Improving generalization in reinforcement learning with mixture regularization
Wang, K., Kang, B., Shao, J., and Feng, J · 2020
Closest in time.
Automatic curriculum learning through value disagreement
Zhang, Y., Abbeel, P., and Pinto, L · 2020
Closest in time.
RTFM: generalising to new environment dynamics via reading
Zhong, V., Rocktäschel, T., and Grefenstette, E · 2020
Closest in time.
Automatic data augmentation for generalization in reinforcement learning, 2021
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R · 2021
Closest in time.
Rank the episodes: A simple approach for exploration in procedurally-generated environments
Zha, D., Ma, W., Yuan, L., Hu, X., and Liu, J · 2021
Closest in time.
Leveraging Procedural Generation to Benchmark Reinforcement Learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2056
Closest in time.