The NetHack Learning Environment
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T. (2020) · 2020
Later among the works it cites.
Reinforcement learming with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A. (2020) · 2020
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2020) · 2020
Later among the works it cites.
On gradient-based learning in continuous games
Mazumdar, E., Ratliff, L. J., and Sastry, S. S. (2020) · 2020
Later among the works it cites.
Active domain randomization
Mehta, B., Diaz, M., Golemo, F., Pal, C. J., and Paull, L. (2020) · 2020
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S. (2020) · 2020
Later among the works it cites.
Automated curriculum generation through setter-solver interactions
Racaniere, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., and Lillicrap, T. (2020) · 2020
Later among the works it cites.
Increasing generality in machine learning through procedural content generation
Risi, S. and Togelius, J. (2020) · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B. (2020) · 2020
Later among the works it cites.
The unexpected consequence of incremental design changes
Sturtevant, N., Decroocq, N., Tripodi, A., and Guzdial, M. (2020) · 2020
Later among the works it cites.
Enhanced POET: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
Wang, R., Lehman, J., Rawal, A., Zhi, J., Li, Y., Clune, J., and Stanley, K. (2020) · 2020
Later among the works it cites.
Augmented world models facilitate zero-shot dynamics generalization from a single offline environment
Ball, P. J., Lu, C., Parker-Holder, J., and Roberts, S. J. (2021) · 2021
Later among the works it cites.
Learning with AMIGo: Adversarially motivated intrinsic goals
Campero, A., Raileanu, R., Kuttler, H., Tenenbaum, J. B., Rocktäschel, T., and Grefenstette, E. (2021) · 2021
Later among the works it cites.
Self-paced context evaluation for contextual reinforcement learning
Eimer, T., Biedenkapp, A., Hutter, F., and Lindauer, M. (2021) · 2021
Later among the works it cites.
Adaptive procedural task generation for hard-exploration problems
Fang, K., Zhu, Y., Savarese, S., and Li, F.-F. (2021) · 2021
Later among the works it cites.
On the importance of environments in human-robot coordination
Fontaine, M. C., Hsu, Y., Zhang, Y., Tjanaka, B., and Nikolaidis, S. (2021) · 2021
Later among the works it cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability
Original
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S. (2021) · 2021
Later among the works it cites.
Adversarial environment generation for learning to navigate the web
Gur, I., Jaques, N., Malta, K., Tiwari, M., Lee, H., and Faust, A. (2021) · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Original
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R. (2021) · 2021
Later among the works it cites.
Deep learning for procedural content generation
Liu, J., Snodgrass, S., Khalifa, A., Risi, S., Yannakakis, G. N., and Togelius, J. (2021) · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D. (2021) · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
OpenAI, O., Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D’Sa, R., Petron, A., de Oliveira Pinto, H. P., Paino, A., Noh, H., Weng, L., Yuan, Q., Chu, C., and Zaremba, W. (2021) · 2021
Later among the works it cites.
Automatic data augmentation for generalization in deep reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. (2021) · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021) · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Original
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., and Czarnecki, W. M. (2021) · 2021
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Original
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M. (2022) · 2022
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Closest in time.