Fetching the paper…
Reading the bibliography…
Benchmarks play a crucial role in the development and analysis of reinforcement learning (RL) algorithms.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
Evolutionary robotics and the radical envelope-of-noise hypothesis
Jakobi, N · 1997
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
Lehman, J. and Stanley, K · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. G · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Open-endedness: The last grand challenge you’ve never heard of
Stanley, K. O., Lehman, J., and Soros, L · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A. J., and Klimov, O · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2018
Cited alongside, same era.
The starcraft multi-agent challenge
Samvelyan, M., Rashid, T., De Witt, C. S., Farquhar, G., Nardelli, N., Rudner, T. G., Hung, C.-M., Torr, P. H., Foerster, J., and Whiteson, S · 2019
Cited alongside, same era.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Cited alongside, same era.
Minihack the planet: A sandbox for open-ended reinforcement learning research, 2021
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
Insights from the neurips 2021 nethack challenge
Hambro, E. et al · 2022
Later among the works it cites.
Exploration via elliptical episodic bonuses
Henaff, M., Raileanu, R., Jiang, M., and Rocktäschel, T · 2022
Later among the works it cites.
Grounding aleatoric uncertainty for unsupervised environment design
Jiang, M., Dennis, M., Parker-Holder, J., Lupu, A., Küttler, H., Grefenstette, E., Rocktäschel, T., and Foerster, J · 2022
Later among the works it cites.
gymnax: A JAX-based reinforcement learning environment library, 2022
Lange, R. T · 2022
Later among the works it cites.
Discovered policy optimisation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
You, K., Long, M., Wang, J., and Jordan, M. I · 2019
Cited alongside, same era.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Young, K. and Tian, T · 2019
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Bard, N., Foerster, J. N., Chandar, S., Burch, N., Lanctot, M., Song, H. F., Parisotto, E., Dumoulin, V., Moitra, S., Hughes, E., et al · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Cited alongside, same era.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvári, C., Singh, S., Van Roy, B., Sutton, R., Silver, D., and van Hasselt, H · 2020
Cited alongside, same era.
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2021
Cited alongside, same era.
Lu, C., Kuba, J., Letcher, A., Metz, L., Schroeder de Witt, C., and Foerster, J · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J., Grefenstette, E., and Rocktäschel, T · 2022
Later among the works it cites.
Jumanji: a diverse suite of scalable reinforcement learning environments in jax, 2023
Bonnet, C., Luo, D., Byrne, D., Surana, S., Coyette, V., Duckworth, P., Midgley, L. I., Kalloniatis, T., Abramowitz, S., Waters, C. N., Smit, A. P., Grinsztajn, N., Sob, U. A. M., Mahjoub, O., Tegegn, E., Mimouni, M. A., Boige, R., de Kock, R., Furelos-Blanco, D., Le, V., Pretorius, A., and Laterre, A · 2023
Later among the works it cites.
Chevalier-Boisvert, M., Dai, B., Towers, M., de Lazcano, R., Willems, L., Lahlou, S., Pal, S., Castro, P. S., and Terry, J · 2023
Later among the works it cites.
A study of global and episodic bonuses for exploration in contextual mdps
Henaff, M., Jiang, M., and Raileanu, R · 2023
Later among the works it cites.
minimax: Efficient baselines for autocurricula in jax
Jiang, M., Dennis, M., Grefenstette, E., and Rocktäschel, T · 2023
Later among the works it cites.
Pgx: Hardware-accelerated parallel game simulators for reinforcement learning
Koyamada, S., Okano, S., Nishimori, S., Murata, Y., Habara, K., Kita, H., and Ishii, S · 2023
Later among the works it cites.
XLand-minigrid: Scalable meta-reinforcement learning environments in JAX
Nikulin, A., Kurenkov, V., Zisman, I., Sinii, V., Agarkov, A., and Kolesnikov, S · 2023
Later among the works it cites.
Jaxmarl: Multi-agent rl environments in jax
Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C. S., Souly, A., Bandyopadhyay, S., Samvelyan, M., Jiang, M., Lange, R. T., Whiteson, S., Lacerda, B., Hawes, N., Rocktaschel, T., Lu, C., and Foerster, J. N · 2023
Later among the works it cites.
Jaxued: A simple and useable ued library in jax
Coward, S., Beukman, M., and Foerster, J · 2024
Closest in time.
Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning
Ellis, B., Cook, J., Moalla, S., Samvelyan, M., Sun, M., Mahajan, A., Foerster, J., and Whiteson, S · 2024
Closest in time.