Fetching the paper…
Reading the bibliography…
Advances in large models, reinforcement learning, and open-endedness have accelerated progress toward autonomous agents that can learn and interact in the real world.
Learning to compete, compromise, and cooperate in repeated general-sum games
Crandall, J. W. and Goodrich, M. A · 2005
Earlier work this paper cites.
Programming in Lua
Ierusalimschy, R · 2006
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
AI2-THOR: An interactive 3D environment for visual AI
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al · 2017
Earlier work this paper cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A., and Munos, R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Earlier work this paper cites.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., et al · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Earlier work this paper cites.
MineRL: A large-scale dataset of minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C., Veloso, M., and Salakhutdinov, R · 2019
Earlier work this paper cites.
ViZDoom Competitions: Playing Doom from Pixels
Wydmuch, M., Kempka, M., and Jaśkowski, W · 2019
Earlier work this paper cites.
Agent57: Outperforming the atari human benchmark
Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z. D., and Blundell, C · 2020
Earlier work this paper cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., et al · 2020
Earlier work this paper cites.
Griddly: A platform for ai research in games
Bamford, C., Huang, S., and Lucas, S · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Earlier work this paper cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., and Levine, S · 2020
Earlier work this paper cites.
The NetHack Learning Environment
Küttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Earlier work this paper cites.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
ThreeDWorld: A platform for interactive multi-modal physical simulation
Gan, C., Schwartz, J., Alter, S., Mrowca, D., Schrimpf, M., Traer, J., Freitas, J. D., Kubilius, J., Bhandwaldar, A., Haber, N., Sano, M., Kim, K., Wang, E., Lingelbach, M., Curtis, A., Feigelis, K. T., Bear, D., Gutfreund, D., Cox, D. D., Torralba, A., DiCarlo, J. J., Tenenbaum, J. B., Mcdermott, J., and Yamins, D. L · 2021
Cited alongside, same era.
Evocraft: A new challenge for open-endedness
Grbic, D., Palm, R. B., Najarro, E., Glanois, C., and Risi, S · 2021
Cited alongside, same era.
Scalable evaluation of multi-agent reinforcement learning with Melting Pot
Leibo, J. Z., Dueñez-Guzman, E. A., Vezhnevets, A., Agapiou, J. P., Sunehag, P., Koster, R., Matyas, J., Beattie, C., Mordatch, I., and Graepel, T · 2021
Cited alongside, same era.
Stable-Baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
Learning diverse risk preferences in population-based self-play
Jiang, Y., Liu, Q., Ma, X., Li, C., Yang, Y., Yang, J., Liang, B., and Zhao, Q · 2024
Closest in time.
Position: Benchmarking is limited in reinforcement learning research
Jordan, S. M., White, A., Silva, B. C. D., White, M., and Thomas, P. S · 2024
Closest in time.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y., and Lee, Y. J · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Self-composing policies for scalable continual reinforcement learning
Malagon, M., Ceberio, J., and Lozano, J. A · 2024
Closest in time.
Craftax: A lightning-fast benchmark for open-ended reinforcement learning
Matthews, M., Beukman, M., Ellis, B., Samvelyan, M., Jackson, M. T., Coward, S., and Foerster, J. N · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MiniHack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., and Rocktäschel, T · 2021
Cited alongside, same era.
PettingZoo: Gym for multi-agent reinforcement learning
Terry, J., Black, B., Grammel, N., Jayakumar, M., Hari, A., Sullivan, R., Santos, L. S., Dieffendahl, C., Horsch, C., Perez-Vicente, R., et al · 2021
Cited alongside, same era.
Continual world: A robotic benchmark for continual reinforcement learning
Wołczyk, M., Zajac, M., Pascanu, R., Kuciński, Ł., and Miłoś, P · 2021
Cited alongside, same era.
ProcTHOR: Large-scale embodied AI using procedural generation
Deitke, M., VanderBilt, E., Herrasti, A., Weihs, L., Ehsani, K., Salvador, J., Han, W., Kolve, E., Kembhavi, A., and Mottaghi, R · 2022
Cited alongside, same era.
MineDojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Cited alongside, same era.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2022
Cited alongside, same era.
The 37 implementation details of proximal policy optimization
Huang, S., Dossa, R. F. J., Raffin, A., Kanervisto, A., and Wang, W · 2022
Cited alongside, same era.
Closest in time.
Position: A call for embodied AI
Paolo, G., Gonzalez-Billandon, J., and Kégl, B · 2024
Closest in time.
Dreaming of many worlds: Learning contextual world models aids zero-shot generalization
Prasanna, S., Farid, K., Rajan, R., and Biedenkapp, A · 2024
Closest in time.
Habitat 3.0: A co-habitat for humans, avatars, and robots
Puig, X., Undersander, E., Szot, A., Cote, M. D., Yang, T.-Y., Partsey, R., Desai, R., Clegg, A., Hlavac, M., Min, S. Y., Vondruš, V., Gervet, T., Berges, V.-P., Turner, J. M., Maksymets, O., Kira, Z., Kalakrishnan, M., Malik, J., Chaplot, D. S., Jain, U., Batra, D., Rai, A., and Mottaghi, R · 2024
Closest in time.
Scaling instructable agents across many simulated worlds
Raad, M. A., Ahuja, A., Barros, C., Besse, F., Bolt, A., Bolton, A., Brownfield, B., Buttimore, G., Cant, M., Chakera, S., et al · 2024
Closest in time.
Reward-free curricula for training robust world models
Rigter, M., Jiang, M., and Posner, I · 2024
Closest in time.
MAMBA: an effective world model approach for meta-reinforcement learning
Rimon, Z., Jurgenson, T., Krupnik, O., Adler, G., and Tamar, A · 2024
Closest in time.
JaxMARL: Multi-agent rl environments in JAX
Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C. S., Souly, A., Bandyopadhyay, S., Samvelyan, M., Jiang, M., Lange, R. T., Whiteson, S., Lacerda, B., Hawes, N., Rocktaschel, T., Lu, C., and Foerster, J. N · 2024
Closest in time.
Gymnasium: A standard interface for reinforcement learning environments
Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al · 2024
Closest in time.
RoboGen: Towards unleashing infinite data for automated robot learning via generative simulation
Wang, Y., Xian, Z., Chen, F., Wang, T.-H., Wang, Y., Fragkiadaki, K., Erickson, Z., Held, D., and Gan, C · 2024
Closest in time.
Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem
Wołczyk, M., Cupiał, B., Ostaszewski, M., Bortkiewicz, M., Zajkac, M., Pascanu, R., Kuciński, Ł., and Miłoś, P · 2024
Closest in time.
OMNI-EPIC: Open-endedness via models of human notions of interestingness with environments programmed in code
Faldor, M., Zhang, J., Cully, A., and Clune, J · 2025
Closest in time.
VoxeLibre, a voxel-based sandbox game for luanti
Fleckenstein, L., Wuzzy, davedevils, and contributors · 2025
Closest in time.
Luanti’s modding API reference
Luanti Team · 2025
Closest in time.
Luanti’s main page
Luanti Team · 2025
Closest in time.
Luanti’s wiki FAQ page
Luanti Wiki · 2025
Closest in time.
ContentDB: a content database for Luanti mods, games, and more
Ward, A · 2025
Closest in time.
Luanti modding book
Ward, A · 2025
Closest in time.