Fetching the paper…
Reading the bibliography…
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge.
Guss, W. H., Codel, C., Hofmann, K., Houghton, B., Kuno, N., Milani, S., Mohanty, S., Liebana, D. P., Salakhutdinov, R., Topin, N., et al · 1904
Earlier work this paper cites.
Unit testing: test early, test often
Olan, M · 2003
Earlier work this paper cites.
Open-ended artificial evolution
Standish, R. K · 2003
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Lee, H., et al · 2016
Earlier work this paper cites.
Fighting zombies in minecraft with deep reinforcement learning
Udagawa, H., Narasimhan, T., and Lee, S.-Y · 2016
Earlier work this paper cites.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Shu, T., Xiong, C., and Socher, R · 2017
Earlier work this paper cites.
Open-endedness: The last grand challenge you’ve never heard of
Stanley, K. O., Lehman, J., and Soros, L · 2017
Earlier work this paper cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D., and Mannor, S · 2017
Earlier work this paper cites.
Deep reinforcement learning with model learning and monte carlo tree search in minecraft
Alaniz, S · 2018
Earlier work this paper cites.
Deep reinforcement learning from policy-dependent human feedback
Arumugam, D., Lee, J. K., Saskin, S., and Littman, M. L · 2019
Earlier work this paper cites.
Teacher–student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2019
Earlier work this paper cites.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Trott, A., Zheng, S., Xiong, C., and Socher, R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Earlier work this paper cites.
High precision control and deep learning-based corn stand counting algorithms for agricultural robot
Zhang, Z., Kayacan, E., Thompson, B., and Chowdhary, G · 2020
Earlier work this paper cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Open-ended learning leads to generally capable agents
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., et al · 2021
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2022
Earlier work this paper cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Earlier work this paper cites.
In-ide code generation from natural language: Promise and challenges
Xu, F. F., Vasilescu, B., and Neubig, G · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Agentcoder: Multi-agent-based code generation with iterative testing and optimisation
Huang, D., Bu, Q., Zhang, J. M., Luck, M., and Cui, H · 2023
Cited alongside, same era.
Mini-behavior: A procedurally generated benchmark for long-horizon decision-making in embodied ai
Jin, E., Hu, J., Huang, Z., Zhang, R., Wu, J., Fei-Fei, L., and Martín-Martín, R · 2023
Cited alongside, same era.
Prd: Peer rank and discussion improve large language model based evaluations
Li, R., Patel, T., and Du, X · 2023
SWE-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Closest in time.
Challenges, evaluation and opportunities for open-world learning
Kejriwal, M., Kildebeck, E., Steininger, R., and Shrivastava, A · 2024
Closest in time.
Openvla: An open-source vision-language-action model
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Closest in time.
Optimus-1: Hybrid multimodal memory empowered agents excel in long-horizon tasks
Li, Z., Xie, Y., Shao, R., Chen, G., Jiang, D., and Nie, L · 2024
Closest in time.
Showui: One vision-language-action model for generalist gui agent
Lin, K. Q., Li, L., Gao, D., Yang, Z., Bai, Z., Lei, W., Wang, L., and Shou, M. Z · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Steve-1: A generative model for text-to-behavior in minecraft
Lifshitz, S., Paster, K., Chan, H., Ba, J., and McIlraith, S · 2023
Cited alongside, same era.
Bedd: The minerl basalt evaluation and demonstrations dataset for training and benchmarking agents that solve fuzzy tasks
Milani, S., Kanervisto, A., Ramanauskas, K., Schulhoff, S., Houghton, B., and Shah, R · 2023
Cited alongside, same era.
Open-world machine learning: applications, challenges, and opportunities
Parmar, J., Chouhan, S., Raychoudhury, V., and Rathore, S · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
Yuan, H., Zhang, C., Wang, H., Xie, F., Cai, P., Dong, H., and Lu, Z · 2023
Cited alongside, same era.
Creative agents: Empowering agents with imagination for creative tasks
Zhang, C., Cai, P., Fu, Y., Yuan, H., and Lu, Z · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2023
Cited alongside, same era.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
Nasiriany, S., Maddukuri, A., Zhang, L., Parikh, A., Lo, A., Joshi, A., Mandlekar, A., and Zhu, Y · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents
Pan, J., Zhang, Y., Tomlin, N., Zhou, Y., Levine, S., and Suhr, A · 2024
Closest in time.
Mr. steve: Instruction-following agents in minecraft with what-where-when memory
Park, J., Cho, J., and Ahn, S · 2024
Closest in time.
mineflayer: Create Minecraft bots with a powerful, stable, and high level JavaScript API
PrismarineJS · 2024
Closest in time.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Qin, Y., Zhou, E., Liu, Q., Yin, Z., Sheng, L., Zhang, R., Qiao, Y., and Shao, J · 2024
Closest in time.
Scaling instructable agents across many simulated worlds
Raad, M. A., Ahuja, A., Barros, C., Besse, F., Bolt, A., Bolton, A., Brownfield, B., Buttimore, G., Cant, M., Chakera, S., et al · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Wang, J., Xu, H., Ye, J., Yan, M., Shen, W., Zhang, J., Huang, F., and Sang, J · 2024
Closest in time.
Holodeck: Language guided generation of 3d embodied ai environments
Yang, Y., Sun, F.-Y., Weihs, L., VanderBilt, E., Herrasti, A., Han, W., Wu, J., Haber, N., Krishna, R., Liu, L., et al · 2024
Closest in time.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al · 2024
Closest in time.
Pre-training goal-based models for sample-efficient reinforcement learning
Yuan, H., Mu, Z., Xie, F., and Lu, Z · 2024
Closest in time.
Agent-as-a-judge: Evaluate agents with agents
Zhuge, M., Zhao, C., Ashley, D., Wang, W., Khizbullin, D., Xiong, Y., Liu, Z., Chang, E., Krishnamoorthi, R., Tian, Y., et al · 2024
Closest in time.
Visual grounding for object-level generalization in reinforcement learning
Jiang, H. and Lu, Z · 2025
Closest in time.
Reinforcement learning friendly vision-language model for minecraft
Jiang, H., Yue, J., Luo, H., Ding, Z., and Lu, Z · 2025
Closest in time.
Li, M., Wang, Z., He, K., Ma, X., and Liang, Y · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Qin, Y., Ye, Y., Fang, J., Wang, H., Liang, S., Tian, S., Zhang, J., Li, J., Li, Y., Huang, S., et al · 2025
Closest in time.
Octopus: Embodied vision-language programmer from environmental feedback
Yang, J., Dong, Y., Liu, S., Li, B., Wang, Z., Tan, H., Jiang, C., Kang, J., Zhang, Y., Zhou, K., et al · 2025
Closest in time.