Fetching the paper…
Reading the bibliography…
Recently, action-based decision-making in open-world environments has gained significant attention.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. Veloso, and R. Salakhutdinov · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, S. Huang, K. Rasul, and Q. Gallouédec · 2020
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Earlier work this paper cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Earlier work this paper cites.
Objectvla: End-to-end open-world object manipulation without demonstration, 2025
M. Zhu, Y. Zhu, J. Li, Z. Zhou, J. Wen, X. Liu, C. Shen, Y. Peng, and F. Feng · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction
S. Cai, Z. Wang, X. Ma, A. Liu, and Y. Liang · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
Mcu: A task-centric framework for open-ended agent evaluation in minecraft
H. Lin, Z. Wang, J. Ma, and Y. Liang · 2023
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models
A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, et al · 2023
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue, 2023
OpenAI · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Cited alongside, same era.
Steve-1: A generative model for text-to-behavior in minecraft
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. McIlraith · 2024
Later among the works it cites.
Selecting large language model to fine-tune via rectified scaling law
H. Lin, B. Huang, H. Ye, Q. Chen, Z. Wang, S. Li, J. Ma, X. Wan, J. Zou, and Y. Liang · 2024
Later among the works it cites.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024
Meta · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Describe, explain, plan and select: interactive planning with large language models enables open-world multi-task agents
Z. Wang, S. Cai, G. Chen, A. Liu, X. Ma, Y. Liang, and T. CraftJarvis · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2023
Cited alongside, same era.
Foundation models for decision making: Problems, methods, and opportunities
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans · 2023
Cited alongside, same era.
Steve-eye: Equipping llm-based embodied agents with visual perception in open worlds
S. Zheng, J. Liu, Y. Feng, and Z. Lu · 2023
Cited alongside, same era.
Introducing the next generation of claude, 2024
Anthropic · 2024
Cited alongside, same era.
Edgevla: Efficient vision-language-action models
P. Budzianowski, W. Maa, M. Freed, J. Mo, A. Xie, V. Tipnis, and B. Bolte · 2024
Cited alongside, same era.
Exploring large language model based intelligent agents: Definitions, methods, and prospects
Y. Cheng, C. Zhang, Z. Zhang, X. Meng, S. Hong, W. Li, Z. Wang, Z. Wang, F. Yin, J. Zhao, et al · 2024
Cited alongside, same era.
Later among the works it cites.
Hirt: Enhancing robotic control with hierarchical robot transformers
J. Zhang, Y. Guo, X. Chen, Y.-J. Wang, Y. Hu, C. Shi, and J. Chen · 2024
Later among the works it cites.
Optimizing latent goal by learning from trajectory preference
G. Zhao, K. Lian, H. Lin, H. Fu, Q. Fu, S. Cai, Z. Wang, and Y. Liang · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan · 2024
Later among the works it cites.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
R. Zheng, Y. Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang · 2024
Later among the works it cites.
Minedreamer: Learning to follow instructions via chain-of-imagination for simulated-world control
E. Zhou, Y. Qin, Z. Yin, Y. Huang, R. Zhang, L. Sheng, Y. Qiao, and J. Shao · 2024
Later among the works it cites.
P. Chen, P. Bu, Y. Wang, X. Wang, Z. Wang, J. Guo, Y. Zhao, Q. Zhu, J. Song, S. Yang, et al · 2025
Closest in time.
Open-world skill discovery from unsegmented demonstrations
J. Deng, Z. Wang, S. Cai, A. Liu, and Y. Liang · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.
Up-vla: A unified understanding and prediction model for embodied agent
J. Zhang, Y. Guo, Y. Hu, X. Chen, X. Zhu, and J. Chen · 2025
Closest in time.
Dexgraspvla: A vision-language-action framework towards general dexterous grasping, 2025
Y. Zhong, X. Huang, R. Li, C. Zhang, Y. Liang, Y. Yang, and Y. Chen · 2025
Closest in time.
Chatvla: Unified multimodal understanding and robot control with vision-language-action model
Z. Zhou, Y. Zhu, M. Zhu, J. Wen, N. Liu, Z. Xu, W. Meng, R. Cheng, Y. Peng, C. Shen, et al · 2025
Closest in time.