Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges.
M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. M. Veloso, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. A. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. J. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
Zson: Zero-shot object-goal navigation using multimodal goal embeddings
A. Majumdar, G. Aggarwal, B. Devnani, J. Hoffman, and D. Batra · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Earlier work this paper cites.
Gpt-4 technical report
O. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Open-world multi-task control through goal-aware representation learning and adaptive horizon prediction
S. Cai, Z. Wang, X. Ma, A. Liu, and Y. Liang · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Cited alongside, same era.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Cited alongside, same era.
An embodied generalist agent in 3d world
J. Huang, S. Yong, X. Ma, X. Linghu, P. Li, Y. Wang, Q. Li, S.-C. Zhu, B. Jia, and S. Huang · 2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. B. Girshick · 2023
Cited alongside, same era.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
H. Yuan, C. Zhang, H. Wang, F. Xie, P. Cai, H. Dong, and Z. Lu · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Steve-eye: Equipping llm-based embodied agents with visual perception in open worlds
S. Zheng, Y. Feng, Z. Lu, et al · 2023
Later among the works it cites.
Exploring large language model based intelligent agents: Definitions, methods, and prospects
Y. Cheng, C. Zhang, Z. Zhang, X. Meng, S. Hong, W. Li, Z. Wang, Z. Wang, F. Yin, J. Zhao, et al · 2024
Closest in time.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models, 2024
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Lifshitz, K. Paster, H. Chan, J. Ba, and S. A. McIlraith · 2023
Cited alongside, same era.
Mcu: A task-centric framework for open-ended agent evaluation in minecraft
H. Lin, Z. Wang, J. Ma, and Y. Liang · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Cited alongside, same era.
Interactive language: Talking to robots in real time
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2023
Cited alongside, same era.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Y. Qin, E. Zhou, Q. Liu, Z. Yin, L. Sheng, R. Zhang, Y. Qiao, and J. Shao · 2023
Cited alongside, same era.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. H. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, and K. Hausman · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Groot: Learning to follow instructions by watching gameplay videos
S. Cai, B. Zhang, Z. Wang, X. Ma, A. Liu, and Y. Liang
Cited in the paper.
Closest in time.
RL-GPT: Integrating reinforcement learning and code-as-policy
S. Liu, H. Yuan, M. Hu, Y. Li, Y. Chen, S. Liu, Z. Lu, and J. Jia · 2024
Closest in time.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Closest in time.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Closest in time.
Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches
P. Sundaresan, Q. Vuong, J. Gu, P. Xu, T. Xiao, S. Kirmani, T. Yu, M. Stark, A. Jain, K. Hausman, et al · 2024
Closest in time.
Minedreamer: Learning to follow instructions via chain-of-imagination for simulated-world control
E. Zhou, Y. Qin, Z. Yin, Y. Huang, R. Zhang, L. Sheng, Y. Qiao, and J. Shao · 2024
Closest in time.
Rocket-2: Steering visuomotor policy via cross-view goal alignment
S. Cai, Z. Mu, A. Liu, and Y. Liang · 2025
Closest in time.