Fetching the paper…
Reading the bibliography…
Humans act with context and intention, with reasoning playing a central role.
Dynamic time warping algorithm review
P. Senin · 2008
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Embodied intelligence via learning and evolution
A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei · 2021
Earlier work this paper cites.
Creating multimodal interactive agents with imitation and self-supervised learning
D. I. A. Team, J. Abramson, A. Ahuja, A. Brussee, F. Carnevale, M. Cassin, F. Fischer, P. Georgiev, A. Goldin, M. Gupta, et al · 2021
Earlier work this paper cites.
Ocid-ref: A 3d robotic dataset with embodied language for clutter scene grounding
K.-J. Wang, Y.-H. Liu, H.-T. Su, J.-W. Wang, Y.-S. Wang, W. H. Hsu, and W.-C. Chen · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Egoplan-bench: Benchmarking multimodal large language models for human-level planning
Y. Chen, Y. Ge, Y. Ge, M. Ding, B. Li, R. Wang, R. Xu, Y. Shan, and X. Liu · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, et al · 2023
Earlier work this paper cites.
What’s" up" with vision-language models? investigating their struggle with spatial reasoning
A. Kamath, J. Hessel, and K.-W. Chang · 2023
Earlier work this paper cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
Libero: Benchmarking knowledge transfer for lifelong robot learning
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2023
Earlier work this paper cites.
Scaling data-constrained language models
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel · 2023
Earlier work this paper cites.
Paco: Parts and attributes of common objects
V. Ramanathan, A. Kalia, V. Petrovic, Y. Wen, B. Zheng, B. Guo, R. Wang, A. Marquez, R. Kovvuri, A. Kadian, et al · 2023
Earlier work this paper cites.
Waypoint-based imitation learning for robotic manipulation
L. X. Shi, A. Sharma, T. Z. Zhao, and C. Finn · 2023
Earlier work this paper cites.
Reinforcement learning with foundation priors: Let the embodied agent efficiently learn on its own
W. Ye, Y. Zhang, H. Weng, X. Gu, S. Wang, T. Zhang, M. Wang, P. Abbeel, and Y. Gao · 2023
Earlier work this paper cites.
Multimodal chain-of-thought reasoning in language models
Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al · 2023
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
EmbSpatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models
M. Du, B. Wu, Z. Li, X. Huang, and Z. Wei · 2024
Cited alongside, same era.
Blink: Multimodal large language models can see but not perceive
X. Fu, Y. Hu, B. Li, Y. Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna · 2024
Cited alongside, same era.
Learning with 3d rotations, a hitchhiker’s guide to so (3)
A. R. Geist, J. Frey, M. Zhobro, A. Levina, and G. Martius · 2024
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, et al · 2025
Closest in time.
C. Cheang, S. Chen, Z. Cui, Y. Hu, L. Huang, T. Kong, H. Li, Y. Li, Y. Liu, X. Ma, et al · 2025
Closest in time.
Co-evolving embodied intelligence with design for artificial intelligence architecture
Y. Dai and J. Wang · 2025
Closest in time.
Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al · 2025
Closest in time.
Knowledge insulating vision-language-action models: Train fast, run fast, generalize better
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Minicpm: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Cited alongside, same era.
Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y. Deng, S. Xu, Y. Zhang, et al · 2024
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al · 2024
Cited alongside, same era.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Cited alongside, same era.
Sat: Dynamic spatial aptitude training for multimodal language models
A. Ray, J. Duan, E. Brown, R. Tan, D. Bashkirova, R. Hendrix, K. Ehsani, A. Kembhavi, B. A. Plummer, R. Krishna, et al · 2024
Cited alongside, same era.
Dino-x: A unified vision model for open-world object detection and understanding
T. Ren, Y. Chen, Q. Jiang, Z. Zeng, Y. Xiong, W. Liu, Z. Ma, J. Shen, Y. Gao, X. Jiang, et al · 2024
Cited alongside, same era.
D. Driess, J. T. Springenberg, B. Ichter, L. Yu, A. Li-Bell, K. Pertsch, A. Z. Ren, H. Walke, Q. Vuong, L. X. Shi, et al · 2025
Closest in time.
Robix: A unified model for robot interaction, reasoning and planning
H. Fang, M. Zhang, H. Dong, W. Li, Z. Wang, Q. Zhang, X. Tian, Y. Hu, and H. Li · 2025
Closest in time.
Towards human-level intelligence via human-like whole-body manipulation
G. Gao, J. Wang, J. Zuo, J. Jiang, J. Zhang, X. Zeng, Y. Zhu, L. Ma, K. Chen, M. Sheng, et al · 2025
Closest in time.
Thinkact: Vision-language-action reasoning via reinforced visual latent planning
C.-P. Huang, Y.-H. Wu, M.-H. Chen, Y.-C. F. Wang, and F.-E. Yang · 2025
Closest in time.
π _ 0.5 \pi\_0.5 : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Robobrain: A unified brain model for robotic manipulation from abstract to concrete
Y. Ji, H. Tan, J. Shi, X. Hao, Y. Zhang, H. Zhang, P. Wang, M. Zhao, Y. Mu, P. An, et al · 2025
Closest in time.
Molmoact: Action reasoning models that can reason in space
J. Lee, J. Duan, H. Fang, Y. Deng, S. Liu, B. Li, B. Fang, J. Zhang, Y. R. Wang, S. Lee, et al · 2025
Closest in time.
Llava-c: Continual improved visual instruction tuning
W. Liu, F. Zhu, H. Guo, L. Wei, and C.-L. Liu · 2025
Closest in time.
Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning
J. Luo, C. Xu, J. Wu, and S. Levine · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
RoboSpatial: Teaching spatial understanding to 2D and 3D vision-language models for robotics
C. H. Song, V. Blukis, J. Tremblay, S. Tyree, Y. Su, and S. Birchfield · 2025
Closest in time.
Dspv2: Improved dense policy for effective and generalizable whole-body mobile manipulation
Y. Su, C. Zhang, S. Chen, L. Tan, Y. Tang, J. Wang, and X. Liu · 2025
Closest in time.
Spacevista: All-scale visual spatial reasoning from mm to km
P. Sun, S. Lang, D. Wu, Y. Ding, K. Feng, H. Liu, Z. Ye, R. Liu, Y.-H. Liu, J. Wang, and X. Yue · 2025
Closest in time.
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, et al · 2025
Closest in time.
Z. Yuan, T. Wei, L. Gu, P. Hua, T. Liang, Y. Chen, and H. Xu · 2025
Closest in time.
A vision-language-action-critic model for robotic real-world reinforcement learning
S. Zhai, Q. Zhang, T. Zhang, F. Huang, H. Zhang, M. Zhou, S. Zhang, L. Liu, S. Lin, and J. Pang · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, et al · 2025
Closest in time.
Roborefer: Towards spatial referring with reasoning in vision-language models for robotics
E. Zhou, J. An, C. Chi, Y. Han, S. Rong, C. Zhang, P. Wang, Z. Wang, T. Huang, L. Sheng, et al · 2025
Closest in time.