Fetching the paper…
Reading the bibliography…
Reasoning is central to purposeful action, yet most robotic foundation models map perception and instructions directly to control, which limits adaptability, generalization, and semantic grounding.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
FigureQA: An annotated figure dataset for visual reasoning
S. E. Kahou, V. Michalski, A. Atkinson, Á. Kádár, A. Trischler, and Y. Bengio · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
DVQA: Understanding data visualizations via question answering
K. Kafle, B. Price, S. Cohen, and C. Kanan · 2018
Earlier work this paper cites.
Roboturk: A crowdsourcing platform for robotic skill learning through imitation
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, et al · 2018
Earlier work this paper cites.
TallyQA: Answering complex counting questions
M. Acharya, K. Kafle, and C. Kanan · 2019
Earlier work this paper cites.
Robot learning of shifting objects for grasping in cluttered environments
L. Berscheid, P. Meißner, and T. Kröger · 2019
Earlier work this paper cites.
Scene text visual question answering
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
PlotQA: Reasoning over scientific plots
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar · 2020
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Earlier work this paper cites.
DocVQA: A dataset for VQA on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
A survey of embodied ai: From simulators to research tasks
J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2022
Earlier work this paper cites.
Large language models can self-improve
J. Huang, S. S. Gu, L. Hou, Y. Wu, X. Wang, H. Yu, and J. Han · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
InfographicVQA
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
A-OKVQA: A benchmark for visual question answering using world knowledge
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi · 2022
Earlier work this paper cites.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot
H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu · 2023
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al · 2024
Later among the works it cites.
The colosseum: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Later among the works it cites.
From llms to actions: Latent codes as bridges in hierarchical robot control
Y. Shentu, P. Wu, A. Rajeswaran, and P. Abbeel · 2024
Later among the works it cites.
Yell at your robot: Improving on-the-fly from language corrections
L. X. Shi, Z. Hu, T. Z. Zhao, A. Sharma, K. Pertsch, J. Luo, S. Levine, and C. Finn · 2024
Later among the works it cites.
Dolma: An open corpus of three trillion tokens for language model pretraining research
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches, 2023
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, P. Sundaresan, P. Xu, H. Su, K. Hausman, C. Finn, Q. Vuong, and T. Xiao · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
P. Lu, L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan · 2023
Cited alongside, same era.
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Cited alongside, same era.
Rail-only: A low-cost high-performance network for training llms with trillion parameters
W. Wang, M. Ghobadi, K. Shakeri, Y. Zhang, and N. Hasani · 2023
Cited alongside, same era.
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, et al · 2024
Later among the works it cites.
Q. Sun, P. Hong, T. D. Pala, V. Toh, U. Tan, D. Ghosal, S. Poria, et al · 2024
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
A. Xie, L. Lee, T. Xiao, and C. Finn · 2024
Later among the works it cites.
Robopoint: A vision-language model for spatial affordance prediction for robotics
W. Yuan, J. Duan, V. Blukis, W. Pumacay, R. Krishna, A. Murali, A. Mousavian, and D. Fox · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2024
Later among the works it cites.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
R. Zheng, Y. Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang · 2024
Later among the works it cites.
Perception tokens enhance visual reasoning in multimodal language models
M. Bigverdi, Z. Luo, C.-Y. Hsieh, E. Shen, D. Chen, L. G. Shapiro, and R. Krishna · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Worldvla: Towards autoregressive action world model
J. Cen, C. Yu, H. Yuan, Y. Jiang, S. Huang, J. Guo, X. Li, Y. Song, H. Luo, F. Wang, et al · 2025
Closest in time.
Action-free reasoning for policy generalization
J. Clark, S. Mirchandani, D. Sadigh, and S. Belkhale · 2025
Closest in time.
H. Fang, M. Grotz, W. Pumacay, Y. R. Wang, D. Fox, R. Krishna, and J. Duan · 2025
Closest in time.
Foundation models in robotics: Applications, challenges, and the future
R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y. Zhu, S. Song, A. Kapoor, K. Hausman, et al · 2025
Closest in time.
Thinkact: Vision-language-action reasoning via reinforced visual latent planning
C.-P. Huang, Y.-H. Wu, M.-H. Chen, Y.-C. F. Wang, and F.-E. Yang · 2025
Closest in time.
Nora: A small open-sourced generalist vision language action model for embodied tasks
C.-Y. Hung, Q. Sun, P. Hong, A. Zadeh, C. Li, U. Tan, N. Majumder, S. Poria, et al · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization, 2025
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky · 2025
Closest in time.
Hamster: Hierarchical action models for open-world robot manipulation
Y. Li, Y. Deng, J. Zhang, J. Jang, M. Memmel, R. Yu, C. R. Garrett, F. Ramos, D. Fox, A. Li, et al · 2025
Closest in time.
Towards generalist robot policies: What matters in building vision-language-action models
H. Liu, X. Li, P. Li, M. Liu, D. Wang, J. Liu, B. Kang, X. Ma, T. Kong, and H. Zhang · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots, 2025
NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, et al · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, et al · 2025
Closest in time.
Your body thinks as much as your mind
B. Tversky · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, et al · 2025
Closest in time.