Fetching the paper…
Reading the bibliography…
Robot chain-of-thought reasoning (CoT) -- wherein a model predicts helpful intermediate representations before choosing actions -- provides an effective method for improving the generalization and performance of robot policies, especially vision-language-action models (VLAs).
Don’t stop pretraining: Adapt language models to domains and tasks, 2020
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith · 2004
Earlier work this paper cites.
Neural discrete representation learning, 2017
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets, 2021
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2021
Earlier work this paper cites.
Latent plans for task agnostic offline reinforcement learning
E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard · 2022
Earlier work this paper cites.
Affordance learning from play for sample-efficient policy learning
J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard · 2022
Earlier work this paper cites.
Skill induction and planning with latent language, 2022
P. Sharma, A. Torralba, and J. Andreas · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models, 2022
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter · 2022
Earlier work this paper cites.
Co-training improves prompt-based learning for large language models, 2022
H. Lang, M. Agrawal, Y. Kim, and D. Sontag · 2022
Earlier work this paper cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022
A. Kumar, A. Raghunathan, R. Jones, T. Ma, and P. Liang · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale, 2023
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V. Myers, K. Fang, C. Finn, and S. Levine · 2023
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Earlier work this paper cites.
Large language models are zero-shot reasoners, 2023
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2023
Earlier work this paper cites.
Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Earlier work this paper cites.
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto · 2023
Earlier work this paper cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Earlier work this paper cites.
Grounding language with visual affordances over unstructured data
O. Mees, J. Borja-Diaz, and W. Burgard · 2023
Earlier work this paper cites.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches, 2023
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, P. Sundaresan, P. Xu, H. Su, K. Hausman, C. Finn, Q. Vuong, and T. Xiao · 2023
Earlier work this paper cites.
Language-driven representation learning for robotics, 2023
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model, 2023
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Earlier work this paper cites.
What makes pre-trained visual representations successful for robust manipulation?, 2023
K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman · 2023
Earlier work this paper cites.
How to prepare your task head for finetuning, 2023
Y. Ren, S. Guo, W. Bae, and D. J. Sutherland · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training, 2023
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware, 2023
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers, 2023
W. Merrill and A. Sabharwal · 2023
Cited alongside, same era.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Cited alongside, same era.
Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation
R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Cited alongside, same era.
Bridging language and action: A survey of language-based robot manipulation
H. Zhou, X. Yao, O. Mees, Y. Meng, T. Xiao, Y. Bisk, J. Oh, E. Johns, M. Shridhar, D. Shah, J. Thomason, K. Huang, J. Chai, Z. Bing, and A. Knoll · 2024
Later among the works it cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?, 2024
A. Majumdar, K. Yadav, S. Arnaud, Y. J. Ma, C. Chen, S. Silwal, A. Jain, V.-P. Berges, P. Abbeel, J. Malik, D. Batra, Y. Lin, O. Maksymets, A. Rajeswaran, and F. Meier · 2024
Later among the works it cites.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities, 2024
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia · 2024
Later among the works it cites.
What do learning dynamics reveal about generalization in llm reasoning?, 2024
K. Kang, A. Setlur, D. Ghosh, J. Steinhardt, C. Tomlin, S. Levine, and A. Kumar · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
3d-vla: A 3d vision-language-action generative world model
H. Zhen, X. Qiu, P. Chen, J. Yang, X. Yan, Y. Du, Y. Hong, and C. Gan · 2024
Cited alongside, same era.
Robotic control via embodied chain-of-thought reasoning
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2024
Cited alongside, same era.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
R. Zheng, Y. Liang, S. Huang, J. Gao, H. Daumé III, A. Kolobov, F. Huang, and J. Yang · 2024
Cited alongside, same era.
Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer, 2024
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai · 2024
Cited alongside, same era.
Paligemma 2: A family of versatile vlms for transfer, 2024
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y. Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, S. Qin, R. Ingle, E. Bugliarello, S. Kazemzadeh, T. Mesnard, I. Alabdulmohsin, L. Beyer, and X. Zhai · 2024
Cited alongside, same era.
Let’s think dot by dot: Hidden computation in transformer language models, 2024
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Later among the works it cites.
Think before you speak: Training language models with pause tokens, 2024
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Later among the works it cites.
M. Brief, O. Ovadia, G. Shenderovitz, N. B. Yoash, R. Lemberg, and E. Sheetrit · 2024
Later among the works it cites.
Quest: Self-supervised skill abstractions for learning continuous control, 2024
A. Mete, H. Xue, A. Wilcox, Y. Chen, and A. Garg · 2024
Later among the works it cites.
Tensorrt-llm
NVIDIA · 2024
Later among the works it cites.
Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models, 2024
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y. Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y. Chou, A. Chheda, J. Sparks, S. Skjonsberg, M. Schmitz, A. Sarnat, B. Bischoff, P. Walsh, C. Newell, P. Wolters, T. Gupta, K.-H. Zeng, J. Borchardt, D. Groeneveld, C. Nam, S. Lebrecht, C. Wittlif, C. Schoenick, O. Michel, R. Krishna, L. Weihs, N. A. Smith, H. Hajishirzi, R. Girshick, A. Farhadi, and A. Kembhavi · 2024
Later among the works it cites.
Dinov2: Learning robust visual features without supervision, 2024
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Later among the works it cites.
Qwen2.5 technical report, 2024
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Beyond sight: Finetuning generalist robot policies with heterogeneous sensors via language grounding
J. Jones, O. Mees, C. Sferrazza, K. Stachowicz, P. Abbeel, and S. Levine · 2025
Closest in time.
A0: An affordance-aware hierarchical model for general robotic manipulation
R. Xu, J. Zhang, M. Guo, Y. Wen, H. Yang, M. Lin, J. Huang, Z. Li, K. Zhang, L. Wang, et al · 2025
Closest in time.
Otter: A vision-language-action model with text-aware visual feature extraction
H. Huang, F. Liu, L. Fu, T. Wu, M. Mukadam, J. Malik, K. Goldberg, and P. Abbeel · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models, 2025
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Dexvla: Vision-language model with plug-in diffusion expert for general robot control
J. Wen, Y. Zhu, J. Li, Z. Tang, C. Shen, and F. Feng · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y. Liu, D. Xiang, G. Wetzstein, and T.-Y. Lin · 2025
Closest in time.
Action-free reasoning for policy generalization, 2025
J. Clark, S. Mirchandani, D. Sadigh, and S. Belkhale · 2025
Closest in time.
Vision-language models provide promptable representations for reinforcement learning
W. Chen, O. Mees, A. Kumar, and S. Levine · 2025
Closest in time.
A taxonomy for evaluating generalist robot policies
J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh · 2025
Closest in time.
A taxonomy for evaluating generalist robot policies, 2025
J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh · 2025
Closest in time.
Tensorrt-openvla, 2025
W. Chen, M. Zawalski, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success, 2025
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.