Fetching the paper…
Reading the bibliography…
Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability.
Dual processes in reasoning?
P. C. Wason and J. S. B. Evans · 1974
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
M. MacMahon, B. Stankiewicz, and B. Kuipers · 2006
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
H. Mei, M. Bansal, and M. Walter · 2016
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
D. Misra, J. Langford, and Y. Artzi · 2017
Earlier work this paper cites.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. Singh, H. Lee, and P. Kohli · 2017
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
J. Andreas, D. Klein, and S. Levine · 2017
Earlier work this paper cites.
A survey of reinforcement learning informed by natural language
J. Luketina, N. Nardelli, G. Farquhar, J. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel · 2019
Earlier work this paper cites.
Language as an abstraction for hierarchical deep reinforcement learning
Y. Jiang, S. S. Gu, K. P. Murphy, and C. Finn · 2019
Earlier work this paper cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
J. Lee, Y. Lee, J. Kim, A. Kosiorek, S. Choi, and Y. W. Teh · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Robots that use language
S. Tellex, N. Gopalan, H. Kress-Gazit, and C. Matuszek · 2020
Earlier work this paper cites.
Jointly improving parsing and perception for natural language commands through human-robot dialog
J. Thomason, A. Padmakumar, J. Sinapov, N. Walker, Y. Jiang, H. Yedidsion, J. Hart, P. Stone, and R. Mooney · 2020
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2020
Earlier work this paper cites.
Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards
P. Goyal, S. Niekum, and R. Mooney · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, et al · 2022
Earlier work this paper cites.
Flingbot: The unreasonable effectiveness of dynamic manipulation for cloth unfolding
H. Ha and S. Song · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Later among the works it cites.
Vint: A foundation model for visual navigation
D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine · 2023
Later among the works it cites.
Toward general-purpose robots via foundation models: A survey and meta-analysis
Y. Hu, Q. Xie, V. Jain, J. Francis, J. Patrikar, N. Keetha, S. Kim, Y. Xie, T. Zhang, Z. Zhao, et al · 2023
Later among the works it cites.
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, et al · 2023
Later among the works it cites.
Agile catching with whole-body mpc and blackbox policy learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, M. Attarian, B. Ichter, K. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, et al · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Cited alongside, same era.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, S. Kirmani, B. Zitkovich, F. Xia, et al · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large scale language model society
G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem · 2023
Cited alongside, same era.
Liv: Language-image representations and rewards for robotic control
Y. J. Ma, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Cited alongside, same era.
S. Abeyruwan, A. Bewley, N. M. Boffi, K. M. Choromanski, D. B. D’Ambrosio, D. Jain, P. R. Sanketi, A. Shankar, V. Sindhwani, S. Singh, et al · 2023
Later among the works it cites.
Doremi: Grounding language model by detecting and recovering from plan-execution misalignment
Y. Guo, Y.-J. Wang, L. Zha, Z. Jiang, and J. Chen · 2023
Later among the works it cites.
Text2motion: From natural language instructions to feasible plans
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
Gensim: Generating robotic simulation tasks via large language models
L. Wang, Y. Ling, Z. Yuan, M. Shridhar, C. Bao, Y. Qin, B. Wang, H. Xu, and X. Wang · 2023
Later among the works it cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable task planning
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. Reid, and N. Suenderhauf · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Later among the works it cites.
Saytap: Language to quadrupedal locomotion
Y. Tang, W. Yu, J. Tan, H. Zen, A. Faust, and T. Harada · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi · 2024
Closest in time.
Mrest: Multi-resolution sensing for real-time control with vision-language models
S. Saxena, M. Sharma, and O. Kroemer · 2024
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P. Luo · 2024
Closest in time.
Rt-h: Action hierarchies using language
S. Belkhale, T. Ding, T. Xiao, P. Sermanet, Q. Vuong, J. Tompson, Y. Chebotar, D. Dwibedi, and D. Sadigh · 2024
Closest in time.