Fetching the paper…
Reading the bibliography…
Generalist robots that can perform a range of different tasks in open-world settings must be able to not only reason about the steps needed to accomplish their goals, but also process complex instructions, prompts, and even feedback during task execution.
A computational model for the alignment of hierarchical scene representations in human-robot interaction
Swadzba, A., Vorwerg, C., Wachsmuth, S., and Rickheit, G · 2009
Earlier work this paper cites.
Thinking, fast and slow
Kahneman, D · 2011
Earlier work this paper cites.
Learning to parse natural language commands to a robot control system
Matuszek, C., Herbst, E., Zettlemoyer, L., and Fox, D · 2013
Earlier work this paper cites.
Decoupled weight decay regularization, 2017
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Inferring compact representations for efficient natural language understanding of robot instructions
Patki, S., Daniele, A. F., Walter, M. R., and Howard, T. M · 2019
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
Stepputtis, S., Campbell, J., Phielipp, M., Lee, S., Baral, C., and Ben Amor, H · 2020
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Earlier work this paper cites.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning, 2023
Hu, Y., Lin, F., Zhang, T., Yi, L., and Gao, Y · 2023
Earlier work this paper cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Fei-Fei, L · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Earlier work this paper cites.
Interactive robot learning from verbal correction
Liu, H., Chen, A., Zhu, Y., Swaminathan, A., Kolobov, A., and Cheng, C.-A · 2023
Earlier work this paper cites.
Is feedback all you need? leveraging natural language feedback in goal-conditioned rl
McCallum, S., Taylor-Davies, M., Albrecht, S., and Suglia, A · 2023
Cited alongside, same era.
Learning neuro-symbolic programs for language guided robot manipulation
Namasivayam, K., Singh, H., Bindal, V., Tuli, A., Agrawal, V., Jain, R., Singla, P., and Paul, R · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Cited alongside, same era.
Open-world object manipulation using pre-trained vision-language models
Stone, A., Xiao, T., Lu, Y., Gopalakrishnan, K., Lee, K.-H., Vuong, Q., Wohlhart, P., Kirmani, S., Zitkovich, B., Xia, F., et al · 2023
Cited alongside, same era.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
Nasiriany, S., Xia, F., Yu, W., Xiao, T., Liang, J., Dasgupta, I., Xie, A., Driess, D., Wahid, A., Xu, Z., et al · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Xu, C., Luo, J., Kreiman, T., Tan, Y., Chen, L. Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S · 2024
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., Jain, A., et al · 2024
Later among the works it cites.
Open-vocabulary mobile manipulation in unseen dynamic environments with 3d semantic maps
Qiu, D., Ma, W., Pan, Z., Xiong, H., and Liang, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Cited alongside, same era.
Rt-h: Action hierarchies using language
Belkhale, S., Ding, T., Xiao, T., Sermanet, P., Vuong, Q., Tompson, J., Chebotar, Y., Dwibedi, D., and Sadigh, D · 2024
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer
Beyer, L., Steiner, A., Pinto, A. S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., Bugliarello, E., et al · 2024
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Cited alongside, same era.
Automating robot failure recovery using vision-language models with optimized prompts
Chen, H., Yao, Y., Liu, R., Liu, C., and Ichnowski, J · 2024
Cited alongside, same era.
Racer: Rich language-guided failure recovery policies for imitation learning
Dai, Y., Lee, J., Fazeli, N., and Chai, J · 2024
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Fu, Z., Zhao, T. Z., and Finn, C · 2024
Cited alongside, same era.
Shah, R., Yu, A., Zhu, Y., Zhu, Y., and Martín-Martín, R · 2024
Later among the works it cites.
Yell at your robot: Improving on-the-fly from language corrections
Shi, L. X., Hu, Z., Zhao, T. Z., Sharma, A., Pertsch, K., Luo, J., Levine, S., and Finn, C · 2024
Later among the works it cites.
Lgr2: Language guided reward relabeling for accelerating hierarchical reinforcement learning
Singh, U., Bhattacharyya, P., and Namboodiri, V. P · 2024
Later among the works it cites.
Rlvf: Learning from verbal feedback without overgeneralization
Stephan, M., Khazatsky, A., Mitchell, E., Chen, A. S., Hsu, S., Sharma, A., and Finn, C · 2024
Later among the works it cites.
Llmˆ 3: Large language model-based task and motion planning with motion failure reasoning
Wang, S., Han, M., Jiao, Z., Zhang, Z., Wu, Y. N., Zhu, S.-C., and Liu, H · 2024
Later among the works it cites.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Wen, J., Zhu, Y., Li, J., Zhu, M., Wu, K., Xu, Z., Liu, N., Cheng, R., Shen, C., Peng, Y., et al · 2024
Later among the works it cites.
Robi butler: Remote multimodal interactions with household robot assistant
Xiao, A., Janaka, N., Hu, T., Gupta, A., Li, K., Yu, C., and Hsu, D · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S · 2024
Later among the works it cites.
Closed-loop open-vocabulary mobile manipulation with gpt-4v
Zhi, P., Zhang, Z., Han, M., Zhang, Z., Li, Z., Jiao, Z., Jia, B., and Huang, S · 2024
Later among the works it cites.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., and Levine, S · 2025
Closest in time.
Universal actions for enhanced embodied foundation models
Zheng, J., Li, J., Liu, D., Zheng, Y., Wang, Z., Ou, Z., Liu, Y., Liu, J., Zhang, Y.-Q., and Zhan, X · 2025
Closest in time.