Fetching the paper…
Reading the bibliography…
The ability to reflect on and correct failures is crucial for robotic systems to interact stably with real-life objects.Observing the generalization and reasoning capabilities of Multimodal Large Language Models (MLLMs), previous approaches have aimed to utilize these models to enhance robotic systems accordingly.However, these methods typically focus on high-level planning corrections using an additional MLLM, with limited utilization of failed samples to correct low-level contact poses which is particularly prone to occur during articulated object manipulation.To address this gap, we propose an Autonomous Interactive Correction (AIC) MLLM, which makes use of previous low-level interaction experiences to correct SE(3) pose predictions for articulated object.
PartNet: A large-scale benchmark for fine-grained and hierarchical part-level 3D object understanding
K. Mo, S. Zhu, A. X. Chang, L. Yi, S. Tripathi, L. J. Guibas, and H. Su · 2019
Earlier work this paper cites.
SAPIEN: A simulated part-based interactive environment
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su · 2020
Earlier work this paper cites.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Sample efficient grasp learning using equivariant models
X. Zhu, D. Wang, O. Biza, G. Su, R. Walters, and R. Platt · 2022
Earlier work this paper cites.
Edge grasp network: A graph-based se (3)-invariant approach to grasp detection
H. Huang, D. Wang, X. Zhu, R. Walters, and R. Platt · 2022
Earlier work this paper cites.
Devnet: Self-supervised monocular depth learning via density volume construction
K. Zhou, L. Hong, C. Chen, H. Xu, C. Ye, Q. Hu, and Z. Li · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Universal manipulation policy network for articulated objects
Z. Xu, Z. He, and S. Song · 2022
Earlier work this paper cites.
Flowbot3d: Learning 3d articulation flow to manipulate articulated objects
B. Eisner, H. Zhang, and D. Held · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Manydepth2: Motion-aware self-supervised monocular depth estimation in dynamic scenes
K. Zhou, J.-W. Bian, Q. Xie, J.-Q. Zheng, N. Trigoni, and A. Markham · 2023
Earlier work this paper cites.
Reflect: Summarizing robot experiences for failure explanation and correction
Z. Liu, A. Bahety, and S. Song · 2023
Earlier work this paper cites.
Distilling and retrieving generalizable knowledge for robot manipulation via language corrections
L. Zha, Y. Cui, L.-H. Lin, M. Kwon, M. G. Arenas, A. Zeng, F. Xia, and D. Sadigh · 2023
Earlier work this paper cites.
Bootstrap your own skills: Learning to solve new tasks with large language model guidance
J. Zhang, J. Zhang, K. Pertsch, Z. Liu, X. Ren, M. Chang, S.-H. Sun, and J. J. Lim · 2023
Cited alongside, same era.
Programmatically grounded, compositionally generalizable robotic manipulation
R. Wang, J. Mao, J. Hsu, H. Zhao, J. Wu, and Y. Gao · 2023
Cited alongside, same era.
Lidar-llm: Exploring the potential of large language models for 3d lidar understanding
S. Yang, J. Liu, R. Zhang, M. Pan, Z. Guo, X. Li, Z. Chen, P. Gao, Y. Guo, and S. Zhang · 2023
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Partmanip: Learning cross-category generalizable part manipulation policy from point cloud observations
H. Geng, Z. Li, Y. Geng, J. Chen, H. Dong, and H. Wang · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Flowbot++: Learning generalized articulated objects manipulation via articulation projection
H. Zhang, B. Eisner, and D. Held · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Hicrisp: A hierarchical closed-loop robotic intelligent self-correction planner
C. Ming, J. Lin, P. Fong, H. Wang, X. Duan, and J. He · 2023
Cited alongside, same era.
Doremi: Grounding language model by detecting and recovering from plan-execution misalignment
Y. Guo, Y.-J. Wang, L. Zha, Z. Jiang, and J. Chen · 2023
Cited alongside, same era.
M. Skreta, N. Yoshikawa, S. Arellano-Rubach, Z. Ji, L. B. Kristensen, K. Darvish, A. Aspuru-Guzik, F. Shkurti, and A. Garg · 2023
Cited alongside, same era.
Making large multimodal models understand arbitrary visual prompts
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y. Chai, D. Park, and Y. J. Lee · 2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Cited alongside, same era.
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
X. Li, M. Zhang, Y. Geng, H. Geng, Y. Long, Y. Shen, R. Zhang, J. Liu, and H. Dong · 2023
Cited alongside, same era.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Later among the works it cites.
H. Geng, S. Wei, C. Deng, B. Shen, H. Wang, and L. Guibas · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Later among the works it cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Scanet: Correcting lego assembly errors with self-correct assembly network
Y. Wan, K. Zhou, J. Chen, and H. Dong · 2024
Closest in time.
Learning to learn faster from human feedback with language model predictive control
J. Liang, F. Xia, W. Yu, A. Zeng, M. G. Arenas, M. Attarian, M. Bauza, M. Bennice, A. Bewley, A. Dostmohamed, et al · 2024
Closest in time.
Yell at your robot: Improving on-the-fly from language corrections
L. X. Shi, Z. Hu, T. Z. Zhao, A. Sharma, K. Pertsch, J. Luo, S. Levine, and C. Finn · 2024
Closest in time.
Self-corrected multimodal large language model for end-to-end robot manipulation
J. Liu, C. Li, G. Wang, L. Lee, K. Zhou, S. Chen, C. Xiong, J. Ge, R. Zhang, and S. Zhang · 2024
Closest in time.
Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want
W. Lin, X. Wei, R. An, P. Gao, B. Zou, Y. Luo, S. Huang, S. Zhang, and H. Li · 2024
Closest in time.
Scaffolding coordinates to promote vision-language coordination in large multi-modal models
X. Lei, Z. Yang, X. Chen, P. Li, and Y. Liu · 2024
Closest in time.