Fetching the paper…
Reading the bibliography…
The ability to grasp objects in-the-wild from open-ended language instructions constitutes a fundamental challenge in robotics.
Design and use paradigms for gazebo, an open-source multi-robot simulator
N. P. Koenig and A. Howard · 2004
Earlier work this paper cites.
Efficient grasping from rgbd images: Learning using a new rectangle representation
Y. Jiang, S. Moseson, and A. Saxena · 2011
Earlier work this paper cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Earlier work this paper cites.
Jacquard: A large scale dataset for robotic grasp detection
A. Depierre, E. Dellandréa, and L. Chen · 2018
Earlier work this paper cites.
Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach
D. Morrison, P. Corke, and J. Leitner · 2018
Earlier work this paper cites.
How can we know what language models know?
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig · 2019
Earlier work this paper cites.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, P. Lewis, A. Bakhtin, Y. Wu, A. H. Miller, and S. Riedel · 2019
Earlier work this paper cites.
Antipodal robotic grasping using generative residual convolutional neural network
S. Kumra, S. Joshi, and F. Sahin · 2019
Earlier work this paper cites.
Easylabel: A semi-automatic pixel-wise object annotation tool for creating robotic rgb-d datasets
M. Suchi, T. Patten, and M. Vincze · 2019
Earlier work this paper cites.
6-dof graspnet: Variational grasp generation for object manipulation
A. Mousavian, C. Eppner, and D. Fox · 2019
Earlier work this paper cites.
6-dof grasping for target-driven object manipulation in clutter
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox · 2019
Earlier work this paper cites.
Learning grasp affordance reasoning through semantic relations
P. Ardón, É. Pairet, R. P. A. Petrick, S. Ramamoorthy, and K. S. Lohan · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, and et. al · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
H. Fang, C. Wang, M. Gou, and C. Lu · 2020
Earlier work this paper cites.
ACRONYM: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2020
Earlier work this paper cites.
Same object, different grasps: Data and semantic knowledge for task-oriented grasping
A. Murali, W. Liu, K. Marino, S. Chernova, and A. K. Gupta · 2020
Earlier work this paper cites.
Unseen object instance segmentation for robotic environments
C. Xie, Y. Xiang, A. Mousavian, and D. Fox · 2020
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Earlier work this paper cites.
Mdetr - modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Earlier work this paper cites.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Y. Yao, A. Zhang, Z. Zhang, Z. Liu, T. seng Chua, and M. Sun · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
End-to-end trainable deep neural network for robotic grasp detection and semantic segmentation from rgb
S. Ainetter and F. Fraundorfer · 2021
Earlier work this paper cites.
Volumetric grasping network: Real-time 6 dof grasp detection in clutter
M. Breyer, J. J. Chung, L. Ott, R. Y. Siegwart, and J. I. Nieto · 2021
Earlier work this paper cites.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox · 2021
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, and A. R. et. al · 2022
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
C. H. Song, J. Wu, C. Washington, B. M. Sadler, W.-L. Chao, and Y. Su · 2022
Cited alongside, same era.
Robot task planning and situation handling in open worlds
Y. Ding, X. Zhang, S. Amiri, N. Cao, H. Yang, C. Esselink, and S. Zhang · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, and C. F. et. al · 2022
Cited alongside, same era.
Instruct2act: Mapping multi-modality instructions to robotic actions with large language model
S. Huang, Z. Jiang, H.-W. Dong, Y. J. Qiao, P. Gao, and H. Li · 2023
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. F. C. Bucker, and A. Kapoor · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Later among the works it cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. yue Li, and J. Gao · 2023
Later among the works it cites.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. R. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, N. Brown, T. Jackson, L. Luu, S. Levine, K. Hausman, and B. Ichter · 2022
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Cited alongside, same era.
Adapt: Vision-language navigation with modality-aligned action prompts
B. Lin, Y. Zhu, Z. Chen, X. Liang, J. zhuo Liu, and X. Liang · 2022
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. R. Florence, and A. Zeng · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. S. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. R. Florence · 2022
Cited alongside, same era.
Simple open-vocabulary object detection with vision transformers
M. Minderer, A. A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, X. Wang, X. Zhai, T. Kipf, and N. Houlsby · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Cited alongside, same era.
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao · 2023
Later among the works it cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. A. Li, P. Fung, and S. C. H. Hoi · 2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. B. Girshick · 2023
Later among the works it cites.
Language-guided robot grasping: Clip-based referring grasp synthesis in clutter
G. Tziafas, Y. XU, A. Goel, M. Kasaei, Z. Li, and H. Kasaei · 2023
Later among the works it cites.
What does clip know about a red circle? visual prompt engineering for vlms
A. Shtedritski, C. Rupprecht, and A. Vedaldi · 2023
Later among the works it cites.
L. Yang, Y. Wang, X. Li, X. Wang, and J. Yang · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, K. Choromanski, T. Ding, D. Driess, C. Finn, P. R. Florence, and C. F. et. al · 2023
Later among the works it cites.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P. Luo · 2023
Later among the works it cites.
Robotgpt: Robot manipulation learning from chatgpt
Y. Jin, D. Li, Y. A, J. Shi, P. Hao, F. Sun, J. Zhang, and B. Fang · 2023
Later among the works it cites.
Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration
N. Wake, A. Kanehira, K. Sasabuchi, J. Takamatsu, and K. Ikeuchi · 2023
Later among the works it cites.
Instance-wise grasp synthesis for robotic grasping
Y. Xu, M. M. Kasaei, S. H. M. Kasaei, and Z. Li · 2023
Later among the works it cites.
Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes
Y. Lu, Y. Fan, B. Deng, F. Liu, Y. Li, and S. Wang · 2023
Later among the works it cites.
Graspgpt: Leveraging semantic knowledge from a large language model for task-oriented grasping
C. Tang, D. Huang, W. Ge, W. Liu, and H. Zhang · 2023
Later among the works it cites.
Segment everything everywhere all at once
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Gao, and Y. J. Lee · 2023
Later among the works it cites.
Semantic-sam: Segment and recognize anything at any granularity
F. Li, H. Zhang, P. Sun, X. Zou, S. Liu, J. Yang, C. Li, L. Zhang, and J. Gao · 2023
Later among the works it cites.
Moka: Open-world robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Closest in time.
Language-driven grasp detection
V. D. An, M. N. Vu, B. Huang, N. Nguyen, H. Le, T. D. Vo, and A. Nguyen · 2024
Closest in time.