Fetching the paper…
Reading the bibliography…
Despite significant progress in robotic systems for operation within human-centric environments, existing models still heavily rely on explicit human commands to identify and manipulate specific objects.
Graspit! a versatile simulator for robotic grasping
A. T. Miller and P. K. Allen · 2004
Earlier work this paper cites.
Cloth grasp point detection based on multiple-view geometric cues with application to robotic towel folding
J. Maitin-Shepard, M. Cusumano-Towner, J. Lei, and P. Abbeel · 2010
Earlier work this paper cites.
Efficient grasping from rgbd images: Learning using a new rectangle representation
Y. Jiang, S. Moseson, and A. Saxena · 2011
Earlier work this paper cites.
Efficient grasping from rgbd images: Learning using a new rectangle representation
Y. Jiang, S. Moseson, and A. Saxena · 2011
Earlier work this paper cites.
Fast graspability evaluation on single depth maps for bin picking with general grippers
Y. Domae, H. Okuda, Y. Taguchi, K. Sumi, and T. Hirai · 2014
Earlier work this paper cites.
Deep learning for detecting robotic grasps
I. Lenz, H. Lee, and A. Saxena · 2015
Earlier work this paper cites.
Grasp quality measures: review and performance
M. A. Roa and R. Suárez · 2015
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
Jacquard: A large scale dataset for robotic grasp detection
A. Depierre, E. Dellandréa, and L. Chen · 2018
Earlier work this paper cites.
Interactive text2pickup networks for natural language-based human–robot collaboration
H. Ahn, S. Choi, N. Kim, G. Cha, and S. Oh · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Learning robust, real-time, reactive robotic grasping
D. Morrison, P. Corke, and J. Leitner · 2020
Earlier work this paper cites.
Antipodal robotic grasping using generative residual convolutional neural network
S. Kumra, S. Joshi, and F. Sahin · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
H.-S. Fang, C. Wang, M. Gou, and C. Lu · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
A joint network for grasp detection conditioned on natural language commands
Y. Chen, R. Xu, Y. Lin, and P. A. Vela · 2021
Earlier work this paper cites.
Attribute-based robotic grasping with one-grasp adaptation
Y. Yang, Y. Liu, H. Liang, X. Lou, and C. Choi · 2021
Earlier work this paper cites.
Invigorate: Interactive visual grounding and grasping in clutter
H. Zhang, Y. Lu, C. Yu, D. Hsu, X. La, and N. Zheng · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
End-to-end trainable deep neural network for robotic grasp detection and semantic segmentation from rgb
S. Ainetter and F. Fraundorfer · 2021
Cited alongside, same era.
Acronym: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
VL-Grasp: a 6-Dof Interactive Grasp Policy for Language-Oriented Objects in Cluttered Indoor Scenes
Y. Lu, Y. Fan, B. Deng, F. Liu, Y. Li, and S. Wang · 2023
Later among the works it cites.
Language Guided Robotic Grasping with Fine-Grained Instructions
Q. Sun, H. Lin, Y. Fu, Y. Fu, and X. Xue · 2023
Later among the works it cites.
Task-oriented grasp prediction with visual-language inputs
C. Tang, D. Huang, L. Meng, W. Liu, and H. Zhang · 2023
Later among the works it cites.
A joint modeling of vision-language-action for target-oriented grasping in clutter
K. Xu, S. Zhao, Z. Zhou, Z. Li, H. Pi, Y. Zhu, Y. Wang, and R. Xiong · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Learning 6-dof object poses to grasp category-level objects by language instructions
C. Cheang, H. Lin, Y. Fu, and X. Xue · 2022
Cited alongside, same era.
Human-in-the-loop robotic grasping using bert scene representation
Y. Song, P. Sun, P. Fang, L. Yang, Y. Xiao, and Y. Zhang · 2022
Cited alongside, same era.
Interactive robotic grasping with attribute-guided disambiguation
Y. Yang, X. Lou, and C. Choi · 2022
Cited alongside, same era.
Metagraspnet: A large-scale benchmark dataset for scene-aware ambidextrous bin picking via physics-based metaverse synthesis
M. Gilles, Y. Chen, T. R. Winter, E. Z. Zeng, and A. Wong · 2022
Cited alongside, same era.
Language-guided robot grasping: Clip-based referring grasp synthesis in clutter
G. Tziafas, X. Yucheng, A. Goel, M. Kasaei, Z. Li, and H. Kasaei · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Cited alongside, same era.
Later among the works it cites.
Graspgpt: Leveraging semantic knowledge from a large language model for task-oriented grasping
C. Tang, D. Huang, W. Ge, W. Liu, and H. Zhang · 2023
Later among the works it cites.
Lan-grasp: Using large language models for semantic object grasping
R. Mirjalili, M. Krawez, S. Silenzi, Y. Blei, and W. Burgard · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Grasp-anything: Large-scale grasp dataset from foundation models
A. D. Vuong, M. N. Vu, H. Le, B. Huang, B. Huynh, T. Vo, A. Kugi, and A. Nguyen · 2023
Later among the works it cites.
Rt-grasp: Reasoning tuning robotic grasping via multi-modal large language model
J. Xu, S. Jin, Y. Lei, Y. Zhang, and L. Zhang · 2024
Closest in time.
Interactive planning using large language models for partially observable robotic tasks
L. Sun, D. K. Jha, C. Hori, S. Jain, R. Corcodel, X. Zhu, M. Tomizuka, and D. Romeres · 2024
Closest in time.
Rlingua: Improving reinforcement learning sample efficiency in robotic manipulations with large language models
L. Chen, Y. Lei, S. Jin, Y. Zhang, and L. Zhang · 2024
Closest in time.
Llava-1.6: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Scaling open-vocabulary object detection
M. Minderer, A. Gritsenko, and N. Houlsby · 2024
Closest in time.