Fetching the paper…
Reading the bibliography…
Robots operating in human-centric environments require the integration of visual grounding and grasping capabilities to effectively manipulate objects based on user instructions.
A fast and accurate one-stage approach to visual grounding
Z. Yang, B. Gong, L. Wang, W. Huang, D. Yu, and J. Luo · 1908
Earlier work this paper cites.
Design and use paradigms for gazebo, an open-source multi-robot simulator
N. P. Koenig and A. Howard · 2004
Earlier work this paper cites.
Efficient grasping from rgbd images: Learning using a new rectangle representation
Y. Jiang, S. Moseson, and A. Saxena · 2011
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
D. Misra, J. Sung, K. Lee, and A. Saxena · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. andre Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O.-M. Camburu, A. L. Yuille, and K. P. Murphy · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O.-M. Camburu, A. L. Yuille, and K. P. Murphy · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2015
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2015
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
H. de Vries, F. Strub, A. P. S. Chandar, O. Pietquin, H. Larochelle, and A. C. Courville · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Interactively picking real-world objects with unconstrained spoken language instructions
J. Hatori, Y. Kikuchi, S. Kobayashi, K. Takahashi, Y. Tsuboi, Y. Unno, W. K. H. Ko, and J. Tan · 2017
Earlier work this paper cites.
Comprehension-guided referring expressions
R. Luo and G. Shakhnarovich · 2017
Earlier work this paper cites.
Shape completion enabled robotic grasping
J. Varley, C. DeChant, A. Richardson, J. Ruales, and P. Allen · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach
D. Morrison, P. Corke, and J. Leitner · 2018
Earlier work this paper cites.
Interactive visual grounding of referring expressions for human-robot interaction
M. Shridhar and D. Hsu · 2018
Cited alongside, same era.
Jacquard: A large scale dataset for robotic grasp detection
A. Depierre, E. Dellandréa, and L. Chen · 2018
Cited alongside, same era.
Antipodal robotic grasping using generative residual convolutional neural network
S. Kumra, S. Joshi, and F. Sahin · 2019
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2019
Cited alongside, same era.
Sun-spot: An rgb-d dataset with spatial referring expressions
C. Mauceri, M. Palmer, and C. Heckman · 2019
Cited alongside, same era.
Ocid-ref: A 3d robotic dataset with embodied language for clutter scene grounding
K.-J. Wang, Y.-H. Liu, H.-T. Su, J.-W. Wang, Y.-S. Wang, W. H. Hsu, and W.-C. Chen · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Later among the works it cites.
Encoder fusion network with co-attention embedding for referring image segmentation
G. Feng, Z. Hu, L. Zhang, and H. Lu · 2021
Later among the works it cites.
Referring transformer: A one-step approach to multi-task visual grounding
M. Li and L. Sigal · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey of reinforcement learning informed by natural language
J. Luketina, N. Nardelli, G. Farquhar, J. N. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel · 2019
Cited alongside, same era.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
R. Liu, C. Liu, Y. Bai, and A. L. Yuille · 2019
Cited alongside, same era.
Zero-shot grounding of objects from natural language queries
A. Sadhu, K. Chen, and R. Nevatia · 2019
Cited alongside, same era.
Cross-modal relationship inference for grounding referring expressions
S. Yang, G. Li, and Y. Yu · 2019
Cited alongside, same era.
Easylabel: A semi-automatic pixel-wise object annotation tool for creating robotic rgb-d datasets
M. Suchi, T. Patten, and M. Vincze · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Language-conditioned imitation learning for robot manipulation tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. B. Amor · 2020
Cited alongside, same era.
3dvg-transformer: Relation modeling for visual grounding on point clouds
L. Zhao, D. Cai, L. Sheng, and D. Xu · 2021
Later among the works it cites.
Mdetr - modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, I. Misra, G. Synnaeve, and N. Carion · 2021
Later among the works it cites.
Acronym: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2021
Later among the works it cites.
Vision-language transformer and query generation for referring segmentation
H. Ding, C. Liu, S. Wang, and X. Jiang · 2021
Later among the works it cites.
Meta-learning regrasping strategies for physical-agnostic objects
R. Chen, N. Gao, N. A. Vien, H. Ziesche, and G. Neumann · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. S. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. R. Florence · 2022
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. R. Florence, and A. Zeng · 2022
Later among the works it cites.
Cris: Clip-driven referring image segmentation
Z. Wang, Y. Lu, Q. Li, X. Tao, Y. Guo, M. Gong, and T. Liu · 2022
Later among the works it cites.
Multi-view transformer for 3d visual grounding
S. Huang, Y. Chen, J. Jia, and L. Wang · 2022
Later among the works it cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Later among the works it cites.
Deep learning approaches to grasp synthesis: A review
R. Newbury, M. Gu, L. Chumbley, A. Mousavian, C. Eppner, J. Leitner, J. Bohg, A. Morales, T. Asfour, D. Kragic, et al · 2022
Later among the works it cites.
Reclip: A strong zero-shot baseline for referring expression comprehension
S. Subramanian, W. Merrill, T. Darrell, M. Gardner, S. Singh, and A. Rohrbach · 2022
Later among the works it cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2022
Later among the works it cites.
Mvgrasp: Real-time multi-view 3d object grasping in highly cluttered environments
H. Kasaei and M. Kasaei · 2023
Closest in time.
Task-oriented grasp prediction with visual-language inputs
C. Tang, D. Huang, L. Meng, W. Liu, and H. Zhang · 2023
Closest in time.
Instance-wise grasp synthesis for robotic grasping
Y. Xu, M. M. Kasaei, S. H. M. Kasaei, and Z. Li · 2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Closest in time.