Fetching the paper…
Reading the bibliography…
Robotic grasping faces new challenges in human-robot-interaction scenarios.
Segmentation of unknown objects in indoor environments
Andreas Richtsfeld, Thomas Mörwald, Johann Prankl, Michael Zillich, and Markus Vincze · 2012
Earlier work this paper cites.
Generating expressions that refer to visible objects
Margaret Mitchell, Kees Van Deemter, and Ehud Reiter · 2013
Earlier work this paper cites.
Voxel cloud connectivity segmentation-supervoxels for point clouds
Jeremie Papon, Alexey Abramov, Markus Schoeler, and Florentin Worgotter · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Multimodal integration learning of robot behavior using deep neural networks
Kuniaki Noda, Hiroaki Arie, Yuki Suga, and Tetsuya Ogata · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Multi-modal sensor fusion for indoor mobile robot pose estimation
Yassen Dobrev, Sergio Flores, and Martin Vossiek · 2016
Earlier work this paper cites.
Lstm: A search space odyssey
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville · 2017
Earlier work this paper cites.
Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Grasp pose detection in point clouds
Andreas ten Pas, Marcus Gualtieri, Kate Saenko, and Robert Platt · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Key-word-aware network for referring expression image segmentation
Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu · 2018
Cited alongside, same era.
Interactive visual grounding of referring expressions for human-robot interaction
Mohit Shridhar and David Hsu · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Grounding referring expressions in images by variational context
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang · 2018
Cited alongside, same era.
Gqa: a new dataset for compositional question answering over real-world images
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
A joint network for grasp detection conditioned on natural language commands
Yiye Chen, Ruinian Xu, Yunzhi Lin, and Patricio A Vela · 2021
Later among the works it cites.
Deep learning for image and point cloud fusion in autonomous driving: A review
Yaodong Cui, Ren Chen, Wenbo Chu, Long Chen, Daxin Tian, Ying Li, and Dongpu Cao · 2021
Later among the works it cites.
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Encoder fusion network with co-attention embedding for referring image segmentation
Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu · 2021
Later among the works it cites.
Dex-nerf: Using a neural radiance field to grasp transparent objects
Jeffrey Ichnowski, Yahav Avigal, Justin Kerr, and Ken Goldberg · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pointnetgpd: Detecting grasp configurations from point sets
Hongzhuo Liang, Xiaojian Ma, Shuang Li, Michael Görner, Song Tang, Bin Fang, Fuchun Sun, and Jianwei Zhang · 2019
Cited alongside, same era.
Sun-spot: An rgb-d dataset with spatial referring expressions
Cecilia Mauceri, Martha Palmer, and Christoffer Heckman · 2019
Cited alongside, same era.
Cross-modal self-attention network for referring image segmentation
Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang · 2019
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Cited alongside, same era.
Cops-ref: A new dataset and task on compositional referring expression comprehension
Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K Wong, and Qi Wu · 2020
Cited alongside, same era.
Graspnet-1billion: A large-scale benchmark for general object grasping
Hao-Shu Fang, Chenxi Wang, Minghao Gou, and Cewu Lu · 2020
Cited alongside, same era.
Self-supervised learning for alignment of objects and sound
Xinzhu Liu, Xiaoyu Liu, Di Guo, Huaping Liu, Fuchun Sun, and Haibo Min · 2020
Cited alongside, same era.
Locate then segment: A strong pipeline for referring image segmentation
Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan · 2021
Later among the works it cites.
Referring transformer: A one-step approach to multi-task visual grounding
Muchen Li and Leonid Sigal · 2021
Later among the works it cites.
Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images
Haolin Liu, Anran Lin, Xiaoguang Han, Lei Yang, Yizhou Yu, and Shuguang Cui · 2021
Later among the works it cites.
Cross-modal progressive comprehension for referring segmentation
Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, and Guanbin Li · 2021
Later among the works it cites.
Collision-aware target-driven object grasping in constrained environments
Xibai Lou, Yang Yang, and Changhyun Choi · 2021
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2021
Later among the works it cites.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
Martin Sundermeyer, Arsalan Mousavian, Rudolph Triebel, and Dieter Fox · 2021
Later among the works it cites.
Graspness discovery in clutters for fast and accurate grasp detection
Chenxi Wang, Hao-Shu Fang, Minghao Gou, Hongjie Fang, Jin Gao, and Cewu Lu · 2021
Later among the works it cites.
Ocid-ref: A 3d robotic dataset with embodied language for clutter scene grounding
Ke-Jyun Wang, Yun-Hsuan Liu, Hung-Ting Su, Jen-Wei Wang, Yu-Siang Wang, Winston Hsu, and Wen-Chin Chen · 2021
Later among the works it cites.
Invigorate: Interactive visual grounding and grasping in clutter
Hanbo Zhang, Yunfan Lu, Cunjun Yu, David Hsu, Xuguang La, and Nanning Zheng · 2021
Later among the works it cites.
Visual manipulation relationship detection based on gated graph neural network for robotic grasping
Mengyuan Ding, Yaxin Liu, Chenjie Yang, and Xuguang Lan · 2022
Later among the works it cites.
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
Hao-Shu Fang, Chenxi Wang, Hongjie Fang, Minghao Gou, Jirong Liu, Hengxu Yan, Wenhai Liu, Yichen Xie, and Cewu Lu · 2022
Later among the works it cites.
Hybrid physical metric for 6-dof grasp pose detection
Yuhao Lu, Beixing Deng, Zhenyu Wang, Peiyuan Zhi, Yali Li, and Shengjin Wang · 2022
Later among the works it cites.
Audio-visual object classification for human-robot collaboration
A Xompero, YL Pang, T Patten, A Prabhakar, B Calli, and A Cavallaro · 2022
Later among the works it cites.
Improving visual grounding with visual-linguistic verification and iterative reasoning
Li Yang, Yan Xu, Chunfeng Yuan, Wei Liu, Bing Li, and Weiming Hu · 2022
Later among the works it cites.
Overview of robotic grasp detection from 2d to 3d
Zhiyun Yin and Yujie Li · 2022
Later among the works it cites.