Fetching the paper…
Reading the bibliography…
Robots interacting with humans through natural language can unlock numerous applications such as Referring Grasp Synthesis (RGS).
Efficient grasping from rgbd images: Learning using a new rectangle representation
Yun Jiang, Stephen Moseson, and Ashutosh Saxena · 2011
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg · 2016
Earlier work this paper cites.
Modulating early visual processing by language
Harm de Vries, Florian Strub, Jérémie Mary, H. Larochelle, Olivier Pietquin, and Aaron C. Courville · 2017
Earlier work this paper cites.
Extreme clicking for efficient object annotation
Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller, and Vittorio Ferrari · 2017
Earlier work this paper cites.
Jacquard: A large scale dataset for robotic grasp detection
Amaury Depierre, Emmanuel Dellandréa, and Liming Chen · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Earlier work this paper cites.
Learning robotic grasping strategy based on natural-language object descriptions
Achyutha Bharath Rao, Krishna Krishnan, and Hongsheng He · 2018
Earlier work this paper cites.
Interactive visual grounding of referring expressions for human-robot interaction
Mohit Shridhar and David Hsu · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Earlier work this paper cites.
Sun-spot: An rgb-d dataset with spatial referring expressions
Cecilia Mauceri, Martha Palmer, and Christoffer Heckman · 2019
Earlier work this paper cites.
ReferIt3D: Neural listeners for fine-grained 3d object identification in real-world scenes
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas J. Guibas · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Earlier work this paper cites.
Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review
Guoguang Du, Kai Wang, Shiguo Lian, and Kaiyong Zhao · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
Hao-Shu Fang, Chenxi Wang, Minghao Gou, and Cewu Lu · 2020
Earlier work this paper cites.
Gsnet: Joint vehicle pose and shape reconstruction with geometrical and scene-aware supervision
Lei Ke, Shichao Li, Yanan Sun, Yu-Wing Tai, and Chi-Keung Tang · 2020
Earlier work this paper cites.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
Shuran Song, Andy Zeng, Johnny Lee, and Thomas Funkhouser · 2020
Earlier work this paper cites.
End-to-end trainable deep neural network for robotic grasp detection and semantic segmentation from rgb
Stefan Ainetter and Friedrich Fraundorfer · 2021
Earlier work this paper cites.
A joint network for grasp detection conditioned on natural language commands
Yiye Chen, Ruinian Xu, Yunzhi Lin, and Patricio A. Vela · 2021
Earlier work this paper cites.
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Earlier work this paper cites.
Acronym: A large-scale grasp dataset based on simulation
Clemens Eppner, Arsalan Mousavian, and Dieter Fox · 2021
Earlier work this paper cites.
Encoder fusion network with co-attention embedding for referring image segmentation
Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu · 2021
Earlier work this paper cites.
Rgb matters: Learning 7-dof grasp poses on monocular rgbd images
Minghao Gou, Hao-Shu Fang, Zhanda Zhu, Sheng Xu, Chenxi Wang, and Cewu Lu · 2021
Earlier work this paper cites.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, and Si Liu · 2021
Earlier work this paper cites.
Vln bert: A recurrent vision-and-language bert for navigation
Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould · 2021
Cited alongside, same era.
Refer-it-in-rgbd: A bottom-up approach for 3d visual grounding in rgbd images
Haolin Liu, Anran Lin, Xiaoguang Han, Lei Yang, Yizhou Yu, and Shuguang Cui · 2021
Cited alongside, same era.
Cross-modal progressive comprehension for referring segmentation
Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, and Guanbin Li · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Real-time deep learning approach to visual servo control and grasp detection for autonomous robotic manipulation
Eduardo Godinho Ribeiro, Raul de Queiroz Mendes, and Valdir Grassi · 2021
Cited alongside, same era.
Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes
Yuhao Lu, Yixuan Fan, Beixing Deng, Fangfu Liu, Yali Li, and Shengjin Wang · 2023
Later among the works it cites.
Uniteam: Open vocabulary mobile manipulation challenge
Andrew Melnik, Michael Büttner, Leon Harz, Lyon Brown, Gora Chand Nandi, PS Arjun, Gaurav Kumar Yadav, Rahul Kala, and Robert Haschke · 2023
Later among the works it cites.
Lan-grasp: Using large language models for semantic object grasping
Reihaneh Mirjalili, Michael Krawez, Simone Silenzi, Yannik Blei, and Wolfram Burgard · 2023
Later among the works it cites.
Towards open-world interactive disambiguation for robotic grasping
Yuchen Mo, Hanbo Zhang, and Tao Kong · 2023
Later among the works it cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning
Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Invigorate: Interactive visual grounding and grasping in clutter
Hanbo Zhang, Yunfan Lu, Cunjun Yu, David Hsu, Xuguang Lan, and Nanning Zheng · 2021
Cited alongside, same era.
3DVG-Transformer: Relation modeling for visual grounding on point clouds
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Cited alongside, same era.
3dvg-transformer: Relation modeling for visual grounding on point clouds
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter · 2022
Cited alongside, same era.
Text2pos: Text-to-point-cloud cross-modal localization
Manuel Kolmet, Qunjie Zhou, Aljosa Osep, and Laura Leal-Taix’e · 2022
Cited alongside, same era.
Hybrid physical metric for 6-dof grasp pose detection
Yuhao Lu, Beixing Deng, Zhenyu Wang, Peiyuan Zhi, Yali Li, and Shengjin Wang · 2022
Cited alongside, same era.
Later among the works it cites.
Language embedded radiance fields for zero-shot task-oriented grasping
Adam Rashid, Satvik Sharma, Chung Min Kim, Justin Kerr, Lawrence Yunliang Chen, Angjoo Kanazawa, and Ken Goldberg · 2023
Later among the works it cites.
Robots that ask for help: Uncertainty alignment for large language model planners
Allen Z. Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, Zhenjia Xu, Dorsa Sadigh, Andy Zeng, and Anirudha Majumdar · 2023
Later among the works it cites.
Distilled feature fields enable few-shot language-guided manipulation
William Shen, Ge Yang, Alan Yu, Jansen Wong, Leslie Pack Kaelbling, and Phillip Isola · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2023
Later among the works it cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su · 2023
Later among the works it cites.
Open-world object manipulation using pre-trained vision-language models
Austin Stone, Ted Xiao, Yao Lu, Keerthana Gopalakrishnan, Kuang-Huei Lee, Quan Vuong, Paul Wohlhart, Sean Kirmani, Brianna Zitkovich, Fei Xia, Chelsea Finn, and Karol Hausman · 2023
Later among the works it cites.
Language guided robotic grasping with fine-grained instructions
Qiang Sun, Haitao Lin, Ying Fu, Yanwei Fu, and Xiangyang Xue · 2023
Later among the works it cites.
Task-oriented grasp prediction with visual-language inputs
Chao Tang, Dehao Huang, Lingxiao Meng, Weiyu Liu, and Hong Zhang · 2023
Later among the works it cites.
Language-guided robot grasping: Clip-based referring grasp synthesis in clutter
Georgios Tziafas, XU Yucheng, Arushi Goel, Mohammadreza Kasaei, Zhibin Li, and Hamidreza Kasaei · 2023
Later among the works it cites.
Gpt-4v(ision) for robotics: Multimodal task planning from human demonstration
Naoki Wake, Atsushi Kanehira, Kazuhiro Sasabuchi, Jun Takamatsu, and Katsushi Ikeuchi · 2023
Later among the works it cites.
Segment every reference object in spatial and temporal spaces
Jiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu, Zehuan Yuan, and Ping Luo · 2023
Later among the works it cites.
A joint modeling of vision-language-action for target-oriented grasping in clutter
Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yifeng Zhu, Yue Wang, and Rong Xiong · 2023
Later among the works it cites.
Universal instance perception as object discovery and retrieval
Bin Yan, Yi Jiang, Jiannan Wu, Dong Wang, Zehuan Yuan, Ping Luo, and Huchuan Lu · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang · 2023
Later among the works it cites.
Gliner: Generalist model for named entity recognition using bidirectional transformer, 2023
Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois · 2023
Later among the works it cites.
Grounding llms for robot task planning using closed-loop state feedback
Vineet Bhat, Ali Umut Kaypak, Prashanth Krishnamurthy, Ramesh Karri, and Farshad Khorrami · 2024
Closest in time.
Reasoning grasping via multimodal large language model
Shiyu Jin, Jinxuan Xu, Yutian Lei, and Liangjun Zhang · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine · 2024
Closest in time.
Ok-robot: What really matters in integrating open-knowledge models for robotics
Peiqi Liu, Yaswanth Orru, Chris Paxton, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang · 2024
Closest in time.
Language-driven grasp detection
An Dinh Vuong, Minh Nhat Vu, Baoru Huang, Nghia Nguyen, Hieu Le, Thieu Vo, and Anh Nguyen · 2024
Closest in time.
Closed-loop open-vocabulary mobile manipulation with gpt-4v, 2024
Peiyuan Zhi, Zhiyuan Zhang, Muzhi Han, Zeyu Zhang, Zhitian Li, Ziyuan Jiao, Baoxiong Jia, and Siyuan Huang · 2024
Closest in time.