Fetching the paper…
Reading the bibliography…
Object grounding tasks aim to locate the target object in an image through verbal communications.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Pragmatic aspects of representational gestures: Do speakers use them to clarify verbal ambiguity for the listener?
Judith Holler and Geoffrey Beattie. 2003 · 2003
Earlier work this paper cites.
Image retrieval using scene graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3668–3678
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval. In Proceedings of the fourth workshop on vision and language . 70–80
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Spatial references and perspective in natural language instructions for collaborative manipulation. In 2016 25th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IEEE, 44–51
Shen Li, Rosario Scalise, Henny Admoni, Stephanie Rosenthal, and Siddhartha S Srinivasa. 2016 · 2016
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1115–1124
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Graph-structured representations for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1–9
Damien Teney, Lingqiao Liu, and Anton van Den Hengel. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Reducing errors in object-fetching interactions through social feedback. In 2017 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 1006–1013
David Whitney, Eric Rosen, James MacGlashan, Lawson LS Wong, and Stefanie Tellex. 2017 · 2017
Cited alongside, same era.
Scene graph generation by iterative message passing. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5410–5419
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. 2017 · 2017
Cited alongside, same era.
Interactively picking real-world objects with unconstrained spoken language instructions. In 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 3774–3781
Jun Hatori, Yuta Kikuchi, Sosuke Kobayashi, Kuniyuki Takahashi, Yuta Tsuboi, Yuya Unno, Wilson Ko, and Jethro Tan. 2018 · 2018
Cited alongside, same era.
Interactive visual grounding of referring expressions for human-robot interaction
Mohit Shridhar and David Hsu. 2018 · 2018
Cited alongside, same era.
Understanding natural language instructions for fetching daily objects using GAN-based multimodal target–source classification
Aly Magassouba, Komei Sugiura, Anh Trinh Quoc, and Hisashi Kawai. 2019 · 2019
Later among the works it cites.
Miscommunication detection and recovery in situated human–robot dialogue
Matthew Marge and Alexander I Rudnicky. 2019 · 2019
Later among the works it cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Relationship-embedded representation learning for grounding referring expressions
Sibei Yang, Guanbin Li, and Yizhou Yu. 2019b · 2019
Later among the works it cites.
Semantic Image Manipulation Using Scene Graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5213–5222
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A comparison of visualisation methods for disambiguating verbal requests in human-robot interaction. In 2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IEEE, 43–50
Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, and Danica Kragic. 2018 · 2018
Cited alongside, same era.
Graph r-cnn for scene graph generation. In Proceedings of the European conference on computer vision (ECCV) . 670–685
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning. In Proceedings of the European conference on computer vision (ECCV) . 684–699
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2018 · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1307–1315
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018 · 2018
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5831–5840
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Bohan Zhuang, Qi Wu, Chunhua Shen, Ian Reid, and Anton Van Den Hengel. 2018 · 2018
Cited alongside, same era.
Grounding Abstract Spatial Concepts for Language Interaction with Robots
Rohan Paul, Jacob Arkin, Nicholas Roy, and Thomas M Howard. [n.d.]
Cited in the paper.
Cross-modal relationship inference for grounding referring expressions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4145–4154
Sibei Yang, Guanbin Li, and Yizhou Yu. 2019a
Cited in the paper.
Helisa Dhamo, Azade Farshad, Iro Laina, Nassir Navab, Gregory D Hager, Federico Tombari, and Christian Rupprecht. 2020 · 2020
Later among the works it cites.
The impact of adding perspective-taking to spatial referencing during human–robot interaction
Fethiye Irmak Doğan, Sarah Gillet, Elizabeth J Carter, and Iolanda Leite. 2020 · 2020
Later among the works it cites.
Scene Graph Reasoning for Visual Question Answering
Marcel Hildebrandt, Hang Li, Rajat Koner, Volker Tresp, and Stephan Günnemann. 2020 · 2020
Later among the works it cites.
Unbiased scene graph generation from biased training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3716–3725
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. 2020 · 2020
Later among the works it cites.
Graph-structured referring expression reasoning in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9952–9961
Sibei Yang, Guanbin Li, and Yizhou Yu. 2020 · 2020
Later among the works it cites.
Open Challenges on Generating Referring Expressions for Human-Robot Interaction
Fethiye Irmak Doğan and Iolanda Leite. 2021 · 2021
Later among the works it cites.