Fetching the paper…
Reading the bibliography…
3D visual grounding aims at grounding a natural language description about a 3D scene, usually represented in the form of 3D point clouds, to the targeted object region.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Faster r-cnn features for instance search
Amaia Salvador, Xavier Giró-i Nieto, Ferran Marqués, and Shin’ichi Satoh · 2016
Earlier work this paper cites.
Deep sliding shapes for amodal 3d object detection in rgb-d images
Shuran Song and Jianxiong Xiao · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Earlier work this paper cites.
11-1: Invited paper: Towards the ultimate mixed reality experience: Hololens display architecture choices
Bernard C Kress and William J Cummings · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
2d-driven 3d object detection in rgb-d images
Jean Lahoud and Bernard Ghanem · 2017
Earlier work this paper cites.
Comprehension-guided referring expressions
Ruotian Luo and Gregory Shakhnarovich · 2017
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Split-brain autoencoders: Unsupervised learning by cross-channel prediction
Richard Zhang, Phillip Isola, and Alexei A Efros · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Revisiting trends in augmented reality research: A review of the 2nd decade of ismar (2008–2017)
Kangsoo Kim, Mark Billinghurst, Gerd Bruder, Henry Been-Lirn Duh, and Gregory F Welch · 2018
Cited alongside, same era.
Joint 3d proposal generation and object detection from view aggregation
Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven L Waslander · 2018
Cited alongside, same era.
Frustum pointnets for 3d object detection from rgb-d data
Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas · 2018
Cited alongside, same era.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Later among the works it cites.
Iterative answer prediction with pointer-augmented multimodal transformers for textvqa
Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach · 2020
Later among the works it cites.
Pointgroup: Dual-set point grouping for 3d instance segmentation
Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia · 2020
Later among the works it cites.
Attngrounder: Talking to cars with attention
Vivek Mittal · 2020
Later among the works it cites.
Imvotenet: Boosting 3d object detection in point clouds with image votes
Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas · 2020
Later among the works it cites.
Video object grounding using semantic roles in language description
Arka Sadhu, Kan Chen, and Ram Nevatia · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik · 2018
Cited alongside, same era.
Gibson env: Real-world perception for embodied agents
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Cited alongside, same era.
Pointfusion: Deep sensor fusion for 3d bounding box estimation
Danfei Xu, Dragomir Anguelov, and Ashesh Jain · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
3d scene graph: A structure for unified semantics, 3d space, and camera
Iro Armeni, Zhi-Yang He, JunYoung Gwak, Amir R Zamir, Martin Fischer, Jitendra Malik, and Silvio Savarese · 2019
Cited alongside, same era.
Scaling and benchmarking self-supervised visual representation learning
Priya Goyal, Dhruv Mahajan, Abhinav Gupta, and Ishan Misra · 2019
Cited alongside, same era.
3d-sis: 3d semantic instance segmentation of rgb-d scans
Ji Hou, Angela Dai, and Matthias Nießner · 2019
Cited alongside, same era.
Later among the works it cites.
Learning 3d semantic scene graphs from 3d indoor reconstructions
Johanna Wald, Helisa Dhamo, Nassir Navab, and Federico Tombari · 2020
Later among the works it cites.
Improving one-stage visual grounding by recursive sub-query construction
Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo · 2020
Later among the works it cites.
Grounding-tracking-integration
Zhengyuan Yang, Tushar Kumar, Tianlang Chen, Jingsong Su, and Jiebo Luo · 2020
Later among the works it cites.
Where does it exist: Spatio-temporal video grounding for multi-form sentences
Zhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang, Huasheng Liu, and Lianli Gao · 2020
Later among the works it cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Dave Zhenyu Chen, Ali Gholami, Matthias Nießner, and Angel X Chang · 2021
Closest in time.
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Closest in time.
Free-form description guided 3d visual graph network for object grounding in point cloud
Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian · 2021
Closest in time.
Cityflow-nl: Tracking and retrieval of vehicles at city scaleby natural language descriptions
Qi Feng, Vitaly Ablavsky, and Stan Sclaroff · 2021
Closest in time.
Text-guided graph neural networks for referring 3d instance segmentation
Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu · 2021
Closest in time.
Languagerefer: Spatial-language model for 3d visual grounding
Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox · 2021
Closest in time.
Tap: Text-aware pre-training for text-vqa and text-caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2021
Closest in time.
Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Zhen Li, and Shuguang Cui · 2021
Closest in time.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, Si Liu, et al · 2021
Closest in time.