Fetching the paper…
Reading the bibliography…
The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Image question answering: A visual semantic embedding model and a new dataset
Mengye Ren, Ryan Kiros, and Richard S. Zemel · 2015
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Faster r-cnn features for instance search
Amaia Salvador, Xavier Giró-i Nieto, Ferran Marqués, and Shin’ichi Satoh · 2016
Earlier work this paper cites.
Deep sliding shapes for amodal 3d object detection in rgb-d images, 2016
Shuran Song and Jianxiong Xiao · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
2d-driven 3d object detection in rgb-d images
Jean Lahoud and Bernard Ghanem · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee · 2017
Earlier work this paper cites.
A joint speaker-listener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Joint 3d proposal generation and object detection from view aggregation, 2018
Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven Waslander · 2018
Cited alongside, same era.
Frustum pointnets for 3d object detection from rgb-d data
Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas · 2018
Cited alongside, same era.
Pointfusion: Deep sensor fusion for 3d bounding box estimation, 2018
Danfei Xu, Dragomir Anguelov, and Ashesh Jain · 2018
Cited alongside, same era.
Imvotenet: Boosting 3d object detection in point clouds with image votes, 2020
Charles R. Qi, Xinlei Chen, Or Litany, and Leonidas J. Guibas · 2020
Later among the works it cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Later among the works it cites.
Learning 3d semantic scene graphs from 3d indoor reconstructions
Johanna Wald, Helisa Dhamo, Nassir Navab, and Federico Tombari · 2020
Later among the works it cites.
Improving one-stage visual grounding by recursive sub-query construction
Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo · 2020
Later among the works it cites.
D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans, 2021
Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, and Angel X. Chang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L. Berg · 2018
Cited alongside, same era.
3d-sis: 3d semantic instance segmentation of rgb-d scans, 2019
Ji Hou, Angela Dai, and Matthias Nießner · 2019
Cited alongside, same era.
Pointpillars: Fast encoders for object detection from point clouds
Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, K. Li, Yuanyuan Li, and Yun Raymond Fu · 2019
Cited alongside, same era.
Learning cross-modal context graph for visual grounding
Yongfei Liu, Bo Wan, Xiaodan Zhu, and Xuming He · 2019
Cited alongside, same era.
Zero-shot grounding of objects from natural language queries
Arka Sadhu, Kan Chen, and Ram Nevatia · 2019
Cited alongside, same era.
Densefusion: 6d object pose estimation by iterative dense fusion, 2019
Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martín-Martín, Cewu Lu, Li Fei-Fei, and Silvio Savarese · 2019
Cited alongside, same era.
Scan2cap: Context-aware dense captioning in rgb-d scans
Zhenyu Chen, Ali Gholami, Matthias Niessner, and Angel X. Chang · 2021
Later among the works it cites.
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li · 2021
Later among the works it cites.
Free-form description guided 3d visual graph network for object grounding in point cloud
Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian · 2021
Later among the works it cites.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, and Si Liu · 2021
Later among the works it cites.
Text-guided graph neural networks for referring 3d instance segmentation
Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu · 2021
Later among the works it cites.
Looking outside the box to ground language in 3d scenes
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki · 2021
Later among the works it cites.
Looking outside the box to ground language in 3d scenes
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki · 2021
Later among the works it cites.
Tap: Text-aware pre-training for text-vqa and text-caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2021
Later among the works it cites.
Sat: 2d semantics assisted training for 3d visual grounding
Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo · 2021
Later among the works it cites.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui · 2021
Later among the works it cites.
3dvg-transformer: Relation modeling for visual grounding on point clouds
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Later among the works it cites.
3dreftransformer: Fine-grained object identification in real-world scenes using natural language
Ahmed Abdelreheem, Ujjwal Upadhyay, Ivan Skorokhodov, Rawan Al Yahya, Jun Chen, and Mohamed Elhoseiny · 2022
Closest in time.
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Closest in time.
Languagerefer: Spatial-language model for 3d visual grounding
Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox · 2022
Closest in time.