Fetching the paper…
Reading the bibliography…
The two popular datasets ScanRefer [16] and ReferIt3D [3] connect natural language to real-world 3D data.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Amazon mechanical turk: A research tool for organizations and information systems scholars
Kevin Crowston · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Sun3d: A database of big spaces reconstructed using sfm and object labels
Jianxiong Xiao, Andrew Owens, and Antonio Torralba · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Text to 3D Scene Generation
Angel X. Chang · 2015
Earlier work this paper cites.
ShapeNet: An information-rich 3D model repository
Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
SpRL-CWW: Spatial relation classification with independent multi-class models
Eric Nichols and Fadi Botros · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
3D semantic parsing of large-scale indoor spaces
Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese · 2016
Earlier work this paper cites.
Scenenn: A scene meshes dataset with annotations
Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Learning efficient object detection models with knowledge distillation
Guobin Chen, Wongun Choi, Xiang Yu, Tony X. Han, and Manmohan Chandraker · 2017
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Nießner · 2017
Earlier work this paper cites.
Lstm: A search space odyssey
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
PointNet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
Kevin Chen, Christopher B Choy, Manolis Savva, Angel X Chang, Thomas Funkhouser, and Silvio Savarese · 2018
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi · 2018
Cited alongside, same era.
Mattnet: Modular attention network for referring expression comprehension
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
ShapeGlot: Learning language for shape differentiation
Panos Achlioptas, Judy Fan, Robert XD Hawkins, Noah D Goodman, and Leonidas J. Guibas · 2019
Cited alongside, same era.
ScanRefer: 3D object localization in RGB-D scans using natural language
Z. Dave Chen, Angel X. Chang, and Matthias Nießner · 2019
Cited alongside, same era.
Rio: 3d object instance re-localization in changing indoor environments
Wald Johanna, Avetisyan Armen, Navab Nassir, Tombari Federico, and Niessner Matthias · 2019
Cited alongside, same era.
LanguageRefer: Spatial-language model for 3D visual grounding
Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox · 2021
Later among the works it cites.
Language grounding with 3d objects
Jesse Thomason, Mohit Shridhar, Yonatan Bisk, Chris Paxton, and Luke Zettlemoyer · 2021
Later among the works it cites.
SAT: 2d semantics assisted training for 3D visual grounding
Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo · 2021
Later among the works it cites.
InstanceRefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Zhen Li, and Shuguang Cui · 2021
Later among the works it cites.
3DVG-Transformer: Relation modeling for visual grounding on point clouds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Embodied question answering in photorealistic environments with point cloud perception
Erik Wijmans, Samyak Datta, Oleksandr Maksymets, Abhishek Das, Georgia Gkioxari, Stefan Lee, Irfan Essa, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo · 2019
Cited alongside, same era.
A fast and accurate one-stage approach to visual grounding
Zhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang, Dong Yu, and Jiebo Luo · 2019
Cited alongside, same era.
Multi-target embodied question answering
Licheng Yu, Xinlei Chen, Georgia Gkioxari, Mohit Bansal, Tamara L Berg, and Dhruv Batra · 2019
Cited alongside, same era.
ReferIt3D: Neural listeners for fine-grained 3d object identification in real-world scenes
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas J. Guibas · 2020
Cited alongside, same era.
Shapecaptioner: Generative caption network for 3d shapes by learning a mapping from parts detected in multiple views to sentences
Zhizhong Han, Chao Chen, Yu-Shen Liu, and Matthias Zwicker · 2020
Cited alongside, same era.
Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation
David M. W. Powers · 2020
Cited alongside, same era.
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Later among the works it cites.
3DRefTransformer: Fine-grained object identification in real-world scenes using natural language
Ahmed Abdelreheem, Ujjwal Upadhyay, Ivan Skorokhodov, Rawan Al Yahya, Jun Chen, and Mohamed Elhoseiny · 2022
Closest in time.
Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding
Eslam Mohamed Bakr, Yasmeen Alsaedy, and Mohamed Elhoseiny · 2022
Closest in time.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Daigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng, and Dong Xu · 2022
Closest in time.
Unit3d: A unified transformer for 3d dense captioning and visual grounding, 2022
Dave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner, and Angel X. Chang · 2022
Closest in time.
LADIS: Language disentanglement for 3D shape editing
Ian Huang, Panos Achlioptas, Tianyi Zhang, Sergey Tulyakov, Minhyuk Sung, and Guibas Leonidas · 2022
Closest in time.
Multi-view transformer for 3d visual grounding
Shijia Huang, Yilun Chen, Jiaya Jia, and Liwei Wang · 2022
Closest in time.
PartGlot: Learning shape part segmentation from language reference games
Juil Koo, Ian Huang, Panos Achlioptas, Leonidas J. Guibas, and Minhyuk Sung · 2022
Closest in time.
3d compat: Composition of materials on parts of 3d things
Yuchen Li, Ujjwal Upadhyay, Habib Slim, Ahmed Abdelreheem, Arpita Prajapati, Suhail Pothigara, Peter Wonka, and Mohamed Elhoseiny · 2022
Closest in time.
3d-sps: Single-stage 3d visual grounding via referred point progressive selection
Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, and Si Liu · 2022
Closest in time.
Language-grounded indoor 3d semantic segmentation in the wild
Dávid Rozenberszki, Or Litany, and Angela Dai · 2022
Closest in time.
Denserefer3d: A language and 3d dataset for coreference resolution and referring expression comprehension, 2022
Akshit Sharma · 2022
Closest in time.
Spatiality-guided transformer for 3d dense captioning on point clouds
Heng Wang, Chaoyi Zhang, Jianhui Yu, and Weidong (Tom) Cai · 2022
Closest in time.
Eda: Explicit text-decoupling and dense alignment for 3d visual grounding, 2022
Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, and Jian Zhang · 2022
Closest in time.
Toward explainable and fine-grained 3d grounding through referring textual phrases
Zhihao Yuan, Xu Yan, Zhuo Li, Xuhao Li, Yao Guo, Shuguang Cui, and Zhen Li · 2022
Closest in time.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Shuguang Cui, and Zhen Li · 2022
Closest in time.
Contextual modeling for 3d dense captioning on point clouds, 2022
Yufeng Zhong, Long Xu, Jiebo Luo, and Lin Ma · 2022
Closest in time.
ShapeTalk: A language dataset and framework for 3d shape edits and deformations
Panos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov, and Leonidas Guibas · 2023
Closest in time.
Sqa3d: Situated question answering in 3d scenes, 2023
Xiaojian Ma, Silong Yong, Zilong Zheng, Qing Li, Yitao Liang, Song-Chun Zhu, and Siyuan Huang · 2023
Closest in time.
Text to point cloud localization with relation-enhanced transformer
Guangzhi Wang, Hehe Fan, and Mohan Kankanhalli · 2023
Closest in time.