Fetching the paper…
Reading the bibliography…
3D referring segmentation is an emerging and challenging vision-language task that aims to segment the object described by a natural language expression in a point cloud scene.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions. In CVPR . 11–20
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions. In ECCV . Springer, 69–85
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016 · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR . 5828–5839
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks. In CVPR
Christopher Choy, JunYoung Gwak, and Silvio Savarese. 2019 · 2019
Earlier work this paper cites.
Panoptic segmentation. In CVPR . 9404–9413
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. In ECCV . Springer, 422–440
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas. 2020 · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language. In ECCV . Springer, 202–221
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner. 2020 · 2020
Earlier work this paper cites.
Phraseclick: toward achieving flexible interactive segmentation by phrase and click. In ECCV . Springer, 417–435
Henghui Ding, Scott Cohen, Brian Price, and Xudong Jiang. 2020 · 2020
Earlier work this paper cites.
Pointgroup: Dual-set point grouping for 3d instance segmentation. In CVPR . 4867–4876
Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia. 2020 · 2020
Earlier work this paper cites.
Multi-task collaborative network for joint referring expression comprehension and segmentation. In CVPR . 10034–10043
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. 2020 · 2020
Earlier work this paper cites.
Vision-language transformer and query generation for referring segmentation. In ICCV . 16321–16330
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. 2021 · 2021
Earlier work this paper cites.
Free-form description guided 3d visual graph network for object grounding in point cloud. In ICCV . 3722–3731
Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian. 2021 · 2021
Earlier work this paper cites.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding. In ACM MM . 2344–2352
Dailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui, Shaofei Huang, Aixi Zhang, and Si Liu. 2021 · 2021
Earlier work this paper cites.
Text-guided graph neural networks for referring 3d instance segmentation. In AAAI . 1610–1618
Pin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, and Tyng-Luh Liu. 2021 · 2021
Earlier work this paper cites.
Referring transformer: A one-step approach to multi-task visual grounding
Muchen Li and Leonid Sigal. 2021 · 2021
Earlier work this paper cites.
An end-to-end transformer model for 3d object detection. In CVPR . 2906–2917
Ishan Misra, Rohit Girdhar, and Armand Joulin. 2021 · 2021
Cited alongside, same era.
Sat: 2d semantics assisted training for 3d visual grounding. In ICCV . 1856–1866
Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo. 2021 · 2021
Cited alongside, same era.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. In CVPR . 1791–1800
Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui. 2021 · 2021
Cited alongside, same era.
3DVG-Transformer: Relation modeling for visual grounding on point clouds. In ICCV . 2928–2937
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu. 2021 · 2021
Cited alongside, same era.
3dreftransformer: Fine-grained object identification in real-world scenes using natural language. In WACV . 3941–3950
Ahmed Abdelreheem, Ujjwal Upadhyay, Ivan Skorokhodov, Rawan Al Yahya, Jun Chen, and Mohamed Elhoseiny. 2022 · 2022
GREC: Generalized referring expression comprehension
Shuting He, Henghui Ding, Chang Liu, and Xudong Jiang. 2023c · 2023
Later among the works it cites.
Prototype adaption and projection for few-and zero-shot 3d point cloud semantic segmentation
Shuting He, Xudong Jiang, Wei Jiang, and Henghui Ding. 2023d · 2023
Later among the works it cites.
Ns3d: Neuro-symbolic grounding of 3d objects and relations. In CVPR . 2614–2623
Joy Hsu, Jiayuan Mao, and Jiajun Wu. 2023 · 2023
Later among the works it cites.
Mask-Attention-Free Transformer for 3D Instance Segmentation. In ICCV
Xin Lai, Yuhui Yuan, Ruihang Chu, Yukang Chen, Han Hu, and Jiaya Jia. 2023 · 2023
Later among the works it cites.
Multi-modal mutual attention and iterative interaction for referring image segmentation
Chang Liu, Henghui Ding, Yulun Zhang, and Xudong Jiang. 2023b · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Look around and refer: 2d synthetic semantics knowledge distillation for 3d visual grounding. In NeurIPS . 37146–37158
Eslam Bakr, Yasmeen Alsaedy, and Mohamed Elhoseiny. 2022 · 2022
Cited alongside, same era.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds. In CVPR . 16464–16473
Daigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng, and Dong Xu. 2022 · 2022
Cited alongside, same era.
D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding. In ECCV . Springer, 487–505
Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, and Angel X Chang. 2022 · 2022
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation. In CVPR . 1290–1299
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. 2022 · 2022
Cited alongside, same era.
VLT: Vision-language transformer and query generation for referring segmentation
Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. 2022 · 2022
Cited alongside, same era.
Multi-view transformer for 3d visual grounding. In CVPR . 15524–15533
Shijia Huang, Yilun Chen, Jiaya Jia, and Liwei Wang. 2022 · 2022
Cited alongside, same era.
Bottom up top down detection transformers for language grounding in images and point clouds. In ECCV . Springer, 417–433
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki. 2022 · 2022
Cited alongside, same era.
Mask3D: Mask Transformer for 3D Semantic Instance Segmentation. In ICRA . IEEE, 8216–8223
Jonas Schult, Francis Engelmann, Alexander Hermans, Or Litany, Siyu Tang, and Bastian Leibe. 2023 · 2023
Later among the works it cites.
Superpoint transformer for 3d scene instance segmentation. In AAAI . 2393–2401
Jiahao Sun, Chunmei Qing, Junpeng Tan, and Xiangmin Xu. 2023 · 2023
Later among the works it cites.
EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding. In CVPR . 19231–19242
Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, and Jian Zhang. 2023 · 2023
Later among the works it cites.
Multi3DRefer: Grounding Text Description to Multiple 3D Objects. In ICCV . 15225–15236
Yiming Zhang, ZeMing Gong, and Angel X Chang. 2023 · 2023
Later among the works it cites.
3d-vista: Pre-trained transformer for 3d vision and text alignment. In ICCV
Ziyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng, Siyuan Huang, and Qing Li. 2023 · 2023
Later among the works it cites.
Decoupling static and hierarchical motion perception for referring video segmentation. In CVPR . 13332–13341
Shuting He and Henghui Ding. 2024 · 2024
Closest in time.
SegPoint: Segment Any Point Cloud via Large Language Model. In ECCV
Shuting He, Henghui Ding, Xudong Jiang, and Bihan Wen. 2024 · 2024
Closest in time.
OneFormer3D: One transformer for unified point cloud segmentation. In CVPR
Maxim Kolodiazhnyi, Anna Vorontsova, Anton Konushin, and Danila Rukhovich. 2024 · 2024
Closest in time.
Lisa: Reasoning segmentation via large language model. In CVPR
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. 2024 · 2024
Closest in time.
Transformer-based visual segmentation: A survey
Xiangtai Li, Henghui Ding, Haobo Yuan, Wenwei Zhang, Jiangmiao Pang, Guangliang Cheng, Kai Chen, Ziwei Liu, and Chen Change Loy. 2024 · 2024
Closest in time.
PrimitiveNet: decomposing the global constraints for referring segmentation
Chang Liu, Xudong Jiang, and Henghui Ding. 2024a · 2024
Closest in time.
X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural Networks. In AAAI . 4551–4559
Zhipeng Qian, Yiwei Ma, Jiayi Ji, and Xiaoshuai Sun. 2024 · 2024
Closest in time.
Context-Aware Integration of Language and Visual References for Natural Language Tracking. In CVPR
Yanyan Shao, Shuting He, Qi Ye, Yuchao Feng, Wenhan Luo, and Jiming Chen. 2024 · 2024
Closest in time.
Towards robust referring image segmentation
Jianzong Wu, Xiangtai Li, Xia Li, Henghui Ding, Yunhai Tong, and Dacheng Tao. 2024a · 2024
Closest in time.
Towards open vocabulary learning: A survey
Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, et al · 2024
Closest in time.