Fetching the paper…
Reading the bibliography…
Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus.
The interpretation of structure from motion
Shimon Ullman · 1979
Earlier work this paper cites.
Ewa volume splatting
Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross · 2001
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration
Angela Dai, Matthias Nießner, Michael Zollhöfer, Shahram Izadi, and Christian Theobalt · 2017
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
The replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al · 2019
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Earlier work this paper cites.
Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges
Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer · 2020
Earlier work this paper cites.
Mmnet: Multi-stage and multi-scale fusion network for rgb-d salient object detection
Guibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang, and Ge Li · 2020
Earlier work this paper cites.
Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields
Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan · 2021
Earlier work this paper cites.
From multi-view to hollow-3d: Hallucinated hollow-3d r-cnn for 3d object detection
Jiajun Deng, Wengang Zhou, Yanyong Zhang, and Houqiang Li · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Point-based neural rendering with per-view optimization
Georgios Kopanas, Julien Philip, Thomas Leimkühler, and George Drettakis · 2021
Earlier work this paper cites.
Tensorf: Tensorial radiance fields
Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su · 2022
Earlier work this paper cites.
Instant neural graphics primitives with a multiresolution hash encoding
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller · 2022
Earlier work this paper cites.
Nerf: Neural radiance field in 3d vision, a comprehensive review
Kyle Gao, Yina Gao, Hongjie He, Dening Lu, Linlin Xu, and Jonathan Li · 2022
Cited alongside, same era.
Tdrnet: Transformer-based dual-branch restoration network for geometry based point cloud compression artifacts
Xiaoyu Zhang, Guibiao Liao, Wei Gao, and Ge Li · 2022
Cited alongside, same era.
Language-driven semantic segmentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl · 2022
Cited alongside, same era.
Language-augmented pixel embedding for generalized zero-shot learning
Ziyang Wang, Yunhao Gou, Jingjing Li, Lei Zhu, and Heng Tao Shen · 2022
Cited alongside, same era.
Decomposing nerf for editing via feature field distillation
Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann · 2022
Cited alongside, same era.
Lerf: Language embedded radiance fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik · 2023
Later among the works it cites.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2023
Later among the works it cites.
Language embedded 3d gaussians for open-vocabulary scene understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao-Hua Guan · 2023
Later among the works it cites.
Tracking anything with decoupled video segmentation
Ho Kei Cheng, Seoung Wug Oh, Brian Price, Alexander Schwing, and Joon-Young Lee · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis · 2023
Cited alongside, same era.
Zip-nerf: Anti-aliased grid-based neural radiance fields
Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman · 2023
Cited alongside, same era.
Visual language maps for robot navigation
Chenguang Huang, Oier Mees, Andy Zeng, and Wolfram Burgard · 2023
Cited alongside, same era.
Conceptfusion: Open-set multimodal 3d mapping
Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu, Mohd Omama, Ganesh Iyer, Soroush Saryazdi, Tao Chen, Alaa Maalouf, Shuang Li, Nikhil Varma Keetha, Ayush Tewari, Joshua Tenenbaum, Celso de Melo, Madhava Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba · 2023
Cited alongside, same era.
Dense object grounding in 3d scenes
Wencan Huang, Daizong Liu, and Wei Hu · 2023
Cited alongside, same era.
Cross-modal unsupervised domain adaptation for 3d semantic segmentation via bidirectional fusion-then-distillation
Yao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie, Jianping Fan, Zhongchao Shi, and Yanyun Qu · 2023
Cited alongside, same era.
Side adapter network for open-vocabulary semantic segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei, Han Hu, and Xiang Bai · 2023
Cited alongside, same era.
Zhenyu Bao, Guibiao Liao, Zhongyuan Zhao, Kanglin Liu, Qing Li, and Guoping Qiu · 2024
Closest in time.
Guibiao Liao, Kaichen Zhou, Zhenyu Bao, Kanglin Liu, and Qing Li · 2024
Closest in time.
Clipself: Vision transformer distills itself for open-vocabulary dense prediction
Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Xiangtai Li, Wentao Liu, and Chen Change Loy · 2024
Closest in time.
Vlm2scene: Self-supervised image-text-lidar learning with foundation models for autonomous driving scene understanding
Guibiao Liao, Jiankun Li, and Xiaoqing Ye · 2024
Closest in time.
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Zehao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi · 2024
Closest in time.
Dreamgaussian: Generative gaussian splatting for efficient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng · 2024
Closest in time.
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Wang Xinggang · 2024
Closest in time.
Gaussian splatting slam
Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, and Andrew J. Davison · 2024
Closest in time.
Fmgs: Foundation model embedded 3d gaussian splatting for holistic 3d scene understanding
Xingxing Zuo, Pouya Samangouei, Yunwen Zhou, Yan Di, and Mingyang Li · 2024
Closest in time.
Langsplat: 3d language gaussian splatting
Minghan Qin, Wanhua Li, Jiawei Zhou, Haoqian Wang, and Hanspeter Pfister · 2024
Closest in time.