Fetching the paper…
Reading the bibliography…
In contrast to numerous NLP and 2D vision foundational models, learning a 3D foundational model poses considerably greater challenges.
W. E. Lorensen and H. E. Cline, “Marching cubes: A high resolution 3d surface construction algorithm,” in
1998
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
R. Jensen, A. Dahl, G. Vogiatzis, E. Tola, and H. Aanæs, “Large scale multi-view stereopsis evaluation,” in
2014
Earlier work this paper cites.
S. Song, S. P. Lichtenberg, and J. Xiao, “Sun rgb-d: A rgb-d scene understanding benchmark suite,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
I. Armeni, O. Sener, A. R. Zamir, H. Jiang, I. Brilakis, M. Fischer, and S. Savarese, “3d semantic parsing of large-scale indoor spaces,” in
2016
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,”
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,”
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever
2018
Earlier work this paper cites.
B. Graham, M. Engelcke, and L. Van Der Maaten, “3d semantic segmentation with submanifold sparse convolutional networks,” in
2018
Earlier work this paper cites.
Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolutional detection,”
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever
2019
Earlier work this paper cites.
C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: Minkowski convolutional neural networks,” in
2019
Earlier work this paper cites.
C. R. Qi, O. Litany, K. He, and L. J. Guibas, “Deep hough voting for 3d object detection in point clouds,” in
2019
Earlier work this paper cites.
M. Tatarchenko, S. R. Richter, R. Ranftl, Z. Li, V. Koltun, and T. Brox, “What do single-view 3d reconstruction networks learn?” in
2019
Earlier work this paper cites.
H. Thomas, C. R. Qi, J.-E. Deschaud, B. Marcotegui, F. Goulette, and L. J. Guibas, “Kpconv: Flexible and deformable convolution for point clouds,” in
2019
Earlier work this paper cites.
L. N. Smith and N. Topin, “Super-convergence: Very fast training of neural networks using large learning rates,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
2020
Earlier work this paper cites.
S. Xie, J. Gu, D. Guo, C. R. Qi, L. Guibas, and O. Litany, “Pointcontrast: Unsupervised pre-training for 3d point cloud understanding,” in
2020
Earlier work this paper cites.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Earlier work this paper cites.
X. Chen, H. Fan, R. Girshick, and K. He, “Improved baselines with momentum contrastive learning,”
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
K.-A. Aliev, A. Sevastopolsky, M. Kolos, D. Ulyanov, and V. Lempitsky, “Neural point-based graphics,” in
2020
Earlier work this paper cites.
J. Zheng, J. Zhang, J. Li, R. Tang, S. Gao, and Z. Zhou, “Structured3d: A large photo-realistic dataset for structured 3d modeling,” in
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in
2020
Earlier work this paper cites.
Z. Zhang, B. Sun, H. Yang, and Q. Huang, “H3dnet: 3d object detection using hybrid geometric primitives,” in
2020
Earlier work this paper cites.
S. Peng, M. Niemeyer, L. Mescheder, M. Pollefeys, and A. Geiger, “Convolutional occupancy networks,” in
2020
Earlier work this paper cites.
J. Chibane, T. Alldieck, and G. Pons-Moll, “Implicit functions in feature space for 3d shape reconstruction and completion,” in
2020
Earlier work this paper cites.
L. Liu, J. Gu, K. Zaw Lin, T.-S. Chua, and C. Theobalt, “Neural sparse voxel fields,”
2020
Earlier work this paper cites.
L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, and J. Jia, “Pointgroup: Dual-set point grouping for 3d instance segmentation,” in
2020
Earlier work this paper cites.
M. Contributors, “MMDetection3D: OpenMMLab next-generation platform for general 3D object detection,”
2020
Earlier work this paper cites.
S. Vora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” in
2020
Earlier work this paper cites.
H. Tang, Z. Liu, S. Zhao, Y. Lin, J. Lin, H. Wang, and S. Han, “Searching efficient 3d architectures with sparse point-voxel convolution,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
J. Hou, B. Graham, M. Nießner, and S. Xie, “Exploring data-efficient 3d scene understanding with contrastive scene contexts,” in
2021
Earlier work this paper cites.
L. Jiang, S. Shi, Z. Tian, X. Lai, S. Liu, C.-W. Fu, and J. Jia, “Guided point contrastive learning for semi-supervised point cloud semantic segmentation,” in
2021
Earlier work this paper cites.
S. Huang, Y. Xie, S.-C. Zhu, and Y. Zhu, “Spatio-temporal self-supervised representation learning for 3d point clouds,” in
2021
Earlier work this paper cites.
Y. Rao, B. Liu, Y. Wei, J. Lu, C.-J. Hsieh, and J. Zhou, “Randomrooms: Unsupervised pre-training from synthetic shapes and randomized layouts for 3d object detection,” in
2021
Earlier work this paper cites.
Z. Zhang, R. Girdhar, A. Joulin, and I. Misra, “Self-supervised pretraining of 3d features on any point-cloud,” in
2021
Earlier work this paper cites.
H. Wang, Q. Liu, X. Yue, J. Lasenby, and M. J. Kusner, “Unsupervised point cloud pre-training via occlusion completion,” in
2021
Earlier work this paper cites.
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,”
2021
Earlier work this paper cites.
D. Park, R. Ambrus, V. Guizilini, J. Li, and A. Gaidon, “Is pseudo-lidar needed for monocular 3d object detection?” in
2021
Cited alongside, same era.
T. Wang, X. Zhu, J. Pang, and D. Lin, “Fcos3d: Fully convolutional one-stage monocular 3d object detection,” in
2021
Cited alongside, same era.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”
2021
Cited alongside, same era.
A. Yu, R. Li, M. Tancik, H. Li, R. Ng, and A. Kanazawa, “Plenoctrees for real-time rendering of neural radiance fields,” in
2021
Cited alongside, same era.
2021
Cited alongside, same era.
R. Yamada, H. Kataoka, N. Chiba, Y. Domae, and T. Ogata, “Point cloud pre-training with natural 3d structures,” in
2022
Later among the works it cites.
X. Long, C. Lin, P. Wang, T. Komura, and W. Wang, “Sparseneus: Fast generalizable neural surface reconstruction from sparse views,” in
2022
Later among the works it cites.
D. Rozenberszki, O. Litany, and A. Dai, “Language-grounded indoor 3d semantic segmentation in the wild,” in
2022
Later among the works it cites.
X. Lai, J. Liu, L. Jiang, L. Wang, H. Zhao, S. Liu, X. Qi, and J. Jia, “Stratified transformer for 3d point cloud segmentation,” in
2022
Later among the works it cites.
G. Qian, Y. Li, H. Peng, J. Mai, H. Hammoud, M. Elhoseiny, and B. Ghanem, “Pointnext: Revisiting pointnet++ with improved training and scaling strategies,”
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Oechsle, S. Peng, and A. Geiger, “Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction,” in
2021
Cited alongside, same era.
A. Yu, V. Ye, M. Tancik, and A. Kanazawa, “pixelnerf: Neural radiance fields from one or few images,” in
2021
Cited alongside, same era.
Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser, “Ibrnet: Learning multi-view image-based rendering,” in
2021
Cited alongside, same era.
C. Reiser, S. Peng, Y. Liao, and A. Geiger, “Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps,” in
2021
Cited alongside, same era.
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan, “Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields,” in
2021
Cited alongside, same era.
I. Misra, R. Girdhar, and A. Joulin, “An end-to-end transformer model for 3d object detection,” in
2021
Cited alongside, same era.
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun, “Point transformer,” in
2021
Cited alongside, same era.
X. Wu, Y. Lao, L. Jiang, X. Liu, and H. Zhao, “Point transformer v2: Grouped vector attention and partition-based pooling,”
2022
Later among the works it cites.
S. Contributors, “Spconv: Spatially sparse convolution library,”
2022
Later among the works it cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in
2022
Later among the works it cites.
L. Fan, F. Wang, N. Wang, and Z.-X. ZHANG, “Fully sparse 3d object detection,”
2022
Later among the works it cites.
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” in
2022
Later among the works it cites.
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” in
2022
Later among the works it cites.
S. Doll, R. Schulz, L. Schneider, V. Benzin, M. Enzweiler, and H. P. Lensch, “Spatialdetr: Robust scalable transformer-based 3d object detection from multi-view camera images with global cross-sensor attention,” in
2022
Later among the works it cites.
J. Lu, Z. Zhou, X. Zhu, H. Xu, and L. Zhang, “Learning ego 3d representation as ray tracing,” in
2022
Later among the works it cites.
Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao, “Deformable feature aggregation for dynamic multi-modal 3d object detection,” in
2022
Later among the works it cites.
Z. Yang, J. Chen, Z. Miao, W. Li, X. Zhu, and L. Zhang, “Deepinteraction: 3d object detection via modality interaction,”
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Wang, T. Bleja, and L. Agapito, “Go-surf: Neural feature grid optimization for fast, high-fidelity rgb-d surface reconstruction,” in
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
I. Team, “Internlm: A multilingual language model with progressively enhanced capabilities,”
2023
Closest in time.
2023
Closest in time.
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li
2023
Closest in time.
P. Zhang, X. D. B. Wang, Y. Cao, C. Xu, L. Ouyang, Z. Zhao, S. Ding, S. Zhang, H. Duan, H. Yan
2023
Closest in time.
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” in
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Yan, Z. Yang, H. Li, C. Song, L. Guan, H. Kang, G. Hua, and Q. Huang, “Implicit autoencoder for point-cloud self-supervised representation learning,” in
2023
Closest in time.
D. Huang, S. Peng, T. He, H. Yang, X. Zhou, and W. Ouyang, “Ponder: Point cloud pre-training via neural rendering,” in
2023
Closest in time.
X. Wu, X. Wen, X. Liu, and H. Zhao, “Masked scene contrast: A scalable framework for unsupervised 3d representation learning,” in
2023
Closest in time.
H. Yang, T. He, J. Liu, H. Chen, B. Wu, B. Lin, X. He, and W. Ouyang, “Gd-mae: generative decoder for mae pre-training on lidar point clouds,” in
2023
Closest in time.
A. Boulch, C. Sautier, B. Michele, G. Puy, and R. Marlet, “Also: Automotive lidar self-supervision by occupancy estimation,” in
2023
Closest in time.
2023
Closest in time.
H. Liu, Y. Teng, T. Lu, H. Wang, and L. Wang, “Sparsebev: High-performance sparse 3d object detection from multi-camera videos,” in
2023
Closest in time.
2023
Closest in time.
W. Xing, J. Chen, and Y. Guo, “Robust local light field synthesis via occlusion-aware sampling and deep visual feature fusion,”
2023
Closest in time.
2023
Closest in time.
P. Contributors, “Pointcept: A codebase for point cloud perception research,”
2023
Closest in time.
C.-Y. Wu, J. Johnson, J. Malik, C. Feichtenhofer, and G. Gkioxari, “Multiview compressive coding for 3d reconstruction,” in
2023
Closest in time.
H. Yang, W. Wang, M. Chen, B. Lin, T. He, H. Chen, X. He, and W. Ouyang, “Pvt-ssd: Single-stage 3d object detector with point-voxel transformer,” in
2023
Closest in time.
Y. Chen, J. Liu, X. Zhang, X. Qi, and J. Jia, “Voxelnext: Fully sparse voxelnet for 3d object detection and tracking,” in
2023
Closest in time.
——, “Largekernel3d: Scaling up kernels in 3d sparse cnns,” in
2023
Closest in time.
C. Shu, J. Deng, F. Yu, and Y. Liu, “3dppe: 3d point positional encoding for multi-camera 3d object detection transformers,” in
2023
Closest in time.
X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “Futr3d: A unified sensor fusion framework for 3d detection,” in
2023
Closest in time.
Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. L. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,” in
2023
Closest in time.
X. Lai, Y. Chen, F. Lu, J. Liu, and J. Jia, “Spherical transformer for lidar-based 3d recognition,” in
2023
Closest in time.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” in
2023
Closest in time.